
A p-value on its own tells readers very little they need. Report the estimate, its interval and an exact p-value the way medical journals expect.
Three numbers, three different questions
A results section that leans on p-values answers one question, whether the data would be surprising if nothing were going on, and leaves out the answers readers actually need.
Every comparative result carries three numbers, and each does a different job:
| Number | The question it answers | What it cannot tell you |
|---|---|---|
| Effect estimate | How big is the difference or association, and in which direction? | Whether the estimate is precise |
| 95% confidence interval | Which effect sizes are reasonably compatible with these data? | The probability that the true effect lies in this particular interval |
| P-value | How unusual are these data if the null hypothesis and every model assumption were true? | The probability that the hypothesis is true, or how large or important the effect is |
The p-value is easy to misread. It is calculated assuming the null hypothesis is true, so it cannot be the probability that the null hypothesis is true. The confidence interval has its own common misreading. The 95% describes the method: across many studies, 95% of intervals computed this way would contain the true effect if all the assumptions held. Any single interval either contains the true effect or does not. Greenland and colleagues' guide to misinterpretations, published in the European Journal of Epidemiology in 2016, works through both errors in detail and is worth keeping to hand.
A useful rule for drafting: the estimate and interval carry the result; the p-value is optional supporting detail.
Why journals ask for more than a p-value
This is not one statistician's preference. The ICMJE Recommendations, which many medical journals follow, ask authors to quantify findings with indicators of uncertainty such as confidence intervals, and to avoid relying solely on hypothesis tests such as p-values, "which fail to convey important information about effect size and precision of estimates."
The American Statistical Association's 2016 statement on p-values sets out six principles. Two matter most when you write results. A p-value, or statistical significance, does not measure the size of an effect or the importance of a result. And scientific conclusions should not be based only on whether a p-value passes a specific threshold.
The practical consequence: a manuscript that reports "P < 0.05" with no estimate gives readers nothing to use, and nothing a later meta-analysis can pool.
Writing the results sentence
A complete result for a between-group comparison gives the summary for each group with its denominator, the effect estimate with units and direction, a confidence interval for that effect, and, if you report one, an exact p-value.
The figures below are illustrative, not from a real study.
Before: Pain scores were significantly lower in the intervention group (P < 0.05).
After: At 12 weeks, the mean pain score was 3.1 (SD 1.9) in the intervention group (n=112) and 3.9 (SD 2.0) in the control group (n=108); mean difference −0.8 points (95% CI −1.3 to −0.3; P = 0.003).
The second version lets a reader judge size, direction and precision without opening your tables.
Two details that randomised-trial reports often get wrong are spelled out in the explanation for CONSORT 2025 item 26:
- Give the interval for the difference, not for each group. A common error is presenting a separate confidence interval for each arm instead of one for the treatment effect.
- Report binary outcomes on both scales. For binary outcomes, CONSORT 2025 asks for both absolute and relative effect sizes, because neither alone gives a complete picture and readers tend to overestimate effects presented only in relative terms.
For observational studies, STROBE item 16 asks for unadjusted estimates and, where applicable, confounder-adjusted estimates with their precision, plus which confounders were adjusted for and why.
Exact p-values and decimal places
Report exact values. The SAMPL guidelines for basic statistical reporting prefer confidence intervals, but where you give p-values they ask for equalities, such as P = 0.03 or P = 0.22, rather than inequalities such as P < 0.05, and say not to report "NS". The smallest value that need be reported is P < 0.001, except in studies of genetic associations.
Decimal places are a matter of house style, and styles differ:
| Convention | Decimal places | Smallest value | Leading zero |
|---|---|---|---|
| SAMPL guidelines | One or two, for example P = 0.03 | P < 0.001 | Shown in its examples |
| AMA style (JAMA Network journals) | Two; three when below .01 | P < .001 | Omitted |
AMA style adds two useful rules. P-values can never equal 0 or 1, so software output of 0.000 becomes P < .001 and a value that would round to 1 becomes P > .99. And where rounding to two digits would make a value such as .046 look non-significant, keep three places.
Decision rule: check the target journal's instructions for authors first. If they are silent, give exact values to two or three decimal places, use P < 0.001 as the floor, and apply the same format in the abstract, text and tables.
Statistical versus clinical significance
The ICMJE Recommendations ask authors to distinguish between clinical and statistical significance in the discussion, and to avoid nontechnical uses of "significant". In practice, that means using "significant" only in its statistical sense and saying "statistically significant" in full when you do.
Clinical importance cannot be judged after the fact from a p-value. SAMPL asks authors to identify, where possible, the smallest difference considered clinically important. State it in the methods, ideally as it appears in your protocol, then read your confidence interval against it:
| Where the 95% CI falls | Reasonable wording |
|---|---|
| Excludes no effect, and lies entirely beyond the important threshold | Statistically significant and likely to be clinically important |
| Excludes no effect, but straddles the threshold | Statistically significant; clinical importance uncertain |
| Excludes no effect, but lies entirely below the threshold | Statistically significant but probably too small to matter |
| Includes no effect and the threshold | Inconclusive |
| Includes no effect, but excludes important effects | Evidence against an important effect |
Large studies can make trivial effects statistically significant. Small studies can leave important effects non-significant. The interval shows which situation you are in.
When the result is not statistically significant
A large p-value is not evidence that there is no effect. Especially in small studies, a real and important effect can fail to reach statistical significance.
Take a second illustrative example: 18 of 150 participants (12.0%) had the event in one group and 30 of 150 (20.0%) in the other. The risk difference is −8.0 percentage points (95% CI −16.2 to 0.2) and the risk ratio 0.60 (95% CI 0.35 to 1.03), with P = 0.06.
"There was no difference between groups" misstates this result. The interval runs from a 65% relative reduction to a 3% relative increase. A defensible sentence is: "The estimate favoured the intervention, but the confidence interval was wide and included no effect, so the trial could not rule out either a substantial benefit or a small increase in risk."
Two phrasings need more than a p-value. A claim that treatments are equivalent, or that one is non-inferior, requires a prespecified margin, which SAMPL asks you to report. And "a trend towards significance" only restates that the threshold was missed; say what the interval shows instead.
Null and inconclusive results are worth publishing well; the case for doing so is set out in why negative and null results matter.
Wording statistical reviewers flag
| What you wrote | The problem | What to write instead |
|---|---|---|
| P < 0.05, or NS | Hides the actual value | The exact p-value |
| P = 0.000 | A p-value cannot be zero | P < 0.001 |
| "Highly significant" | A p-value does not measure size or importance | The estimate and its interval |
| "Marginally significant" | Presents a near miss as a partial finding | The estimate, interval and what it includes |
| Mean ± SEM to describe the data | The standard error describes precision, not variability | Mean (SD), as SAMPL recommends |
| Significant in one subgroup, not the other | Does not show the subgroup effects differ | The difference in effect between subgroups, with its CI, from an interaction analysis |
| P-values in a trial's baseline table | With intact randomisation, baseline differences are due to chance | Descriptive statistics only |
| Overlapping CIs read as "no difference" | Intervals can overlap while the difference is statistically significant | A direct estimate of the difference |
The subgroup and baseline-table points both come from the CONSORT 2025 explanation and elaboration, which states that significance testing of baseline differences is not recommended and should not be reported. If a reviewer has already raised one of these, the simplest response is usually to fix the analysis or the wording and set out exactly what changed, as in our guide to responding to peer review comments.
What the statistical methods section must let readers check
The ICMJE standard is enough detail for a knowledgeable reader with access to the original data to judge whether the methods suit the study and to verify the results. Drawing on ICMJE and SAMPL, check the methods section against this list before submission:
- The method used for each analysis, not one list of every test in the paper
- Whether tests were one- or two-sided, with a justification for any one-sided test
- The alpha level that defines statistical significance
- Whether and how you adjusted for multiple comparisons
- How you checked each test's assumptions
- The smallest clinically important difference, where one exists
- Which analyses were prespecified and which were exploratory, including subgroups
- The statistical software and version
Senior authors are best placed to run this check, because they know which analyses were planned. Label error bars explicitly in every figure too, a point covered in what peer reviewers wish authors knew.
Before you submit
Read every results sentence and ask whether a reader could state the size of the effect and its plausible range without seeing the p-value. If not, rewrite the sentence before a reviewer asks you to.
Then check the target journal's own instructions for its statistical and number style. The Directive Publications For Authors page links to its author guidelines and article-type requirements, and when the manuscript is ready you can submit it directly.
Frequently asked questions
Should I report exact p-values or just state whether a result is significant?
Report exact p-values, such as P = 0.03 or P = 0.22, rather than thresholds like P < 0.05 or the label NS. The SAMPL guidelines and the CONSORT explanation for randomised trials both prefer actual values, with P < 0.001 as the usual floor. The p-value should accompany an effect estimate and confidence interval, not replace them.
What does a confidence interval actually tell readers?
It shows the range of effect sizes that are reasonably compatible with your data under the statistical model, and so how precise your estimate is. The 95% describes how often intervals built this way would contain the true effect across many studies if every assumption held. It does not mean there is a 95% probability that this particular interval contains the true value.
Is a non-significant result the same as no effect?
No. A p-value above 0.05 only means the data did not cross the chosen threshold for incompatibility with the null hypothesis, and they may be just as compatible with effects that would matter. If the confidence interval spans both no effect and clinically important effects, the result is inconclusive rather than negative. Only a narrow interval that excludes important effects supports saying any effect is likely to be small.
How many decimal places should p-values be reported to?
Follow the target journal's house style, because conventions differ. The SAMPL guidelines suggest one or two decimal places, while AMA style, used by JAMA Network journals, uses two digits, three digits for values below .01, and P < .001 for anything smaller. Never copy software output such as P = 0.000; write P < 0.001 instead.
Why do some journals ask authors not to rely on p-values alone?
A p-value does not measure the size of an effect or the importance of a result, and it says nothing about precision. The ICMJE Recommendations, which many medical journals follow, ask authors to present findings with measures of uncertainty such as confidence intervals and to avoid relying solely on hypothesis tests. The American Statistical Association's 2016 statement also warns against basing conclusions only on whether a p-value passes a threshold.