Medical Statistics

Reporting P-Values Responsibly in Clinical Manuscripts

Published: 2026-02-10

Reporting P-Values Responsibly in Clinical Manuscripts

A practical guide to interpreting and reporting p-values, effect sizes, and confidence intervals — and the difference between statistical and clinical significance.

Reporting P-Values Responsibly in Clinical Manuscripts

"p < 0.05" is not a conclusion. It's a single piece of evidence about one narrow question — whether the observed data would be unusual if there were truly no effect — and treating it as the finish line of an analysis is one of the most persistent and consequential errors in clinical research reporting. Statistical reviewers see it constantly: a result described as "significant" with no effect size, no confidence interval, and no discussion of whether the magnitude of the difference would matter to an actual patient.

This matters for publication, not just for statistical purity. Reporting p-values responsibly is now an explicit expectation at most well-run journals, and getting it wrong is one of the most common reasons manuscripts come back with statistical reviewer comments.

What a P-Value Actually Tells You — and What It Doesn't

A p-value is the probability of observing a result at least as extreme as the one obtained, assuming the null hypothesis is exactly true. That's the whole definition. It is not the probability that the null hypothesis is true, and it is not the probability that your finding happened by chance — both are extremely common misreadings, and both subtly overstate what a single p-value can tell you.

Just as important: a p-value carries no information about the size or clinical importance of an effect. A very large sample size can produce a tiny, clinically meaningless difference with a striking p-value of 0.001, while a smaller, well-designed study can find a large, clinically important difference that doesn't cross the conventional 0.05 threshold simply because it lacked the statistical power to detect it. The p-value tells you about the compatibility of your data with a null hypothesis — nothing about whether the result matters.

The ASA Statement: Principles Every Author Should Know

In 2016, the American Statistical Association published an official statement on statistical significance and p-values in The American Statistician, authored by Ronald Wasserstein and Nicole Lazar — the first time the ASA had taken an organizational position on the topic. It set out six principles, and while the full statement is worth reading directly, the practical takeaways for a clinical author are these:

  • A p-value indicates how incompatible the data are with a specified statistical model (usually the null hypothesis) — it is not a measure of the probability that the studied hypothesis is true, and not a measure of the probability that the data were produced by random chance alone.
  • Scientific conclusions and business or policy decisions should not be based only on whether a p-value crosses a particular threshold.
  • Proper inference requires full reporting and transparency — including disclosure of how many analyses were conducted and how the reported ones were selected.
  • A p-value does not measure the size of an effect or the importance of a result.
  • By itself, a p-value does not provide a good measure of evidence regarding a model or hypothesis.

A follow-up 2019 editorial by Wasserstein, Schirm, and Lazar, "Moving to a World Beyond p < 0.05," went further, recommending that authors stop using the words "statistically significant" altogether in favor of reporting effect sizes, confidence intervals, and p-values together, with an interpretation grounded in context rather than a bright-line cutoff. Journals increasingly reflect this in their author guidelines.

Statistical Significance Is Not Clinical Significance

These are two different questions, and conflating them is arguably the single most common statistical misstep in clinical manuscripts.

Statistical significance asks: is the observed difference unlikely to have occurred if there were truly no effect? Clinical significance asks: is the observed difference large enough to change what a clinician or patient would do? A trial with 10,000 participants might detect a statistically significant 1 mmHg difference in blood pressure between two drugs — a difference no clinician would act on. Conversely, a pilot study with 40 participants might show a 15 mmHg difference that fails to reach p < 0.05 purely because the sample was too small to have adequate power, despite the effect being potentially important if confirmed.

This is why many fields now define a minimal clinically important difference (MCID) for key outcomes in advance — a threshold, grounded in what patients or clinicians consider meaningful, against which the observed effect size is judged, independent of the p-value. Reporting your result against a pre-specified MCID, where one exists for your outcome, is far more informative to a reader than the p-value alone.

Why Effect Sizes and Confidence Intervals Matter More Than P-Values Alone

A confidence interval does something a p-value cannot: it communicates direction, magnitude, and precision all in one number range. A 95% confidence interval of 1.2 to 8.4 for a mean difference tells a reader the best estimate, the plausible range of true effects consistent with the data, and — indirectly — the statistical significance, since an interval that excludes the null value corresponds to a p-value below 0.05 at the same alpha level.

A p-value alone tells you none of that. Two studies can both report "p = 0.03" while one has a confidence interval of 1.1 to 1.2 (a precise, narrow, clinically small effect) and the other has 1.1 to 9.8 (an imprecise estimate that could represent almost no effect or a very large one). Reporting both the point estimate and the confidence interval — for every primary and secondary outcome, not just the ones that reached significance — is now standard expectation under reporting guidelines like CONSORT for trials and STROBE for observational studies.

How to Report Statistics the Way Journals Expect

  • Report exact p-values (p = 0.032) rather than only inequalities (p < 0.05), except for very small values where a threshold such as p < 0.001 is appropriate. Exact values let readers judge strength of evidence rather than a binary pass/fail.
  • Report effect sizes with confidence intervals for every outcome, not only the statistically significant ones — omitting non-significant results with their intervals is a form of selective reporting.
  • State your significance threshold and correction plan in the methods, before results are known, particularly when testing multiple outcomes or subgroups. Deciding after the fact which comparisons to correct for is a well-recognized source of bias.
  • Avoid "trend toward significance" language. A p-value of 0.06 is not a near-miss; it's simply a p-value of 0.06. Report it as such and let the confidence interval and effect size carry the interpretation.
  • Distinguish pre-specified primary analyses from exploratory ones explicitly, both in the methods and again when interpreting results in the discussion.

Common P-Value Reporting Mistakes That Draw Reviewer Comments

  • P-hacking — running many statistical tests, subgroup analyses, or outcome definitions and reporting only the ones that reached significance, without disclosing the others.
  • Dichotomizing continuous outcomes into "improved/not improved" purely to obtain a cleaner p-value, discarding information and often inflating apparent significance.
  • Multiple comparisons without correction or acknowledgment — testing ten outcomes at p < 0.05 each carries a substantially higher chance of at least one false positive than testing one, and this needs to be addressed methodologically or discussed as a limitation.
  • Describing a non-significant result as "no difference" or "no effect.” A non-significant p-value means the data don't provide strong evidence against the null — it does not confirm the null is true, especially in an underpowered study.
  • Reporting significance without the underlying test, sample size, or degrees of freedom, which prevents a reader from evaluating the analysis at all.

Getting this right isn't about statistical purism for its own sake — it's about giving readers, and the clinicians who might change practice based on your paper, an honest and complete picture of what your data actually show.

Frequently Asked Questions

What does p < 0.05 really mean?

It means that, if the null hypothesis were exactly true, a result this extreme or more extreme would occur less than 5% of the time by chance alone. It does not mean there is a 95% probability the effect is real, and it says nothing about how large or clinically important the effect is.

Is a non-significant p-value evidence that there is no effect?

No. A non-significant result means the data did not provide strong evidence against the null hypothesis — this is especially true in underpowered studies, where a real effect can easily fail to reach significance simply due to small sample size.

Should I report exact p-values or just state "significant"?

Report exact p-values (for example, p = 0.032) rather than only stating significance, except for very small values where a threshold like p < 0.001 is conventional. Exact values let readers judge the strength of evidence rather than relying on a binary cutoff.

What's the difference between clinical and statistical significance?

Statistical significance concerns whether a result is unlikely under the null hypothesis. Clinical significance concerns whether the size of the effect is large enough to matter to patient care. A result can be statistically significant but clinically trivial, or clinically important but not statistically significant in a small study.

Do I still need confidence intervals if I report p-values?

Yes. Confidence intervals convey the magnitude and precision of an effect, information a p-value alone cannot provide, and are expected by most reporting guidelines including CONSORT and STROBE for every primary and secondary outcome.

What is p-hacking?

P-hacking refers to running multiple analyses, outcome definitions, or subgroups and selectively reporting the ones that reach statistical significance, without disclosing the full set of analyses conducted. It inflates the apparent strength of evidence and is considered a serious reporting problem.

References

  1. Wasserstein RL, Lazar NA. The ASA's Statement on p-Values: Context, Process, and Purpose. The American Statistician. 2016;70(2):129–133.
  2. Wasserstein RL, Schirm AL, Lazar NA. Moving to a World Beyond "p < 0.05". The American Statistician. 2019;73(sup1):1–19.
  3. Schulz KF, Altman DG, Moher D; CONSORT Group. CONSORT 2010 Statement: updated guidelines for reporting parallel group randomised trials. Annals of Internal Medicine. 2010.
  4. von Elm E, Altman DG, Egger M, et al.; STROBE Initiative. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement. The Lancet. 2007.
WhatsApp