Quick answer
A p-value measures how incompatible the observed data are with a specified statistical model, often one assuming no difference; it is not the probability that the treatment works or that chance caused the result. A confidence interval gives a range of effect estimates compatible with the data under the analysis assumptions and helps show precision. Read both with the point estimate, units, no-effect value, prespecified threshold, clinically meaningful difference, sample size, missing data, and number of analyses performed.
Key takeaways
- ✓Statistical significance does not establish clinical importance, approval, or individual benefit.
- ✓Effect size answers how much; the confidence interval shows uncertainty around that estimate.
- ✓A wide interval can be compatible with meaningful benefit, little effect, or harm even when a headline says no difference.
- ✓Multiple outcomes, time points, subgroups, and repeated looks can increase false-positive risk unless planned and adjusted.
- ✓Use the protocol and SAP to distinguish prespecified analyses from later exploration.
01
Begin with the outcome and effect estimate
Before reading a p-value, identify what was measured, in which participants, against which comparator, at what time, and in what units. A mean difference in body weight, a risk ratio for an adverse event, a laboratory biomarker, and a patient-reported symptom score are not interchangeable. The point estimate is the study's best estimate under its analysis, but it is not a guarantee of what another person will experience.
Ask whether the outcome was primary, secondary, exploratory, or a subgroup analysis and whether it was prespecified. Relative percentages can make changes look larger than absolute differences, while change from baseline within one arm does not answer the between-group treatment question. A statistical label cannot repair a mismatch between the measured outcome and the marketing claim.
02
What a p-value can and cannot tell you
The American Statistical Association states that a p-value can indicate how incompatible data are with a specified statistical model. It does not measure the probability that the hypothesis is true, the probability that random chance alone produced the data, the size of an effect, or the importance of a result. Treat p less than 0.05 as a convention used in a defined analysis, not a universal truth switch.
The p-value depends on sample size, variability, model choices, outcome definition, and assumptions. A very large study can produce a small p-value for a clinically minor difference, while a small study can miss a potentially important effect. Read the exact value when reported and the prespecified analysis threshold. P equals 0.049 and p equals 0.051 are not radically different bodies of evidence merely because they fall on opposite sides of a convention.
03
Use confidence intervals to see effect size and precision together
A confidence interval surrounds the estimated effect with a range reflecting sampling uncertainty under the model. Narrower intervals indicate greater precision than wider ones, all else equal. For a difference in means, the no-effect value is generally zero; for ratios such as risk ratios or hazard ratios, it is generally one. The units and direction must be checked because lower can be better for one outcome and worse for another.
Do not translate a 95% confidence interval into a 95% probability that the true effect lies inside this particular interval. The frequentist interpretation concerns the long-run performance of the interval procedure under repeated samples and model assumptions. For practical reading, compare both ends with no effect and with a meaningful-benefit or harm threshold. If the range includes materially different decisions, the evidence remains imprecise even when the point estimate looks attractive.
04
Separate statistical significance from clinical importance
Clinical importance asks whether the magnitude matters to patients or decisions. It depends on the outcome, baseline risk, burdens, adverse effects, duration, alternatives, and values—not solely on a p-value. A statistically detectable change in a laboratory marker can be too small or too indirect to support a claim about function, symptoms, or long-term health.
Look for a prespecified minimal important difference or another justified threshold, but examine who defined it and for which population. For binary outcomes, absolute risk difference and number needed to treat or harm can add context to a relative ratio. Those measures also need confidence intervals and appropriate follow-up. Never turn an average group effect into a promise that each participant benefited.
05
Check multiplicity, subgroups, and missing data
Testing many outcomes, doses, time points, subgroups, and models creates more opportunities for a small p-value. A protocol and SAP should identify the primary analysis and explain adjustments or a testing hierarchy. A striking subgroup result deserves extra caution when the interaction was not prespecified, the subgroup is small, many subgroups were searched, or the main result was negative.
Missing data can change both the estimate and uncertainty. Check how many participants were randomized, completed follow-up, and entered each analysis; why data were missing; and which assumptions were used to handle them. Per-protocol analyses may exclude people who stopped treatment or deviated from the protocol, while intention-to-treat approaches generally preserve randomized groups. Neither label by itself proves the implementation was appropriate.
06
Build a result note that resists headline inflation
Write one sentence containing the comparison, population, primary outcome, time point, point estimate, 95% confidence interval, exact p-value when relevant, and main safety limitation. Then add whether the analysis was prespecified and whether the interval rules out no effect, clinically trivial effect, and important harm. This makes it harder to report only the most favorable number.
Finally, compare the result with the registry, protocol, SAP, full results, and FDA status. A statistically significant investigational result does not make a peptide approved, a compounded formulation equivalent, or a provider's broader claim established. Statistics quantify evidence under assumptions; they do not replace study design, replication, safety evaluation, product verification, or individualized clinical judgment.
- →Population and comparator
- →Primary outcome and units
- →Point estimate
- →Confidence interval and no-effect value
- →Exact p-value and threshold
- →Meaningful-difference threshold
- →Multiplicity and missing data
- →Safety and regulatory status
Common questions
Frequently asked questions
Does p less than 0.05 mean a peptide works?
No. It does not give the probability that the treatment works and does not show effect size, clinical importance, study quality, approval, or applicability to another product.
What is the no-effect value in a confidence interval?
It is usually zero for a difference and one for a ratio, but confirm the measure, direction, and units in the paper.
Is a result with p greater than 0.05 proof of no effect?
No. Review the point estimate and confidence interval. A wide interval may leave both meaningful benefit and harm compatible with the data, making the result inconclusive.
What does a narrow confidence interval mean?
It indicates a more precise estimate under the analysis assumptions. Precision does not by itself make the measured effect clinically important or the study unbiased.
Why are subgroup findings risky?
Many subgroup tests create opportunities for chance findings, especially when the analysis was not prespecified, the groups are small, or the overall primary result was negative.
Can a statistically significant trial support a compounded version?
Not automatically. The studied and compounded products may differ in formulation, manufacturing, route, concentration, and evidence, and compounded drugs are not FDA-approved.
Primary sources
- ASA Statement on Statistical Significance and P-ValuesAmerican Statistical Association · checked August 24, 2026
- Hypothesis Testing, P Values, Confidence Intervals, and SignificanceNational Library of Medicine, NCBI Bookshelf · checked August 24, 2026
- The clinician's guide to p values, confidence intervals, and magnitude of effectsEye, via PubMed Central · checked August 24, 2026
- A practical guide for understanding confidence intervals and P valuesOtolaryngology–Head and Neck Surgery, via PubMed · checked August 24, 2026
Continue researching
- Placebo-adjusted weight loss: how to read GLP-1 trial percentages →
- ClinicalTrials.gov protocols and statistical analysis plans: a peptide research guide →
- Randomized, blinded, and placebo-controlled peptide trials: what the terms mean →
- Peptide research: preprint vs. peer-reviewed study vs. press release →
Continue into provider research
Apply this guide’s verification questions to source-backed directory profiles and state coverage pages.
