To judge peptide research, ask three questions: what kind of study is it (cell, animal or human), how well was it designed, and has the finding been confirmed by others? A result in cells or mice can be a genuine scientific clue, but it is a long way from showing that something works, or is safe, in people. This guide explains the levels of evidence, why animal results so often fail to translate, and a practical checklist for reading any study.
Key Takeaways
- Not all evidence is equal. Systematic reviews and well-run randomised controlled trials in humans generally sit at the top of evidence hierarchies [1][2].
- Only about a third of highly cited animal research has been found to translate at the level of human randomised trials [3].
- Small studies, small effects, flexible analyses and financial interests all make a finding less likely to be true [4].
- In lab studies, impurities such as endotoxin or trifluoroacetate (TFA) can create misleading results [5][6].
- Most compounds that enter human trials do not end up approved [7][8].
The Types of Study, From Bench to Bedside
Peptide research usually moves through stages, and each stage answers a different kind of question.
| Study type | What it is | What it can show | Main limitations |
|---|---|---|---|
| In vitro (cells or tissues in a dish) | Experiments outside a living organism | Possible mechanisms; whether a peptide can affect cells at all | No whole-body context; concentrations may not be realistic; contaminants can confound results |
| Animal (in vivo) | Studies in mice, rats or other species | Effects in a living system; early safety signals | Species differences; often small and unblinded; publication bias |
| Observational human studies | Researchers observe people without assigning treatment | Associations in real populations | Cannot easily separate cause from coincidence |
| Randomised controlled trials (RCTs) | People randomly assigned to a treatment or a control | Whether a treatment causes an effect in humans | Costly; may be short or narrow in population |
| Systematic reviews and meta-analyses | Structured summaries of all relevant studies | The overall balance of evidence | Only as good as the studies included |
Evidence Hierarchies: A Useful Starting Point
The Oxford Centre for Evidence-Based Medicine (OCEBM) Levels of Evidence rank study types by how reliably they answer particular clinical questions. For questions about treatment benefits, the highest level is systematic reviews of randomised trials (or n-of-1 trials) [1]. The OCEBM itself warns that hierarchies have sometimes been used too rigidly and should be read together with its accompanying explanatory documents [1].
The GRADE system, widely adopted by guideline organisations, takes a more flexible approach [2]:
- Evidence from randomised trials starts as high quality, but can be downgraded for study limitations, inconsistent results, indirectness (the evidence doesn't quite match the question), imprecision, or reporting bias.
- Observational studies start as low quality, but can be upgraded, for example if the effect is very large or shows a dose-response relationship.
GRADE's authors give a sobering example of why this matters: hormone replacement therapy was widely recommended to reduce cardiovascular risk based on observational data, but randomised trials later showed it did not reduce that risk and may even have increased it [2].
Animal vs. Human Studies: Why Results Often Don't Translate
Animal research is essential and has contributed to many medical advances. But results in animals are an uncertain guide to results in people.
- Translation rates are modest. A widely cited analysis found that only about a third of highly cited animal research translated at the level of human randomised trials [3].
- Big failures happen. In stroke research, about 500 "neuroprotective" treatments were reported to improve outcomes in animal models, but only aspirin and very early intravenous thrombolysis proved effective in patients [3].
- Biology can differ. A large study comparing genomic responses to acute inflammatory stresses in people with those in the corresponding mouse models found that the mouse responses correlated poorly with the human conditions, with matches close to random for many genes [9].
- Publication bias inflates results. In systematic reviews of animal stroke studies, publication bias may account for a third or more of the reported efficacy [3].
This does not mean animal studies are worthless. It means a positive animal result is best treated as a reason to investigate further, not as proof of benefit in humans. To improve quality, the ARRIVE 2.0 guidelines set out what animal studies should report so others can judge their rigour and reproduce them [10]. Asking whether a paper follows ARRIVE is a useful quality check.
Why Cell and Lab Studies Need Extra Caution With Peptides
Peptide experiments have some specific pitfalls:
- Endotoxin contamination. Endotoxins are common contaminants in biologically active preparations and can complicate the study of the main ingredient's effects [5]. Cells may respond to the contaminant rather than the peptide. See endotoxin and sterility testing.
- Counter-ions such as TFA. Peptides purified by HPLC are often TFA salts. One study found that TFA itself reduced cell growth at very low concentrations, which could cause researchers to miss an effect or wrongly attribute one to the peptide [6].
- Identity and purity. If the material wasn't properly characterised, the result may not reflect the intended peptide. Look for reported analytical testing; our guide on how to read a peptide lab report explains what to look for.
Red Flags That Make Findings Less Reliable
In a famous essay, John Ioannidis argued that a research finding is less likely to be true when [4]:
- studies are small,
- effect sizes are small,
- many relationships are tested with little preselection,
- there is flexibility in designs, definitions, outcomes and analyses,
- there are financial or other interests and prejudice, and
- many teams are chasing statistical significance in a "hot" field.
These are not reasons to dismiss research, but they are reasons to look for replication and larger, well-designed studies.
A Practical Checklist for Judging Peptide Research
- Who or what was studied? Cells, animals or humans? How many?
- Was there a control group? Was allocation randomised and were researchers blinded?
- What was measured? A lab marker, or an outcome that matters to people (like symptoms or survival)?
- How big was the effect, and how precise? Look at confidence intervals, not just "significant" p-values.
- Was the material characterised? Identity, purity and contaminants reported?
- Was it registered in advance? Human trials should be registered before they start; registries such as those searchable via the WHO ICTRP portal help reveal unpublished or changed studies [11].
- Who funded it, and are there conflicts of interest?
- Has it been replicated? By independent groups, ideally summarised in a systematic review.
- Where was it published? A peer-reviewed journal is not a guarantee, but press releases, product websites and social media are not evidence.
From Promising to Proven: The Long Road
Even when early evidence is strong, most compounds do not reach approval. The FDA estimates that only about 25–30% of drugs in Phase 3 move on to the next stage [7]. A large MIT analysis of more than 21,000 compounds found success rates varied widely by disease area, estimating 3.4% for oncology drugs in its sample [8]. Read more in our guide to investigational compounds and clinical trial phases.
Approval decisions are made by regulators in each jurisdiction, such as the FDA (United States), EMA and European Commission (European Union), MHRA (UK), TGA (Australia) and Health Canada. See what "approved" actually means. Regulatory status varies by jurisdiction and may change over time. Consult the relevant regulatory authority for current information.
Frequently Asked Questions
Why don't animal studies always apply to humans?
Species differ in biology, metabolism and disease processes, and animal studies are often small and affected by publication bias. Only about a third of highly cited animal findings have translated in human randomised trials [3][9].
What is the strongest type of evidence?
For questions about whether a treatment works, systematic reviews of well-conducted randomised controlled trials generally rank highest, though quality within each type still matters [1][2].
Are cell studies useless?
No. They help explain mechanisms and generate hypotheses. But they cannot show that a peptide is safe or effective in people, and contaminants can affect results [5][6].
What does "peer-reviewed" mean?
It means other experts reviewed a paper before publication. It improves quality but does not guarantee a finding is correct or will be replicated [4].
How can I tell if a human trial is legitimate?
Check whether it is registered in a public trial registry, had ethics approval, and reports its methods and results transparently [11].
References
- OCEBM Levels of Evidence Working Group. The Oxford Levels of Evidence 2. Oxford Centre for Evidence-Based Medicine. https://www.cebm.ox.ac.uk/resources/levels-of-evidence/ocebm-levels-of-evidence ↗
- Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924-926. https://doi.org/10.1136/bmj.39489.470347.AD ↗
- van der Worp HB, Howells DW, Sena ES, et al. Can animal models of disease reliably inform human studies? PLoS Med. 2010;7(3):e1000245. https://doi.org/10.1371/journal.pmed.1000245 ↗
- Ioannidis JPA. Why most published research findings are false. PLoS Med. 2005;2(8):e124. https://doi.org/10.1371/journal.pmed.0020124 ↗
- Magalhães PO, Lopes AM, Mazzola PG, et al. Methods of endotoxin removal from biological preparations: a review. J Pharm Pharm Sci. 2007;10(3):388-404. https://pubmed.ncbi.nlm.nih.gov/17727802/ ↗
- Cornish J, Callon KE, Lin CQ, et al. Trifluoroacetate, a contaminant in purified proteins, inhibits proliferation of osteoblasts and chondrocytes. Am J Physiol. 1999;277(5):E779-E783. https://doi.org/10.1152/ajpendo.1999.277.5.E779 ↗
- U.S. Food and Drug Administration. Step 3: Clinical Research. https://www.fda.gov/patients/drug-development-process/step-3-clinical-research ↗
- Wong CH, Siah KW, Lo AW. Estimation of clinical trial success rates and related parameters. Biostatistics. 2019;20(2):273-286. https://doi.org/10.1093/biostatistics/kxx069 ↗
- Seok J, Warren HS, Cuenca AG, et al. Genomic responses in mouse models poorly mimic human inflammatory diseases. Proc Natl Acad Sci U S A. 2013;110(9):3507-3512. https://doi.org/10.1073/pnas.1222878110 ↗
- Percie du Sert N, Hurst V, Ahluwalia A, et al. The ARRIVE guidelines 2.0: updated guidelines for reporting animal research. BMJ Open Sci. 2020;4(1):e100115. https://doi.org/10.1136/bmjos-2020-100115 ↗
- World Health Organization. International Clinical Trials Registry Platform (ICTRP). https://www.who.int/tools/clinical-trials-registry-platform ↗
This article is for educational purposes only and is not medical advice. For health decisions, consult a qualified healthcare professional, and for regulatory questions, consult the medicines regulator in your jurisdiction.
