How we grade evidence
The five tiers, what each one can and cannot tell you, and why evidence tier, regulatory status and claim verdicts are three separate things.
Why a grade at all
Every report in this atlas carries a tier from L1 to L5. The tier answers one question and one question only: what kind of study is the best evidence that exists?
It is not a score for how well something works. It is not a safety rating. It is not a recommendation. A compound can sit at L5 because someone ran a 1,100-patient randomised trial — and that trial can have shown the compound does nothing. Thymosin α1 is exactly that case, and it is graded L5.
The reason to grade study design rather than outcome is that design is objective and checkable. Whether a trial was randomised is a fact. Whether a compound “works” is a conclusion that depends on the claim, the population and the endpoint.
The five tiers
L1 — Cells and test tubes
The compound did something measurable in cultured cells, isolated tissue or a biochemical assay.
What it tells you: a mechanism is chemically possible. These experiments are how real pharmacology begins, and a good one can be quite precise about a molecular pathway.
What it cannot tell you: anything about a living body. Cells in a dish have no liver to metabolise the compound, no kidneys to excrete it, no immune system, no barriers to cross and no competing physiology. Concentrations used in culture are frequently unachievable in a person.
L2 — Animal studies
Usually rats or mice, sometimes larger species.
What it tells you: the compound does something in an intact organism with circulation, metabolism and immunity. This is a genuine and necessary step.
What it cannot tell you: whether the same happens in humans. Species differences are not a technicality — carnosine is the sharp example in this atlas, where the enzyme that destroys it is confined to the kidney in rodents but active in human blood, so rodent results systematically overstate what oral carnosine can do in people. Doses in animal work are also often far above anything used in humans. One clinician reviewing this field cited the figure that only around 5% of interventions tested in animals reach regulatory approval.
L3 — Early human data
Case reports, open-label pilots, retrospective chart reviews, small uncontrolled series.
What it tells you: the compound has been given to people, and gross acute harm was not observed in that small group. Occasionally a dramatic effect in a rare disease is informative even without a control.
What it cannot tell you: whether the compound caused the improvement. There is no comparison group, so natural history, expectation, regression to the mean and the attention of the clinic cannot be separated out. Most of the human peptide literature lives here, and it is routinely quoted as though it lived at L4.
L4 — Controlled human trials
Randomised, placebo- or comparator-controlled, ideally blinded.
What it tells you: a real estimate of effect, because the control group experienced the same passage of time and the same expectations. This is the first tier where “it works” becomes a defensible statement.
What it cannot tell you: that the result will hold. Single trials get overturned, small trials produce exaggerated effects, and a trial can be randomised and still be too small to mean much — the GHK-Cu trial in this atlas had thirteen completers.
L5 — Replicated or large trials
Multiple independent controlled trials, or trials large enough to settle a question, ideally with a systematic review pooling them.
What it tells you: the most reliable answer medicine produces. Note again: a reliable answer can be negative.
What it cannot tell you: that pooling was done well. Meta-analyses inherit the flaws of what they include, and the collagen literature in this atlas shows how differently the same trials read depending on whether you split them by funding and quality.
Three things that are not the same
Most confusion about peptides comes from collapsing three independent facts into one impression. We report them separately on every page.
A worked example. GHK-Cu is legal to sell and has been used for decades (regulatory status: fine), sits at L3 (evidence tier: thin), and in the one randomised trial with blinded objective assessment it improved nothing the evaluators could measure, while patients did report liking their skin more (claim verdicts: mostly unproven, one supported). All three statements are true at once, and no single label could carry them.
The claim verdicts
In each report’s claims table, every verdict refers to that specific claim, not the compound as a whole.
| Claim as usually stated | Verdict | What the published evidence actually shows |
|---|---|---|
| Supported | Supported | Controlled human evidence, or an uncontroversial mechanistic fact, backs this specific claim. Not “proven beyond doubt” — supported. |
| Partly supported | Partly supported | True under conditions narrower than the claim implies. Oral glutathione raising glutathione levels is partly supported: yes at six months, no at four weeks, no from a single dose. |
| Mixed | Mixed | Real evidence points both ways, or the answer changes with how you analyse it. The collagen skin literature is the archetype. |
| Unproven | Unproven | The claim may well be true. Nobody has tested it adequately. This is the most common verdict in this atlas, and it is not the same as “false”. |
| Not supported | Not supported | Tested and failed, or contradicted by the evidence. Thymosin α1 for sepsis mortality is not supported — a 1,106-patient trial found 23.4% versus 24.1%. |
The distinction between Unproven and Not supported is the one we are most careful about. “Nobody has looked” and “somebody looked and it did not work” are very different states of knowledge, and marketing tends to describe both as “promising”.
What would change a grade
These grades are not permanent. A tier rises when a better study is published — a first randomised trial, an independent replication, a properly powered confirmation. A tier can effectively fall in usefulness when a large trial contradicts a smaller positive one, which is what happened to thymosin α1 in sepsis: the reviews before 2025 described it as a promising adjuvant therapy, and a single well-run trial retired that description.
Each report carries the date it was last reviewed. If regulatory facts here are more than a few months old, check the primary source before relying on them — this field is moving quickly, and the compounding rules in particular were in active flux throughout 2026.