The evidence ceiling

A result cannot be stronger than the design that produced it. A dramatic finding in mice is still a finding in mice; a guideline that summarises fifty trials is still not one of them. This page takes a study you have found and gives you the highest grade its design could ever support, and the reason.

It cannot see you. It does not know your diagnosis or your treatment, it is not told what the study was about, and there is no answer you can give it that will make it say anything about whether something would help you. Its only possible output is a ceiling.

Every answer this page can give

The tool above is a shortcut to a row of this table. Both are the same function, and it is the same one the register runs against itself, so a study you look up here and an entry you read on this site cannot disagree.

What it isMost it could supportWhy
A trial that randomly assigned people to something or notAIt randomised people and measured something that happened to them. Nothing caps this: it is the kind of evidence the top of the ladder is for.
A trial that randomly assigned people to something or not, measuring a marker rather than a personCIt randomised people, which is the strongest thing a study can do, but it measured a marker rather than a person. A marker that moves is not a person who is better, and many markers have moved in trials where nobody lived longer.
A review that pools the numbers from several studies, of trials that randomised peopleAIt pools randomised trials, so it carries what they carry. Pooling does not add strength, but it does not lose it either.
A review that pools the numbers from several studies, of studies that followed what people already didBIt pools observational studies. More of them is not the same as randomising anyone: whatever made people different in one study made them different in all of them.
A review that pools the numbers from several studies, of both kinds togetherBIt pools randomised and observational studies together. The observational part cannot be separated out by a reader, so the whole is read at the level of its weaker half.
A review that gathers studies without pooling their numbers, of trials that randomised peopleBIt gathers randomised trials without pooling their numbers, so it establishes that a literature exists rather than what that literature amounts to.
A review that gathers studies without pooling their numbers, of studies that followed what people already didCIt gathers studies that followed what people already did, and does not pool their numbers. Neither half of that settles anything: nobody was assigned, and no combined estimate is offered.
A guideline or consensus statement from a professional body, of trials that randomised peopleBIt is a guideline. The studies behind it may be randomised and may be excellent, and they are not the thing in front of you: you cannot read them, count them, or check what was left out. That remove is the limit worth knowing.
An evaluation of a literature by an agency, of studies that followed what people already didBIt is an agency reading a literature and stating a conclusion. That is a serious judgement made by people with access to more than you have, and it is still a judgement about other people’s studies rather than a study.
A study that followed people and recorded what they were already doingBIt followed people and recorded what they were already doing. Nobody was assigned anything, so whatever made them do it may be what made them differ. That is the limit no amount of careful adjustment removes.
A study that gave everyone the same thing, with no comparison groupCEveryone in it received the same thing, so there is nobody to compare them with. People who enrol in studies differ from people who do not, and improvement over time is what most illnesses do anyway.
A set of individual cases, written upDIt is a set of cases somebody chose to write up. The ones that did not go well are not in it, and there is no way from inside it to know how many those were.
An experiment in animalsDIt was not done in people. However strong the result, a finding in animals or in a dish cannot show that something helps a person, because the step from one to the other is exactly what has not been tested.
An experiment in cells, or in a dishDIt was not done in people. However strong the result, a finding in animals or in a dish cannot show that something helps a person, because the step from one to the other is exactly what has not been tested.
Anything registered that has not reported a resultDIt has not reported a result. A registered trial is a question somebody is asking, not an answer, and nothing that has not reported can lift anything.

What a ceiling is not

A ceiling is the most a study could show, never what it did show. A randomised trial with a clinical endpoint has a ceiling of A and may have found nothing at all, or found harm. Grades E and X on this site are harm findings and sit outside this table entirely: no design ceiling applies to them, because they are not claims about benefit.

The ladder itself is on the front page, and the grade table has every claim the book grades.