Not all scientific studies on food have the same evidentiary value. Researchers often describe a hierarchy of evidence: systematic reviews of randomized trials sit near the top, while cell and animal studies sit much lower when the question is what humans should eat. People who cite “a study” to support a dramatic nutrition claim are often relying on evidence from the lower half. Seven basic questions are enough to dismantle a large share of the nutrition claims that circulate online.

It happens at the dinner table, in comments under a reel, in WhatsApp groups where somebody has shared the latest food scare. At some point, when the arguments run out, the sentence arrives: “Read the studies.”

It is a powerful phrase. It shuts down the discussion, puts the other person on the defensive and creates a distinction between the person who knows and the person who does not. It has one problem: in many cases, the person saying it has not read a study at all. They read a headline. Watched a reel. Heard something from someone who heard it from someone else. “Read the studies” is not being used as an invitation to learn. It is a rhetorical weapon, useful for winning an argument rather than finding out what is true.

This article is for the next time you hear that sentence and want to answer with a simple question: “Which study? And where does it sit in the evidence hierarchy?”

Because a hierarchy exists. It is one of the ways the scientific community distinguishes research designs by the kinds of questions they can answer and by how vulnerable they are to bias. Not all studies are equal. Understanding the difference is one of the most useful skills you can develop if you want to avoid being manipulated by people who talk about food as though it were a religion.

The source article describes ten broad levels. Near the top are research designs capable of carrying more causal weight. Near the bottom are studies that may be essential for generating hypotheses or understanding mechanisms but cannot, by themselves, settle what happens when humans eat ordinary amounts of a food over years. The higher levels are rarer because good trials are expensive, slow and difficult. Lower-level studies are easier and more numerous. That is part of the reason a guru can almost always find “a study” that appears to support a claim.

Let us start near the bottom and work upward.

At the ground floor are in-vitro studies and animal models. The first are performed on isolated cells or tissues in a laboratory; the second on mice, rats or other animals. They are fundamental for understanding biological mechanisms — how a substance interacts with a cell, what happens in a mouse liver when a particular molecule is administered. But they do not tell you directly what happens to a human being eating normal amounts of food. A mouse is not a person. A cell in a dish is not an organism. And doses used in mechanistic studies can be far higher than ordinary dietary exposure. If someone says “it causes cancer” on the basis of a high-dose mouse experiment, they are making an enormous logical leap — a little like claiming that because a car is destroyed at 300 kilometers per hour, driving it at 50 is inherently dangerous.

A step above are expert opinions, narrative reviews and consensus statements. These can be very useful because they summarize the perspective of people who know a field. But narrative selection is subjective by definition: authors decide which literature to emphasize and which to leave aside. Strong priors or conflicts of interest can shape the story. An expert view is not worthless; it is simply different from a systematic method designed to identify all relevant evidence.

Then come ecological studies and case reports. Ecological studies compare data across whole populations — for example, countries that consume more olive oil may have lower rates of heart disease. These studies can generate important hypotheses, but population-level comparisons hide countless variables: climate, income, healthcare, genetics, work patterns, smoking, social life. If Italy has lower rates of something than northern Europe, that does not prove olive oil caused the difference. Case reports, meanwhile, can alert medicine to something new or unusual, but one patient or a handful of patients cannot establish a general rule.

Cross-sectional studies take a snapshot of a population at one moment and measure exposure and outcome at the same time. They can show that two things coexist, but not which came first. If people who eat more vegetables are also leaner, you cannot automatically conclude that the vegetables caused the lower body weight. It might be partly the reverse, or both could be influenced by other behaviors. Cause and consequence are tangled.

Retrospective cohorts and case-control studies try to reconstruct the past. In one, researchers take a group and look backward at previous exposures; in the other, they begin with people who already have a disease and compare them with people who do not, searching for differences in past behavior. These designs can be more informative than a simple snapshot, but looking backward brings problems of recall. People struggle to remember exactly what they ate last week, let alone years ago. And people who are already ill may, unintentionally, remember or report past behaviors differently.

Prospective cohort studies are one of the stronger observational designs. Researchers enroll a large group of people who are initially free of the outcome of interest, record exposures such as diet and then follow them for years, sometimes decades. The time sequence is clear: exposure first, outcome later. But observation is still not control. People who eat more fish may also have higher income, more education, better healthcare access or a generally more health-conscious lifestyle. Statistical adjustment can reduce these differences, but never guarantees they have all disappeared.

Then there are systematic reviews and meta-analyses of observational studies. Instead of selecting one convenient paper, researchers search systematically for all studies answering a specific question, apply predefined inclusion criteria and combine the results when appropriate. This is far more informative than citing a single cohort. But the synthesis cannot fully escape the limits of the studies underneath it. A meta-analysis of observational studies is still built from observational evidence. Combining many imperfect studies improves precision; it does not magically turn association into randomized causation.

Above that sit non-randomized interventions and community or cluster interventions. Here researchers actively change something — a diet, a school menu, a food environment — and compare outcomes. That is a step closer to causal inference because an intervention actually occurs. But without individual random assignment, groups can still differ in important ways before the intervention begins.

Then we reach the design many people treat as the reference standard: the randomized controlled trial, or RCT. Participants are randomly assigned to different groups, one receiving the intervention and another a control condition, and outcomes are compared. Randomization matters because, when done properly and with enough participants, it tends to distribute both known and unknown characteristics across groups. Income, education, genetics and lifestyle are less likely to line up systematically with one treatment. A crossover trial goes further in a different way: the same people receive more than one condition at different times, so each participant partly acts as their own control. Nutrition creates practical limits even for RCTs. You cannot easily make thousands of people eat the same prescribed diet for twenty years to see who develops disease. Long trials are expensive. Adherence is difficult. Some interventions would be unethical. So nutrition science will always depend on a mixture of randomized evidence, observational cohorts, mechanistic work and clinical judgment.

Near the top are systematic reviews and meta-analyses of randomized controlled trials and, above them in some schematic hierarchies, umbrella reviews that synthesize multiple systematic reviews. These designs can offer the highest level of summary evidence available for a question when the underlying studies are good. They are not infallible — nothing is — but a high-quality systematic review of well-conducted randomized trials has passed through more layers of methodological scrutiny than a single animal experiment or isolated cohort.

That is the hierarchy. Now comes the part that matters in real life.

When someone says “read the studies” and you ask which study, three things can happen. First: they cannot answer because they did not read one. Second: they cite something from much lower in the hierarchy — a mouse experiment, a case report, an expert opinion — and present it as definitive proof. Third, and more rarely: they really do cite a strong study. In that case, good. But the next question still matters: does that result fit the wider body of evidence, or is it an outlier being used to erase everything else?

Food misinformation follows a remarkably repetitive pattern. Take one study — sometimes a real and perfectly legitimate one — and remove it from context. Ignore the rest. Confuse correlation with causation. Use laboratory doses to frighten people about normal dietary exposure. Quote relative numbers that sound enormous even when the absolute difference is tiny. A statistically detectable difference in a huge cohort can make a perfect headline while barely changing an individual's real-world risk.

If you want to protect yourself, seven questions are enough to begin with. They are the same kinds of questions researchers ask when reading a paper. You do not need a degree. You need the willingness not to stop at the headline.

What kind of study is it? How many people were involved? How long did it last? Was it conducted in humans or animals? Who funded it? Have independent studies found similar results? And is the size of the effect meaningful in real life, or merely statistically detectable?

Use those questions consistently and a large proportion of nutrition claims start to look very different. Not because every claim is false — some are true, many are partly true — but because the communication is often distorted. A small piece of truth is inflated until it looks like absolute certainty. Science does not usually work that way. Science accumulates: one study adds a piece, and confidence grows when many pieces point in the same direction.

That brings us back to everything discussed on this site about labels, apps and scoring systems. They are tools. Scientific research is a tool too. No single study settles every question, even when it sits high in a hierarchy. The purpose of research is to accumulate knowledge and periodically synthesize it. People who say “it is proven” are often compressing a complicated picture. People who say “read the studies” often have not. And people who truly understand research are frequently the ones who speak with more caution, because they know caution is not weakness. It is competence.

One more point is worth keeping. “High in the hierarchy” does not automatically mean true. A badly conducted randomized trial with too few participants, flawed methods or serious conflicts can be less informative than a rigorous observational study following tens of thousands of people for twenty years. The hierarchy describes the potential strength of a design, not the quality of every execution. In the same way, “low in the hierarchy” does not mean useless. Cell and animal studies are indispensable for understanding mechanisms and generating hypotheses. The problem begins when somebody uses them to tell you exactly what to put on your plate while skipping all the intermediate steps.

The next time you hear “read the studies,” do not lower your eyes. Raise the question.