← The Signal Report Work with me

Report 056 · Lab Science

What those error bars actually mean

A little vertical line on top of a bar. It might be a standard deviation, a standard error, or a 95% confidence interval, and those three things answer completely different questions. Worse, one of them shrinks automatically as you collect more data, which means a figure can be made to look more convincing without the result becoming any more true.

Spectroscopy taught me to distrust a clean number. You take a measurement, you take it again, and the second one is different. Not because anything changed, but because measurement has a spread, and the spread is data. It is arguably the more informative half of the data. A result reported without it is not a result, it is an anecdote with a decimal point.

Which is why the error bar is such a strange object. It is the part of the figure that carries the honesty, and it is also the part almost nobody defines, and it can mean at least three different things that a reader cannot tell apart by looking.

This one is not a new scandal. It is a piece of craft knowledge that has been written down clearly for nearly twenty years and is still routinely mangled, which is exactly the kind of thing worth explaining rather than reporting.

Three quantities, one picture

The canonical explainer is Geoff Cumming, Fiona Fidler and David Vaux, "Error bars in experimental biology," published in the Journal of Cell Biology in 2007. Their opening problem statement is the entire issue in one sentence: error bars "may show confidence intervals, standard errors, standard deviations, or other quantities."

Take those one at a time, because the distinction is not pedantry. They are answers to different questions.

Standard deviation (SD) describes your data. It says how spread out the individual measurements were. If you measure ten samples and they scatter widely, the SD is large, and that is a fact about the world you are studying. Collect more data and the SD does not systematically shrink. It converges on the true spread, because the spread is real.

Standard error of the mean (SE or SEM) describes your estimate, not your data. It says how precisely you have pinned down the average. The relationship is simple and it is the crux of everything below: SE = SD/√n. Because n is in the denominator, the standard error falls as you add measurements, regardless of whether the underlying variation is large or small. Quadruple your sample size and the SE halves, while the SD sits where it always was.

The 95% confidence interval (CI) is a range constructed so that, over many repetitions of the experiment, intervals built this way would contain the true value 95% of the time. It is roughly two standard errors wide for reasonable sample sizes, and it is the one that most directly answers the question a reader actually has, which is how much the reported number could move if someone did this again.

Now the uncomfortable consequence. SD and SE bars on the same data look identical in kind and differ only in length, and the SE bars are always shorter. Always, for any n greater than one. So a figure drawn with SE bars looks tighter, cleaner and more decisive than the same figure drawn with SD bars, and no rule anywhere forces an author to tell you which one they chose. Cumming and colleagues' first rule of thumb is exactly that mundane: describe in the figure legend what the error bars are.

I am not alleging that authors do this to deceive. Mostly it is convention inherited from a lab's previous papers, which is how methodological habits actually propagate. But the reader-facing effect is the same either way. If you do not know which quantity is drawn, you cannot read the figure, and a small bar has told you nothing.

The failure that matters more

There is a deeper error underneath the labeling problem, and it is the one that turns a figure from unclear into wrong.

Every one of these quantities depends on n. So the question that governs the whole picture is: n of what?

Cumming, Fidler and Vaux are explicit that this is the load-bearing distinction. It is essential, they write, that n, the number of independent results, is carefully distinguished from the number of replicates, which they define as repetition of measurement on one individual in a single condition. And they draw the hard conclusion: for replicates, n = 1, and it is therefore inappropriate to show error bars or statistics at all.

That sounds abstract until you put it in a lab. Suppose you run one culture dish, image it, and measure 300 cells. What is n? The instinct is 300, and 300 is a wonderfully large number that makes error bars small and p-values tiny. But those 300 cells shared one dish, one passage, one batch of medium, one afternoon. They are 300 measurements of a single experiment, not 300 experiments. If something was off about that dish, all 300 are off together, and no amount of counting more cells will reveal it.

Run the experiment three times on three separate days, and n = 3. Small, honest, and it is the number that tells a reader whether the result would recur.

This is pseudoreplication, and in cell biology it is common enough that a 2020 paper in the same journal was written specifically to fix it. Samuel Lord, Katrina Velle, R. Dyche Mullins and Lillian Fritz-Laylin titled the preprint version bluntly: "If your P value looks too good to be true, it probably is." Their diagnosis, in the abstract of the published version, is that the sample size n used for statistical tests should represent biological replicates, meaning independent measurements of the population from separate experiments. Their preprint states the problem even more directly, that the literature is littered with erroneously tiny P values, often the result of evaluating individual cells as independent samples.

Note the direction of the error. Pseudoreplication does not add noise. It manufactures confidence. Counting 300 cells instead of 3 experiments divides your standard error by roughly ten, shrinks the bars to almost nothing, and drives the p-value through the floor. The figure becomes more persuasive precisely as it becomes less trustworthy, and nothing about the picture warns you.

Reading the overlap

The other thing people do with error bars is eyeball two of them and decide whether the difference is real. There are rough rules for this, and they are more constrained than the folk version.

The folk version says that if the bars overlap, there is no significant difference. That is not right, and how wrong it is depends entirely on which quantity is drawn. Cumming and colleagues give rules of eye for reasonably large samples: with n of at least 10, if 95% confidence interval bars overlap by about half the average arm length, the p-value is in the neighborhood of 0.05. For standard error bars the geometry is different, and a visible gap of about one SE corresponds to roughly p = 0.05, with a gap of about two SE corresponding to roughly p = 0.01.

Read that carefully, because it inverts the intuition. With SE bars you can have a clear gap between the bars and still be sitting right at the conventional significance threshold, not comfortably past it. And with CI bars, overlapping bars can still be a significant difference. These are approximations for independent groups at decent sample size, not a substitute for the actual test, and the authors present them as rules of thumb rather than law. The useful takeaway is not the specific geometry. It is that no visual overlap rule survives without knowing which bar you are looking at and how big n is.

What I would actually do with this

Three questions, in order, and they take about ten seconds on any figure.

First: does the legend say what the bars are? If it does not, the figure is unreadable and you should treat the size of the bars as carrying no information. Not "probably fine." No information.

Second: what is n, and n of what? Look for whether the number counts independent experiments or observations pooled from within one. If a paper reports n in the hundreds for something that is obviously run in batches, cultures, dishes, litters, or sessions, that is the flag. This is the same move I have described for a reagent nobody validated and for a cell line nobody authenticated: separate what was actually independent from what was assumed to be.

Third: if the bars are SE, mentally lengthen them. Multiply by roughly √n to recover a sense of the actual spread in the data, then ask whether the effect still looks like something you would bet on. Often it does. Sometimes the sight of the real scatter changes the story entirely, and that scatter was in the experiment the whole time. It was simply not the thing that got drawn.

Honest limits

None of this means SE bars are illegitimate. They answer a real question, precision of the estimate, and for that question they are the right tool. The problem is exclusively one of unlabeled ambiguity and of n being defined at the wrong level.

The overlap rules above come with conditions I have stated but want to underline: they assume independent groups, they assume n of at least about 10, and they are eye rules rather than tests. Small samples behave differently. And I have not surveyed how often published figures get this wrong, because I have not verified a current prevalence figure I would stand behind, so I am describing a well-documented failure mode rather than claiming a rate.

The signal

An error bar is a claim about how much you should expect the number to move. Almost every other element of a figure is designed to convince you. The error bar is the one element whose job is to tell you how much not to be convinced, which makes it the most valuable mark on the chart and the easiest one to quietly shrink.

It shrinks in two ways. Honestly, by doing the experiment more times. And artificially, by redefining what counts as a time. The picture looks the same either way. The only defense is to ask, every time, what the bar is and what n counts, and to notice that a paper which does not tell you has not made a weak claim. It has made an unreadable one.

Sources

  1. Geoff Cumming, Fiona Fidler and David L. Vaux, "Error bars in experimental biology," Journal of Cell Biology, 2007 Apr 9;177(1):7–11, DOI 10.1083/jcb.200611141, PMID 17420288. (Primary, peer reviewed, open access via PubMed Central. Source of the statement that error bars "may show confidence intervals, standard errors, standard deviations, or other quantities"; the relationship SE = SD/√n; the distinction between n as the number of independent results and replicates as repetition of measurement on one individual in a single condition, with the conclusion that for replicates n = 1 and error bars are inappropriate; the eight rules of thumb, including the rule that figure legends must describe what the error bars are; and the rules of eye for n ≥ 10 relating 95% CI overlap of about half an average arm length to P ≈ 0.05, and gaps of one and two SE to P ≈ 0.05 and P ≈ 0.01 respectively.)
  2. Samuel J. Lord, Katrina B. Velle, R. Dyche Mullins and Lillian K. Fritz-Laylin, "SuperPlots: Communicating reproducibility and variability in cell biology," Journal of Cell Biology, 2020 Jun 1;219(6):e202001064, DOI 10.1083/jcb.202001064, PMID 32346721. (Primary, peer reviewed. Source of the statement that the sample size n used for statistical tests represents biological replicates, independent measurements of the population from separate experiments.)
  3. Samuel J. Lord, Katrina B. Velle, R. Dyche Mullins and Lillian K. Fritz-Laylin, "If your P value looks too good to be true, it probably is: Communicating reproducibility and variability in cell biology," arXiv:1911.03509. (The open preprint of the paper above, cited separately because the sentence about the literature being littered with erroneously tiny P values from evaluating individual cells as independent samples appears in this version's abstract. Preprint, and labeled as such in the text.)
Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports