← The Signal Report Work with me

Report 157 · Lab Science

An R² of 0.999 is not a linearity test

Most quantitative results in a paper, a validation report or a certificate of analysis rest on a calibration curve, and the curve is usually defended with one number: "linear, R² = 0.999." That number measures how tightly the points hug a line. It does not tell you whether a line is the right shape, and a curve that is visibly bent can still score above 0.99.

A calibration works like this. You prepare standards at known concentrations, measure the instrument's response to each, and fit a relationship between response and concentration. Then you measure an unknown sample and read its concentration back off that relationship. If the relationship you fitted is the wrong shape, every result read off it inherits the error, and nothing about the unknown sample will warn you.

The usual shape is a straight line, fitted by least squares, and the usual evidence that the line is appropriate is the correlation coefficient r, or its square, the coefficient of determination R². Numbers like 0.999 look like certainty. They are a measure of something narrower.

What r measures

The correlation coefficient measures the strength of linear association: how much of the scatter in the response lines up with the concentration. A set of points that rises steadily will produce an r close to 1 whether it rises along a line or along a gentle curve, because in both cases almost all of the variation goes the same direction.

That is the point a 2002 paper in Accreditation and Quality Assurance set out to demonstrate with real instrument data. Joris Van Loco and colleagues took a series of cadmium calibration curves from atomic absorption spectroscopy, a workhorse technique in trace-metal testing. From the publisher's abstract:

All the investigated calibration curves were characterized by a high correlation coefficient (r >0.997) and low quality coefficient (QC <5%), but the straight-line model was systematically rejected at the 95% confidence level on the basis of the Lack-of-fit and Mandel’s fitting test.

Every curve passed the number most labs report. Every curve failed the tests designed to detect curvature. And it mattered for results: the abstract reports that a straight-line model systematically biased predictions for mid-scale standards, while results from a quadratic model did not differ significantly from the expected values. Their conclusion is that a straight-line model with a high correlation coefficient but a lack of fit gives significantly less accurate results than its curved alternative.

A worked example

To see how this happens without any noise at all, here is a small calculation of my own. It is not from the paper; it is an illustration of the mechanism.

Take six standards at 0, 2, 4, 6, 8 and 10 units, and suppose the instrument responds with a slight droop at the top: response = concentration − 0.02 × concentration². At the highest standard the signal is 8.0 instead of 10, which is the kind of flattening you see as a detector approaches saturation. Fit a straight line through those six points and you get:

  • r = 0.9973, R² = 0.9947. Comfortably above the 0.99 that many labs treat as a pass.
  • Residuals of −0.27, +0.05, +0.21, +0.21, +0.05, −0.27. Negative at both ends, positive in the middle: an arch.
  • Read the standards back off the line and the 4-unit standard comes back as 4.27, a 6.7 percent overestimate. The blank comes back as −0.33, a negative concentration.
  • A sample that truly contains 1 unit reads as about 0.89, an 11 percent underestimate, and the relative error keeps growing as you approach the low end of the range.

Nothing here is noise. The data are perfect; the model is wrong. The r value cannot see it, because it rewards any steady rise. The residuals see it immediately, because they have a pattern, and a pattern in the residuals is exactly what a wrong shape leaves behind.

The low end deserves special attention. Near the blank, the fitted intercept dominates the answer, so a curve forced into a line can produce its largest relative errors at the concentrations closest to the detection and quantitation limits. That is often where the results matter most: contaminants, impurities, trace residues. (This is my reasoning from the example, not a finding of the sources, but it follows directly from how the line is fitted.)

What the guidelines actually ask for

The standard references for method validation do not treat r as sufficient, although they are sometimes quoted as if they did.

ICH Q2(R2), the international guideline on validating analytical procedures for pharmaceuticals, revised and adopted in November 2023, does ask for the number. It also asks for more, in the same paragraph:

A plot of the data, the correlation coefficient or coefficient of determination, y-intercept and slope of the regression line should be provided. An analysis of the deviation of the actual data points from the regression line is helpful for evaluating linearity (e.g., for a linear response, the impact of any non-random pattern in the residuals plot from the regression analysis should be assessed).

The same section recommends a minimum of five concentrations. Five points and an R² is what often gets reported. The plot and the residual assessment are the parts that tend to go missing.

The Eurachem guide to method validation, the standard laboratory reference in Europe, is more direct. Its quick-reference procedure for checking the working range says to calculate and plot residuals, and then:

Random distribution of residuals about zero confirms linearity. Systematic trends indicate non-linearity or a change in variance with level.

Note what counts as confirmation in that sentence. It is the residuals, not the correlation coefficient. The same guide recommends 6 to 10 concentrations, measured 2 to 3 times each, and the replicates are not a formality: a formal lack-of-fit test works by comparing how far the points sit from the line against how far replicates at the same concentration sit from each other. Without replicates, you cannot separate a bent model from ordinary scatter.

Why r keeps getting reported anyway

It is one number, every spreadsheet computes it, and it is almost always high. That last property is the problem. Calibration standards are chosen to span a range, and any steady rise across a wide range produces a large r. A number that is nearly always close to 1 cannot do much work separating a good calibration from a bad one.

It also answers a question nobody is really asking. Nobody doubts that the instrument responds more to more analyte. The question is whether the response follows the specific shape you are about to use to convert signals into concentrations. That is a question about the pattern of the misfit, and only the residuals, a lack-of-fit test, or a comparison with a curved model can answer it.

Why this one is personal

My own published research is in microwave spectroscopy. Nothing in this report comes from that work, and no measurement was performed for it; the example above is arithmetic and the evidence is from the sources below. I care about this one because it is the same pattern this beat keeps returning to: a single reassuring number standing in for the thing it was supposed to check. I made a similar argument about what "not detected" means in Report 062, about error bars in Report 056, and about the 260/280 purity ratio in Report 128.

What I could not confirm

I read the Van Loco paper's abstract, not the full text. The article is behind a paywall. Everything I attribute to it (the cadmium AAS curves, r above 0.997, rejection by lack-of-fit and Mandel's test at 95 percent confidence, and the mid-scale bias of the linear model) comes from the abstract published on the journal's page. I have not seen their data, their concentration ranges, or the size of the bias they found.

Two classic references I could not open. The Royal Society of Chemistry's Analytical Methods Committee published "Uses (proper and improper) of correlation coefficients" (Analyst, 1988) and "Is my calibration linear?" (Analyst, 1994), both widely cited on exactly this point. The free technical-brief copy is no longer at its old address and the journal versions are paywalled, so I have not quoted them. The IUPAC harmonized guidelines on single-laboratory validation (2002) were also cited in this context but sat behind a bot check I did not attempt to bypass.

Guidelines are not the whole practice. ICH Q2(R2) is written for pharmaceutical analysis and Eurachem is a general laboratory guide. Individual regulators, accreditation bodies and standard methods can set their own acceptance criteria, and some do specify minimum r values alongside other checks. This report is about what r can and cannot show, not about any specific rulebook.

The signal

When a methods section, validation report or certificate says "linear, R² = 0.999," read it as "the response rises steadily with concentration." That is worth knowing and it is not the same claim as "a straight line is the right model." For that, look for the residual plot, a lack-of-fit or curvature test, how many concentrations were run and whether they were replicated, and whether the unknowns sit near the ends of the range, where a wrong shape does the most damage.

If you are producing the calibration, report the residuals. They take one extra plot. They are the only thing in the usual output that can actually tell you the line is wrong.

Sources

  1. Joris Van Loco, Marc Elskens, Christophe Croux and Hedwig Beernaert, "Linearity of calibration curves: use and misuse of the correlation coefficient," Accreditation and Quality Assurance 7(7):281–285, July 2002, DOI 10.1007/s00769-002-0487-6. (PRIMARY. Paywalled: only the abstract on the publisher's article page was opened and read, and nothing beyond the abstract is attributed to the paper. Source for: the statement that r very close to one might be obtained for a clear curved relationship; lack-of-fit and Mandel's fitting test as more suitable checks; the cadmium AAS calibration curves with r above 0.997 and QC below 5% that were systematically rejected at 95% confidence, quoted verbatim; the systematic bias of the linear model for mid-scale standards versus the quadratic model; and the conclusion that a high-r straight line with lack of fit yields significantly less accurate results.)
  2. International Council for Harmonisation, "Validation of Analytical Procedures Q2(R2)," ICH Harmonised Guideline, final version adopted 1 November 2023 (with error corrections dated 30 November 2023). (PRIMARY, full guideline downloaded and read. Source for: section 3.2.2.1 on linear response, including the requirement to provide a plot, the correlation coefficient or coefficient of determination, the y-intercept and the slope, and the statement on residual analysis and non-random residual patterns, quoted verbatim; and the recommended minimum of five concentrations.)
  3. B. Magnusson and U. Örnemark (eds.), "Eurachem Guide: The Fitness for Purpose of Analytical Methods – A Laboratory Guide to Method Validation and Related Topics," 2nd edition, 2014, ISBN 978-91-87461-59-0. (PRIMARY, full 70-page guide downloaded and read. Source for: the working-range assessment by visual inspection, regression statistics and a residual plot; the Quick Reference procedure of 6 to 10 concentrations measured 2 to 3 times; and the statement, quoted verbatim, that a random distribution of residuals about zero confirms linearity while systematic trends indicate non-linearity or a change in variance with level. The guide's reference list also cites the 1988 Analytical Methods Committee paper on correlation coefficients, which was not opened.)
  4. Onur Oncer, "What 'not detected' actually means," The Signal Report 062; "What those error bars actually mean," The Signal Report 056; and the 260/280 ratio report, The Signal Report 128. (Earlier reports in this beat on single summary numbers standing in for the check they were meant to perform.)

Scope note: this report explains what a correlation coefficient can and cannot show about a calibration curve. The worked example (response = concentration − 0.02 × concentration², six standards from 0 to 10) and every number derived from it are the author's own arithmetic, not data from any source or instrument. No measurement was performed, and no method, laboratory or product is assessed or recommended.

Onur Oncer
Onur Oncer

U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.

← All reports