If you have ever sent RNA off for sequencing, the core facility probably sent back a report with a RIN on it. A 9 or a 10 feels like an A. A 6 starts a conversation. Some labs will not sequence below 7. The number has the feel of a physical measurement, like a concentration or a pH. It isn't one, and knowing what it actually is changes how much weight it can carry.
What it replaced
Before the RIN, RNA quality was judged by eye. You ran the sample on a gel and looked at the two big ribosomal RNA bands, 28S and 18S. A 28S:18S ratio of about 2.0 or higher meant good RNA. The paper that introduced the RIN, published in BMC Molecular Biology in 2006 by Andreas Schroeder and colleagues, lays out why that was a problem. Reading gels by eye "is subjective, hardly comparable from one lab to another," and the ribosomal ratio itself showed only weak correlation with RNA integrity in many cases.
The fix ran on a specific instrument: Agilent's 2100 Bioanalyzer, which separates RNA by size in a microfluidic chip and draws the result as a trace called an electropherogram. Several of the paper's authors were Agilent employees, and some of the training data came from Agilent. That is not a scandal; it is how instrument methods usually get built. But it is worth knowing that the field's standard integrity score was born as a feature of one vendor's software.
How the number was made
This is the part most people who use RIN have never read. The team collected 1,208 RNA traces. About 30% came from known tissues (liver, kidney, colon, spleen, brain, heart and placenta from humans, mice and rats). The origin of the rest "was not traceable," beyond being mammalian cells or cell lines. Then, in the paper's words, "Each of the 1208 samples was assigned manually to one of the categories by experienced expert users," on a scale from 1 (totally degraded) to 10 (fully intact).
Those human grades became the target. The team pulled features out of each trace, ranked them by how much they revealed about the expert's grade, and trained small neural networks to predict that grade. The final model uses a handful of features: the share of the signal in the ribosomal region, the height of the 28S peak, how much material shows up in the "fast region" where short fragments run, and the height of a marker peak. The output of that model is the RIN.
So the RIN is a prediction of what an experienced person would have said about a trace. The authors are candid about what that implies: "Note that the degradation of RNA is a continuous process, which implies that there are no natural integrity categories." The ten steps were drawn by people, and the model learned to reproduce them. The authors report that the model's error was "as low as the natural noise in the target values," meaning it disagreed with the experts about as much as the experts' own borderline calls were uncertain. Its weakest separation was between neighboring grades 9 and 10.
And the thing being graded is mostly ribosomal RNA. The paper notes that the largest peaks in a normal trace are the 18S and 28S bands, and the features the model leans on describe those peaks and the smear of fragments below them. If your experiment is about messenger RNA, the RIN tells you about it indirectly: it reports on the ribosomal RNA in the same tube and assumes the rest aged along with it.
It worked, and here is how well
This is a good method, and it was a real improvement. The paper checked RIN against real-time PCR on four housekeeping genes in samples of varying quality. RIN cleanly separated samples that gave strong expression signals from those that didn't. The old 2.0 ratio rule, applied to the same data, would have rejected about 40 experiments of good quality. Swapping a squinted-at gel for a reproducible, automated score was the right call. I made a similar argument in what the 260/280 ratio measures: a quality number is a tool built for one job, and trouble starts when it gets used for another.
Where the gate leaks
The question a sequencing lab really cares about is not "would an expert grade this a 9?" It is "will this sample give me true gene expression?" Those overlap, but they are not the same question. A 2014 study in BMC Biology by Irene Gallego Romero, Athma Pai, Jenny Tung and Yoav Gilad tested the gap directly.
They took blood cells (PBMCs) from four people and left the samples at room temperature for different lengths of time before extracting RNA, from 0 to 84 hours. The mean RIN fell from 9.3 at the start to 3.8 at 84 hours. At 12 hours the mean RIN was still 7.9, a score that passes many labs' gates. Then they sequenced the samples.
By 12 hours, 608 of 14,094 genes (4%) already showed up as differentially expressed compared with the fresh samples, at a 5% false discovery rate. Nothing biological had been done to the cells. The difference was time on the bench. By 84 hours, 9,998 genes (71%) were affected. Worse, the decay was not even: different transcripts degraded at different rates, linked to features like coding sequence length and GC content. The authors found that "standard normalizations failed to account for the effects of degradation."
They also found a fix, and it is the most useful thing in the paper. When they included each sample's RIN as a variable in their statistical model, the spurious differences mostly disappeared: fewer than 50 genes were still flagged between the fresh samples and any later time point. RIN turned out to be a good measure of how degraded the samples were. It just wasn't a good pass or fail line. The authors attach an important condition: the correction works when RIN is not tangled up with the thing you are studying. If your diseased samples are also your more degraded samples, no model can pull those apart cleanly.
The same paper points out that the field never agreed on a cutoff anyway. Published inclusion thresholds had ranged from a RIN of 8 down to 3.95.
How to read a RIN
1. Treat it as a variable, not only a gate. Record it for every sample and put it in the analysis. That is what rescued the degraded data above.
2. Check whether RIN differs between your groups. If cases average 6 and controls average 9, part of your "biology" may be bench time. Fix the sampling, or at minimum report it.
3. Remember what it looks at. It reads ribosomal RNA peaks on one instrument family. If your workflow throws out ribosomal RNA, which I compared in poly(A) selection versus rRNA depletion, the score still describes the input, not the library.
4. Don't read a 9 versus a 10 as meaningful. The model itself had the most trouble telling those two grades apart, because the experts did too.
What I could not confirm
Both studies are older and specific. The RIN paper describes the original 2006 algorithm. I have not checked whether Agilent has changed the model in later software or newer instruments, some of which report their own integrity scores. The degradation study used one cell type left at room temperature; other tissues and other kinds of damage may decay differently.
I make no claim about any particular lab's cutoff. The "7" mentioned at the top is an example of the kind of gate labs use, not a figure from either paper.
The signal
A RIN is a neural network's estimate of how experienced people graded 1,208 ribosomal RNA traces in 2006. It is a big improvement over eyeballing a gel, and it tracks degradation well. It is not a guarantee: in a controlled test, samples still averaging 7.9 had already shifted 4% of their genes, and the fix was to treat RIN as a measured variable rather than a pass mark. The number on your sequencing report is a good witness. Just don't let it be the judge.
Sources
- Andreas Schroeder, Odilo Mueller, Susanne Stocker, Ruediger Salowsky, Michael Leiber, Marcus Gassmann, Samar Lightfoot, Wolfram Menzel, Martin Granzow and Thomas Ragg, "The RIN: an RNA integrity number for assigning integrity values to RNA measurements," BMC Molecular Biology 7:3, 2006, doi:10.1186/1471-2199-7-3. Open access, PMC1413964. (PRIMARY, full text read via Europe PMC. Source for: the 28S:18S ratio of about 2.0 as the old standard and its weak correlation; the quoted "is subjective, hardly comparable from one lab to another"; the Agilent 2100 Bioanalyzer; Agilent author affiliations and Agilent-supplied data; 1,208 samples, about 30% of known tissue origin, the rest "not traceable"; the quoted manual expert assignment to ten categories; the quoted "no natural integrity categories" sentence; the selected features; the quoted generalization-error phrase; AUC 0.96 between categories 9 and 10; the real-time PCR check on four housekeeping genes and about 40 good experiments rejected by the 2.0 ratio.)
- Irene Gallego Romero, Athma A. Pai, Jenny Tung and Yoav Gilad, "RNA-seq: impact of RNA degradation on transcript quantification," BMC Biology 12:42, 2014, doi:10.1186/1741-7007-12-42. Open access, PMC4071332. (PRIMARY, full text read via Europe PMC. Source for: PBMC samples from four individuals held at room temperature for 0 to 84 hours; mean RIN 9.3 at 0 hours, 7.9 at 12 hours and 3.8 at 84 hours; 608 of 14,094 genes (4%) differentially expressed by 12 hours and 9,998 (71%) by 84 hours at 5% FDR; non-random decay associated with coding sequence length and GC content; the quoted "standard normalizations failed" phrase; fewer than 50 genes after including RIN as a covariate; the condition that RIN and the effect of interest not be associated; published thresholds from 8 to 3.95.)
Scope note: the reading-a-RIN checklist is the author's practical summary of the two papers, not a published guideline.
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.