My own research is microwave spectroscopy, which is a different instrument and a different molecule, but it taught me one habit that transfers everywhere: before you interpret a spectrum, you work out what your setup was incapable of showing you. A spectrometer has a bandwidth. Outside it, the answer is not "no signal." The answer is "no measurement," and those two look identical on the page.
RNA sequencing has a bandwidth too. It is just installed earlier, at the bench, in a step most people think of as cleanup rather than as measurement. I wrote in Report 062 that "not detected" is a statement about an instrument rather than about a sample. This is the same argument moved upstream, into the pipette.
The problem the protocol is solving
When you extract total RNA from a cell, the large majority of it is ribosomal RNA. Sequence that directly and you spend nearly all your reads counting the same few ribosomal species over and over. So every standard protocol removes it first, and there are two dominant strategies.
Poly(A) selection uses oligo-dT to grab RNAs by their poly(A) tail. Mature messenger RNAs have that tail; ribosomal RNA does not. Pull on the tails and you have enriched for mRNA.
Ribosomal depletion goes the other way, using probes complementary to ribosomal sequences to subtract rRNA specifically and keep whatever remains. Guo and colleagues describe the practical consequence precisely: in the total RNA library preparation, only ribosomal RNAs and small RNAs are washed out.
Those two sentences sound like a matter of taste. They are not. Poly(A) selection defines its target by a structural feature, and any RNA lacking that feature is not depleted, it is simply never selected. Ribosomal depletion defines its target by sequence, and anything not matching a probe survives, wanted or not. One is a positive selection with a hard structural criterion; the other is a negative selection with a sequence criterion. They have different passbands, and the difference is not subtle.
The genes with no tail to pull on
The cleanest demonstration comes from a 2015 study by Yan Guo and colleagues at Vanderbilt, published in BioMed Research International. They built both library types from the same two breast cancer cell lines, HS578T and BT549, with RNA integrity numbers of 10, and compared what each one saw.
Expression values agreed well between methods, which is the reassuring part: Pearson correlations of 0.92 for protein-coding RNA, 0.79 for lncRNA and 0.69 for other RNAs. For the genes both methods see, they broadly agree on the numbers.
The disagreement is about membership rather than magnitude. The total RNA libraries detected significantly more RNAs at every detection threshold the authors tried, and the authors went looking for a mechanism behind one specific subset:
It has been shown that not all mRNAs necessarily contain a poly(A) tail at their 3′ ends. For example, the mRNA that encodes histone proteins is nonpolyadenylated.
They pulled the 38 histone-encoding genes out of ENSEMBL and ran enrichment analysis. Histone genes came out strongly enriched in the total RNA data, captured at much higher efficiency than in the poly(A) libraries.
Sit with what that means operationally. Replication-dependent histone genes are not obscure. They are protein-coding, they are biologically central, and their expression is tightly coupled to the cell cycle, which makes them exactly the sort of thing a proliferation study would want. In a poly(A)-selected experiment they arrive attenuated, for a reason that has nothing to do with the biology of your samples and everything to do with the chemistry of your capture step. Nothing downstream flags this. The counts are low. Low counts look like low expression.
This is a close relative of the failure I described in Report 038: a reagent that defines your target by one physical property will be silently blind to every target that lacks it, and the resulting absence is indistinguishable from a real negative unless you already knew to check.
The number that gets quoted, and what it actually counts
Guo's most dramatic figure is about novel transcripts. Running Cufflinks to assemble transcripts not in the existing annotation:
The two poly(A) library samples detected 4122 and 6169 potential new transcripts, and the two total RNA samples detected 53282 and 58111 potential new transcripts, roughly a 10-fold increase.
That is a real result and I would not dismiss it, but it deserves an honest reading, because it is the number most likely to be repeated without its context. These are assembled candidate transcripts, not validated genes. A great many of them will be unspliced pre-mRNA and intronic fragments, which ribosomal depletion retains by design and poly(A) selection largely excludes. The tenfold gap is a true statement about how much more material survives ribosomal depletion. It is not a claim that you discovered fifty thousand genes.
Which brings us to why the second paper reads like a contradiction and is not one.
The paper that recommends the opposite
In 2018, Shanrong Zhao and colleagues at Pfizer published a comparison in Scientific Reports aimed squarely at clinical work. They used human blood, pooled from five healthy male volunteers, and colon tissue from a single donor, with four technical replicates per protocol and 50 million reads sampled from each replicate library. Their conclusion runs against Guo's emphasis:
Our analyses showed that rRNA depletion captured more unique transcriptome features, whereas polyA+ selection outperformed rRNA depletion with higher exonic coverage and better accuracy of gene quantification.
The mechanism is where the reads land. In the blood samples, 78% of reads in the ribosomal-depletion libraries came from outside known exons, against 29% for poly(A) selection. Fully half of all mapped reads sat in introns under ribosomal depletion, compared with 6% under poly(A) selection.
That has two costs. The obvious one is efficiency: reads in introns are not counting your exons, so you need more of them. The authors quantified it, finding that 220% more reads for blood and 50% more for colon would have to be sequenced under ribosomal depletion to reach the same exonic coverage as poly(A) selection. Sequencing depth is money, and this is a direct multiplier on it.
The second cost is the one I would put on a lab wall, because it is an accuracy problem rather than a budget problem. All that intronic signal does not merely sit there being useless. Where one gene's exons overlap another gene's introns, those intronic reads get attributed to the wrong feature, and the authors state that this led to overestimation of the expression levels for the genes that overlapped with the intronic regions of other genes. Ribosomal depletion does not just cost you reads. In specific, predictable places, it inflates specific genes.
Their recommendation is blunt, and correctly fenced: in most cases they strongly recommend poly(A) selection over ribosomal depletion for gene quantification in clinical RNA sequencing. Note the last five words. This is advice about a goal, not a verdict on a method. They also flagged a separate practical nuisance, that a small number of lncRNAs and small RNAs consumed a large fraction of reads in the depleted libraries, and suggested depleting those specifically.
Why both papers are right
Put the two side by side and the apparent conflict dissolves into a question about the question.
Guo asked whether total RNA sequencing is more useful for lncRNA and novel transcript discovery. Answer: yes, substantially, and it also recovers non-polyadenylated protein-coding genes that poly(A) selection misses. Zhao asked whether ribosomal depletion is a better way to quantify protein-coding genes in clinical samples. Answer: no, it is less efficient and less accurate for that purpose. Both are correct. Neither method is better. Each is a filter, and choosing one is a declaration about which RNAs you are treating as the signal and which you are treating as the background.
The trouble is that the declaration is usually made by whoever wrote the core facility's default protocol, years ago, for someone else's project. It is then inherited silently, and by the time a result is being interpreted, nobody in the room is thinking of the kit as a hypothesis. This is the same organizational failure mode I described in Report 098, where a physical property of the plate quietly became part of the result, and in Report 074, where two labs running "the same" protocol were not running the same experiment.
The caveats these two papers carry
I would rather state the limits than have you find them. Guo's comparison rests on two cancer cell lines with pristine RNA, so it is a mechanism demonstration rather than a survey, and its novel-transcript count is annotation-dependent. Zhao's used technical replicates from one pooled blood sample and one colon donor, which is the right design for characterizing protocols and the wrong design for making claims about biological variation between people. Both used Ribo-Zero Gold chemistry (Zhao used Globin-Zero for blood), and both are now several years old. Depletion chemistries, probe sets and aligners have all moved since, and the specific percentages should be treated as well-measured examples of a durable effect rather than as current constants.
What has not changed is the structural fact underneath both papers, because it is a property of the molecules rather than of the kits: histone mRNAs still have no poly(A) tail, and ribosomal depletion still retains unspliced material. Any protocol will trade against one of those two.
What to ask
If you are reading someone else's transcriptomics, or commissioning your own, three questions do most of the work.
First, ask which prep was used, and treat the answer as part of the results rather than part of the methods boilerplate. If a paper reports on histone genes, replication-dependent transcripts, many lncRNAs, or anything non-polyadenylated, and used poly(A) selection, the low numbers may be an artifact of capture. If it reports surprisingly high expression for a gene that happens to sit inside another gene's intron, and used ribosomal depletion, ask about intronic misattribution.
Second, ask whether the comparison being drawn crosses a protocol boundary. Comparing poly(A) samples to depleted samples, or merging public datasets built both ways, imports a systematic difference into what looks like biology. Correlations of 0.92 are high enough to make this feel safe and not high enough to make it safe.
Third, if you are choosing, start from the RNA you care about rather than from the default. Quantifying protein-coding genes with intact samples points toward poly(A) selection, and Zhao's paper is a direct argument for that. Non-coding RNA, nascent transcription, viral RNA, degraded or FFPE material, or anything where you do not yet know what you are looking for points toward ribosomal depletion, with the read budget and the intronic caveat priced in.
The general principle is the one I keep returning to, and it is not specific to sequencing. Every enrichment step is a measurement decision wearing the costume of a cleanup step. The moment a protocol removes something, it has defined what counts as signal, and that definition will silently propagate through every figure downstream. The instrument is not where the measurement starts. The measurement starts at the first thing you throw away.
Sources
- Yan Guo, Shilin Zhao, Quanhu Sheng, Mingsheng Guo, Brian Lehmann, Jennifer Pietenpol, David C. Samuels and Yu Shyr (Vanderbilt University), "RNAseq by Total RNA Library Identifies Additional RNAs Compared to Poly(A) RNA Library," BioMed Research International, vol. 2015, article 862130, 2015. DOI 10.1155/2015/862130, PMID 26543871. (Primary source, open access. Full text retrieved through the Europe PMC REST API and read. Source of: the two breast cancer cell lines HS578T and BT549 and their RIN of 10; the Ribo-Zero Magnetic Gold depletion step; the statement that only ribosomal RNAs and small RNAs are washed out in the total RNA preparation; the Pearson correlations of 0.92 for protein-coding RNA, 0.79 for lncRNA and 0.69 for other RNAs; the finding that total RNA libraries detected significantly more RNAs at all detection thresholds; the block quote on histone mRNA being nonpolyadenylated; the 38 ENSEMBL histone-encoding genes and their enrichment in the total RNA data; and the block quote giving the 4122 / 6169 versus 53282 / 58111 novel transcript counts. The characterization of those counts as annotation-dependent Cufflinks assemblies likely to include unspliced pre-mRNA is my interpretation, not a claim made by the authors.)
- Shanrong Zhao, Ying Zhang, Ramya Gamini, Baohong Zhang and David von Schack (Pfizer Worldwide Research and Development), "Evaluation of two main RNA-seq approaches for gene quantification in clinical RNA sequencing: polyA+ selection versus rRNA depletion," Scientific Reports 8:4781, 19 March 2018. DOI 10.1038/s41598-018-23226-4, PMID 29556074. (Primary source, open access. Full text retrieved through the Europe PMC REST API and read. Source of: the blood pooled from five healthy male volunteers and the single-donor colon tissue; the four technical replicates per protocol and the 50 million reads sampled per replicate; the Ribo-Zero Gold and Globin-Zero kits; the block quote on rRNA depletion capturing more unique transcriptome features while polyA+ selection gave higher exonic coverage and better quantification accuracy; the blood figures of 78% versus 29% of reads outside known exons and 50% versus 6% in introns; the 220% and 50% additional reads required for blood and colon respectively to match exonic coverage; the finding that intronic reads led to overestimation of expression for genes overlapping the intronic regions of other genes; the recommendation of polyA+ selection for gene quantification in clinical RNA sequencing; and the observation that a small number of lncRNAs and small RNAs consumed a large fraction of reads in the depleted libraries.)
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.