How RNA Sequencing Is Used to Study Gene Expression

Every cell in your body carries essentially the same DNA, yet a neuron, a liver cell, and an immune cell look and behave in completely different ways. The difference lies in gene expression: which genes are switched on, how strongly, and when. Reading the genome tells you what a cell could do. Measuring gene expression tells you what it is actually doing at a given moment. RNA sequencing, usually shortened to RNA-seq, has become the leading method for taking that measurement across the entire genome at once.

Making sense of RNA-seq data is a bit like how an insurer assesses risk. An analyst pricing car insurance would never judge a driver from a single trip. The signal comes from patterns across thousands of records. In the same way, an RNA-seq experiment does not look at one gene in isolation. It measures tens of thousands of transcripts simultaneously and uses statistics to find the patterns that separate a diseased tissue from a healthy one, or a treated sample from an untreated one.

Technically, the method is simple in concept. RNA is extracted from cells or tissue, converted into complementary DNA (cDNA), and sequenced on a high-throughput machine. The resulting reads are mapped back to a reference genome or transcriptome, and the number of reads assigned to each gene reflects how active that gene was. More reads generally mean higher expression, although several technical steps are needed before the numbers can be compared reliably.

This article explains how RNA-seq is used to study gene expression, from the biological basis to the laboratory workflow, the computational analysis, and the main applications. It also covers the major variants of the technology, common pitfalls, and best practices, so you can understand what an RNA-seq result means and how much confidence to place in it.

Gene Expression and the Transcriptome

Gene expression is the process by which the information in a gene is used to make a functional product. For most genes, the sequence is transcribed from DNA into messenger RNA (mRNA), which is then translated into protein. Other genes produce non-coding RNAs, such as ribosomal RNA, transfer RNA, microRNAs, and long non-coding RNAs, that carry out regulatory or structural roles without being translated.

The transcriptome is the complete set of RNA molecules present in a cell or tissue at a specific time. Unlike the genome, which is largely fixed, the transcriptome is dynamic. It shifts with cell type, developmental stage, environment, disease, and treatment. A cancer cell exposed to a drug, a plant cell facing drought, and a liver cell after a meal will each show a different transcriptional profile.

Changes in RNA levels are often among the earliest and most sensitive signs of a biological response, which points researchers toward the genes and pathways involved. One caveat is that RNA levels do not always predict protein levels, because translation, protein stability, and chemical modification add further layers of regulation. RNA-seq shows what the cell is transcribing, a strong but incomplete proxy for what it is doing.

From Microarrays to RNA-Seq

Before RNA-seq, most large-scale expression studies used microarrays. A microarray is a chip carrying thousands of short DNA probes, each designed for a known gene. Labeled cDNA from a sample binds to matching probes, and the fluorescence intensity indicates how much of each transcript is present.

Microarrays were transformative, but they could only detect sequences the designers had already placed on the chip. They also had a limited dynamic range and suffered from background noise and cross-hybridization. When next-generation sequencing was applied to RNA in the late 2000s, it avoided most of these problems.

RNA-seq does not require prior knowledge of the sequences, so it can discover new transcripts, gene fusions, and unannotated isoforms. It has a wider dynamic range, allowing very low and very high expression levels to be measured in one experiment. It detects alternative splicing and sequence variants, and it works in any organism, even one without a well-established genome. As sequencing costs fell, RNA-seq largely replaced microarrays for discovery work.

The RNA-Seq Workflow

Experimental Design

Good RNA-seq begins before any sample is collected. The central question is what comparison you want to make, such as disease versus healthy, treated versus control, or a time course after stimulation. That question determines how many samples you need and how they should be handled.

Biological replicates are essential. These are independent samples from separate individuals, cultures, or animals, and they let the analysis estimate natural variability. Three replicates per condition is often treated as a practical minimum, with more needed when variability is high, as in human clinical samples. Technical replicates of the same biological sample add little compared with additional biological replicates.

Samples should also be randomized across processing batches, so that a technical effect such as extraction date is not confused with the biological effect under study. Recording metadata such as age, sex, and processing dates makes it possible to account for these factors later.

Sample Collection and RNA Extraction

RNA is fragile. Enzymes called RNases degrade it quickly, so samples must be preserved rapidly by snap-freezing or with stabilizing reagents. RNA is then extracted and treated to remove contaminating DNA.

Quality control follows. Concentration is measured, and RNA integrity is assessed with a score such as the RNA integrity number (RIN), where higher values indicate less degradation. Many standard protocols expect high-integrity RNA, while specialized methods can handle degraded material, such as that from preserved clinical tissue. Poor-quality input can bias results, so flag or exclude such samples rather than interpreting them as if they were reliable.

Library Preparation

Sequencers do not read RNA directly in most workflows, so the RNA must be converted into a sequencing library. Total RNA is dominated by ribosomal RNA, often more than 80 percent, which would consume most sequencing capacity if left in place. Two main strategies deal with this:

  • Poly(A) selection captures mRNAs by their poly-A tails. It enriches for protein-coding transcripts but misses non-polyadenylated RNAs and works poorly on degraded samples.
  • Ribosomal RNA depletion removes rRNA and retains a broader range of RNA species, including many non-coding RNAs and bacterial transcripts, which lack poly-A tails.

The selected RNA is fragmented, reverse-transcribed into cDNA, and ligated to adapters that allow it to bind to the sequencer and identify the sample. Stranded protocols preserve information about which DNA strand the RNA came from, which helps resolve overlapping genes. Most libraries are amplified by PCR, and some include unique molecular identifiers (UMIs), short random tags that let analysts recognize and correct for duplicated fragments created during amplification.

Sequencing

Most RNA-seq uses short-read sequencing, which produces millions of reads of roughly 50 to 150 bases each. Reads may be single-end or paired-end, and paired-end reads give more information about fragment structure, improving alignment and splicing analysis.

Sequencing depth matters. For standard differential expression in human samples, a commonly cited range is about 20 to 30 million reads per sample, though the ideal depth depends on the question. Detecting lowly expressed genes, rare isoforms, or fusions calls for greater depth. Long-read sequencing can read entire transcripts in one pass, which is particularly useful for identifying full-length isoforms.

Bioinformatics: From Raw Reads to Biological Insight

Quality Control, Alignment, and Quantification

Analysts first check read quality, looking at base quality scores, adapter contamination, GC content, and duplication levels. Adapters and low-quality bases are trimmed so they do not interfere with mapping.

Cleaned reads are then mapped to a reference. Splice-aware aligners map reads to the genome while accounting for the fact that reads spanning exon junctions are split across distant genomic locations. This approach is valuable for discovering novel splice junctions and variants. Alternatively, lightweight pseudoalignment methods map reads directly to a transcriptome. They are much faster and estimate transcript abundance efficiently, which suits large projects where the annotation is well established.

After mapping, reads are assigned to genes or transcripts to produce a count matrix, in which rows are genes, columns are samples, and each value is the number of reads. This matrix is the foundation for almost every downstream analysis.

Normalization

Raw counts cannot be compared directly. A sample sequenced more deeply produces more reads for every gene, and longer genes produce more reads than shorter ones at the same expression level. Normalization corrects for these effects.

Commonly reported units include transcripts per million (TPM), which is generally preferred for comparing genes within a sample, and the older RPKM and FPKM values. For differential expression testing, modern methods work from raw counts and apply their own normalization to correct for differences in library size and composition. This matters when a small number of very highly expressed genes changes between conditions.

Differential Expression Analysis

The most common goal is to find genes whose expression differs between conditions. Statistical models designed for count data account for the fact that RNA-seq measurements are discrete and more variable than a simple Poisson distribution would predict.

For each gene, the analysis reports a log fold change, which describes the size of the difference, and a p-value, which describes how unlikely the difference would be if there were no real effect. Because thousands of genes are tested at once, p-values must be adjusted for multiple testing, usually by controlling the false discovery rate (FDR). Researchers then apply thresholds, such as an adjusted p-value below 0.05 combined with a minimum fold change, to define a list of differentially expressed genes. These thresholds are conventions and should be chosen with the biological question in mind.

Visualization and Interpretation

Visual summaries are central to checking and communicating results. Principal component analysis (PCA) shows whether samples cluster by condition or by an unwanted factor such as batch. Heatmaps display expression patterns across genes and samples. Volcano plots show fold change against statistical significance, and MA plots relate average expression to fold change. These plots help catch outliers, mislabeled samples, and batch effects before conclusions are drawn.

A list of hundreds of genes is hard to interpret on its own. Enrichment analysis asks whether the differentially expressed genes are over-represented in known biological processes or pathways drawn from curated databases. Gene set methods go further by ranking all genes and testing whether predefined groups shift together, even when no single gene passes a strict cutoff. Co-expression network analysis groups genes with similar expression patterns into modules and relates them to traits, while deconvolution methods estimate the proportions of different cell types in a bulk sample.

What RNA-Seq Reveals: Key Applications

Understanding disease. RNA-seq is widely used to compare diseased and healthy tissue. In cancer research, it identifies overexpressed oncogenes, silenced tumor suppressors, gene fusions, and expression-based subtypes that can guide classification and treatment. In autoimmune, neurological, and metabolic diseases, it highlights the pathways that go wrong, suggesting biomarkers and therapeutic targets.

Drug discovery and response. Researchers use RNA-seq to see how cells respond to a compound. The transcriptional signature can reveal the mechanism of action, off-target effects, and toxicity signals. Comparing patient samples before and after treatment can also help explain why some individuals respond and others do not.

Development and differentiation. By sampling tissues across time, researchers can map how gene expression changes as an embryo develops or a stem cell differentiates. These time-course experiments show which regulatory programs switch on and off, informing regenerative medicine and developmental biology.

Infectious disease. RNA-seq can capture the response of a host to an infection and, at the same time, the behavior of the pathogen. This dual view is valuable for studying viruses, bacteria, and parasites, and for identifying immune signatures that distinguish different infections.

Splicing and novel transcripts. A single gene can produce multiple transcript variants through alternative splicing, and these isoforms may have different or even opposing functions. Because reads span exon junctions, RNA-seq can detect and quantify these events, and discover new transcripts, non-coding RNAs, and unannotated genes.

Agriculture, ecology, and evolution. In crops, RNA-seq reveals how genes respond to drought, salinity, pests, and nutrient stress, guiding breeding and engineering. In ecology and evolution, it allows gene expression to be compared between species and populations, including organisms with no complete reference genome.

Variants and allele-specific expression. Because reads come from expressed sequences, RNA-seq can identify genetic variants in transcribed regions and reveal allele-specific expression, where the two copies of a gene are expressed at different levels. This connects genetic variation to functional consequences.

Major Types of RNA-Seq

RNA-seq is a family of related methods, each suited to different questions.

ApproachWhat it measuresBest forMain limitation
Bulk RNA-seqAverage expression across all cells in a sampleComparing conditions, discovery, biomarkersHides differences between cell types
Single-cell RNA-seqExpression in individual cellsCell types, rare populations, developmental trajectoriesHigher cost, sparse data, complex analysis
Spatial transcriptomicsExpression with tissue locationTissue architecture, tumor microenvironmentsLower resolution or coverage, depending on method
Long-read RNA-seqFull-length transcriptsIsoforms, fusions, complex splicingHigher error rates and cost, lower throughput
Small RNA-seqmicroRNAs and other short RNAsRegulatory RNA studiesRequires dedicated library preparation

Single-cell RNA-seq has changed how biologists think about tissues. Rather than an average, it shows that a tissue is a mixture of distinct cell types and states, some of them rare, that bulk sequencing would blur together. Spatial methods add the missing context of where each cell sits. Many projects combine approaches, using bulk RNA-seq for statistical power and single-cell or spatial data for resolution.

Challenges and Pitfalls

RNA-seq is powerful, but poorly designed or carelessly analyzed experiments produce misleading conclusions.

Batch effects. Technical variation from different processing dates, reagent lots, or sequencing runs can overwhelm biological signal. Good design and statistical correction help, but no method can fully rescue a study in which batch is perfectly confounded with condition.

Insufficient replication. Too few replicates leave the analysis underpowered, producing both false positives and missed findings. Cutting replicates to save cost is usually a false economy.

Sample quality and heterogeneity. Degraded RNA, contamination, and mixed cell populations can all distort results. In bulk tissue, a change in cell-type composition can look like a change in gene expression, so interpretation must be careful.

Annotation and reference issues. Results depend on the quality of the reference genome and gene annotation. Different annotation versions can give different counts, so record the versions you use.

Overinterpretation. Statistical significance does not equal biological importance, and transcript changes do not always translate into protein-level or functional effects. Key findings should be validated with an independent method, such as quantitative PCR or a functional experiment.

Data and computing demands. Raw sequencing files are large, and analysis requires storage, computing resources, and bioinformatics expertise. Planning for these needs early prevents delays later.

Best Practices for a Reliable Study

  1. Define the question first, and design the experiment, replicate number, and sequencing depth around it.
  2. Use adequate biological replicates, and randomize samples across batches.
  3. Check RNA quality before library preparation, and set clear inclusion criteria.
  4. Apply quality control at every stage, from raw reads to sample-level clustering.
  5. Document your methods, including software versions, parameters, and reference files.
  6. Account for known covariates, such as batch, sex, or age, in the statistical model.
  7. Validate important findings with an independent method or dataset.
  8. Share data and code, depositing raw and processed data in public repositories with complete metadata so others can reproduce the work.

The Future of Gene Expression Profiling

RNA-seq continues to evolve quickly. Single-cell and spatial technologies are scaling to millions of cells and larger tissue areas, and long-read sequencing is improving in accuracy, bringing full-length isoform analysis into routine use. Multi-omic methods that measure RNA alongside DNA accessibility, proteins, or epigenetic marks in the same cell are connecting regulation to outcome. Machine learning is being applied to cell-type annotation, expression prediction, and the integration of large public datasets. In the clinic, transcriptomic tests are gradually expanding, particularly in oncology, where expression profiles help classify tumors and guide therapy.

Conclusion

RNA sequencing has changed how scientists study gene expression. By converting RNA into sequence data and counting reads across the genome, it provides an unbiased, quantitative view of which genes are active, how strongly, and in which form. From experimental design and library preparation to mapping, normalization, and differential expression testing, each step shapes the reliability of the final result.

Used carefully, RNA-seq helps explain disease mechanisms, reveal drug effects, trace development, and uncover new transcripts and regulatory patterns. Used carelessly, it can produce convincing but misleading results. The technology is only as good as the experiment behind it: sound design, adequate replication, careful quality control, and honest interpretation. With those foundations in place, RNA-seq remains one of the most informative tools available for understanding what our genes are doing.

This version has no hyperlinks and no brand, product, or software names, only generic method descriptions. "Car insurance" still appears once, in the second paragraph. I can swap that analogy for a different framing, or add a meta title and description.


Reply

About Us · User Accounts and Benefits · Privacy Policy · Management Center · FAQs
© 2026 MolecularCloud