In RNA-seq analysis, a widely used standard workflow is to identify differentially expressed genes (DEGs) and then perform GO or pathway analysis. These analyses are useful for organizing large numbers of gene expression changes into biologically interpretable patterns. Instead of examining genes one by one, they can reveal broader changes related to processes such as the cell cycle, inflammatory responses, metabolism, or cell death.
However, the analysis that should come next depends on the research question. Identifying what changes most strongly in a disease, finding sets of genes or biomarker candidates that characterize a cellular state, and searching for factors closer to the cause of a biological phenomenon are different goals and do not necessarily call for the same analyses. Nevertheless, the workflow RNA-seq → DEG → GO analysis → pathway analysis is sometimes treated as if it were simply the standard set of steps to follow in RNA-seq analysis.
In a previous article, “Have RNA-Seq Data Analysis Methods Been Too Biased? | Possibilities Beyond DEG Analysis,” we discussed how standard analysis methods can influence which parts of the data we choose to examine. Here, we focus on commonly used GO and pathway analyses after DEG identification and consider what kinds of research questions they are actually suited for.
If We Already Knew How to Find Causes from Omics Data
Consider this for a moment. If a well-established analysis method could already identify, with high probability, upstream factors close to the cause of a biological phenomenon from RNA-seq or other omics data, molecular biology research would probably have changed much more dramatically over the past decade or so.
RNA-seq and other omics approaches are already used in a vast number of studies. Computational methods and databases have advanced substantially, and many analyses, including GO and pathway analysis, can now be performed relatively easily. Even so, narrowing down candidates that may be close to the cause remains a major research challenge. Researchers still combine multiple types of data and prior knowledge to identify candidates and then test them experimentally.
Finding “what has changed” in omics data and determining “why the phenomenon occurred” are fundamentally different tasks. With that distinction in mind, it may be worth reconsidering how we choose analysis methods.
So what are GO and pathway analyses, which are routinely used in RNA-seq studies, actually examining?
What Does GO Analysis Examine?
A typical GO enrichment analysis tests whether particular GO terms are represented more frequently than expected in an input gene set.
For example, if genes annotated with “DNA repair” occur more frequently among 100 upregulated genes than expected based on the background gene set, GO terms related to DNA repair may be enriched. In other words, GO analysis directly addresses the question, “What known functions are overrepresented in this group of genes?” It is not working backward to determine what caused the phenomenon.
Conventional pathway enrichment analysis works on a similar principle. It uses the relationships between altered genes and known pathways to organize observed changes into biologically meaningful units.
Large Downstream Changes Are Easier to Detect
Within a cell, an upstream change can lead to altered expression of many genes. For example, an upstream signal may change the activity of a transcription factor, which then affects the expression of many target genes and ultimately alters cellular function or phenotype.
RNA-seq is particularly good at detecting parts of this process in which many genes change together. By contrast, a factor close to the cause of the phenomenon may show little or no change in its own mRNA abundance. Its function may instead change through ligand binding to a receptor, protein phosphorylation, altered enzyme activity, nuclear translocation, or protein–protein interactions.
As a result, when conventional GO or pathway analysis is applied to RNA-seq data, large downstream responses produced by a causal event may stand out more clearly than the causal factor itself. This makes these analyses well suited to questions such as which processes are strongly altered in a cell or tissue, or what functional patterns characterize a disease state or treatment condition. If the goal is to identify factors closer to the cause, however, that is not the type of question these analyses are best suited to answer.
GO and pathway analysis are therefore better viewed as analyses selected according to the research objective, rather than as mandatory standard steps that should automatically follow every RNA-seq experiment.
“Upstream” in a Pathway Diagram Does Not Mean Closer to the Cause of the Phenomenon
Pathway diagrams often arrange molecules according to known signaling relationships, for example from receptors to kinases, transcription factors, and target genes. However, being “upstream” in a pathway diagram does not mean being closer to the cause of the phenomenon being studied.
Even if a pathway is enriched, the receptor shown at the top of that pathway does not necessarily explain what caused the biological phenomenon of interest. A factor that is upstream in the causal sense could appear in the middle or even toward the downstream end of a pathway diagram, and that would not be surprising. Neither GO analysis nor pathway analysis is designed to determine experiment-specific causal relationships.
GO and Pathway Analysis Use Existing Biological Knowledge
GO and pathway analyses do not discover biological functions from the data alone. They interpret observed data using frameworks built from accumulated biological knowledge. GO annotations are based on previous experiments, publications, databases, computational inference, and other sources. Pathway databases similarly organize knowledge accumulated through previous research. Their results therefore reflect, at least in part, what has historically been studied.
Fields that have received more research attention, as well as a relatively small number of well-studied genes, tend to accumulate more annotations, while poorly studied genes and phenomena have less annotation available. This does not necessarily reflect differences in biological importance. It reflects differences in how much knowledge has accumulated.
Another problem is that the biological context in which a relationship was observed can become simplified as knowledge is incorporated into databases. For example, an “A → B” relationship demonstrated in a particular cell type or under a specific condition may not operate in the same way in another cell type or condition. Nevertheless, GO annotations or pathway databases may represent that relationship in a more general form. The lack of sufficiently condition-, tissue-, and cell-specific information in pathway knowledge bases has long been recognized as a challenge for pathway analysis.
In addition, the studies and data analyses underlying existing biological knowledge are not all equally reliable. For example, a study that reviewed 186 papers using functional enrichment analysis reported numerous problems, including the use of inappropriate background gene sets or failure to report them, inadequate correction for multiple testing, and insufficient reporting of analysis parameters. In other words, the existing knowledge used in GO and pathway analysis reflects not only biases in research attention and biological context, but also variation in the quality of the analyses through which that knowledge was generated.
In other words, the knowledge systems used in GO and pathway analysis have limitations, including biases in what has been extensively studied, simplification of the contexts in which biological relationships hold, and variation in the quality of the underlying studies and analyses. Furthermore, if the phenomenon being studied is genuinely novel, it will not yet be described in existing GO annotations or pathway databases.
What Are GO and Pathway Analysis Good For?
GO and pathway analyses are well suited to organizing hundreds or thousands of gene expression changes into units of known function or biological process. They can help answer questions such as which processes are strongly altered in a disease, what responses occur after drug treatment, what functions distinguish two cellular states, or what functional characteristics are associated with a biomarker or gene signature.
They are less directly suited to questions such as what caused the phenomenon, which factors are closest to its cause, or what new biological relationships exist that have not yet been described in GO or pathway databases. Rather than treating GO and pathway analysis as standard procedures that should automatically be performed after RNA-seq, they should be considered among the analyses selected according to the research question.
How Can We Search for Factors Closer to the Cause?
At present, there is no general method that can reliably identify factors close to the cause of a biological phenomenon from omics data alone. Rather than relying on RNA-seq alone, combining different types of omics data, such as proteomics, phosphorylation data, chromatin information, and metabolomics, may provide a more promising direction.
For example, one approach is to use ChIP-seq data to search for candidate upstream transcription factors behind gene expression changes observed by RNA-seq. Another is to combine multi-omics data with network analysis to search for candidates that may be difficult to identify using DEG analysis alone. However, these approaches are still possible ways of addressing the question, “How can we find the cause, or factors close to it?” They are not a general answer to that question.
If the goal is to investigate the cause, we cannot rely on a standard procedure that promises, “Run this analysis and you will get the answer.” What needs to be examined in order to move closer to the cause depends on the biological phenomenon itself, and researchers must decide what evidence is needed in each case.
In the AI Era, Should Beginners Still Start with Standard Workflows?
Until relatively recently, it was difficult to find analysis methods appropriate for a research question, understand their principles and algorithms, and then get them working in practice. For that reason, learning a standard workflow first and gradually expanding from it was a practical way to learn data analysis.
Today, however, AI can assist with finding analysis methods, understanding their principles and algorithms, writing R or Python code, diagnosing errors, and solving implementation problems. If so, does a beginner still need to start by memorizing a standard sequence such as “DEG → GO analysis → pathway analysis”? Perhaps the more important starting point is what you actually want to know from your research.
We discuss this way of learning RNA-seq analysis in more detail in “How to Learn RNA-Seq Data Analysis in the AI Era | From Using Tools to Judgment and Verification.”
DEG analysis followed by GO and pathway analysis is widely used as a standard RNA-seq workflow. But if research objectives differ, the analysis methods selected should differ as well. It may be more useful to remain aware that following standard methods can sometimes narrow the range of possibilities that might otherwise be discovered in the data.
We explore this broader issue in “Data Analysis in the AI Era | Are Standard Analysis Procedures Biologically Sound?,” which considers how standardized workflows may shape the range of questions we explore.
First, try asking AI:
“This is what I want to find out. What kinds of analysis methods might be useful?”
But do not accept the first answer uncritically.
Always ask:
“What are the limitations of this method?”
“Under what conditions can this method be applied?”
and check those points as well.