In recent years, AI models known as single-cell foundation models (scFMs), such as Geneformer and scGPT, have attracted considerable attention.
Just as large language models such as ChatGPT are pretrained on massive amounts of text and can apply what they learn to a wide range of tasks, scFMs are pretrained on large-scale scRNA-seq datasets, with the aim of using the resulting representations of cells and genes for various downstream analyses.
Why are scFMs attracting so much attention?
scRNA-seq is an extremely powerful technology that allows gene expression to be measured in individual cells, but the resulting data are also difficult to analyze. The data are highly sparse because relatively little information is obtained from each cell. Differences in sequencing depth between cells and batch effects can also be substantial. In addition, the results can change depending on the methods and parameters chosen for normalization, dimensionality reduction, clustering, batch correction, and other analysis steps. In many cases, there is no complete ground truth that tells us which result best reflects the underlying biology.
We discussed these fundamental difficulties of scRNA-seq data in more detail in our previous article, “What Is Reproducibility in Single-Cell RNA-Seq Analysis? Explaining Data Quality and Its Limitations | Revisiting 2019 Insights in 2026.”
One reason scFMs are attracting attention is the hope that large-scale pretraining might help overcome some of these difficulties. If a model is pretrained on enormous amounts of scRNA-seq data collected from many studies, perhaps it can learn gene-expression patterns shared across many cells and capture more generalizable biological features beyond the noise and batch effects specific to individual experiments.
Recent studies, however, suggest that the problem is not so simple.
Batch effects remain even after learning from massive numbers of cells
A 2025 bioRxiv preprint, “Batch Effects Remain a Fundamental Barrier to Universal Embeddings in Single-Cell Foundation Models,” examined how much batch information remains in cell embeddings generated by several single-cell foundation models. The study showed that batch signals remain in the pretrained embeddings of the scFMs evaluated. It also showed that applying simple post-hoc batch-centering to the embeddings can improve alignment.
The authors argue that future scFMs may need explicit batch-effect correction mechanisms to achieve truly universal cell embeddings. This study is currently a bioRxiv preprint and has not yet undergone peer review, but the problem it raises is important.
Why is it so difficult to remove batch effects even after learning from such enormous numbers of cells? Is this simply because current AI models are not yet powerful enough?
Can we really distinguish biological differences from technical ones?
This question points to a fundamental difficulty inherent in batch correction itself. Observed data contain not only the biological differences we want to study, such as those caused by disease, drug treatment, or cell differentiation, but also technical differences caused by sample preparation, measurement date, reagents, instruments, library preparation, sequencing depth, and other factors. What we actually observe is the combined result of all of these influences. In our previous article, “Limitations of Batch-Effect Correction and Normalization in RNA-Seq | Comparing ComBat, VST, TMM, and Quantile Normalization,” we discussed how correcting the data does not necessarily mean that the corrected data more accurately reflect the underlying biology.
Analysis is possible not because technical effects are absent, but because the biological differences we want to investigate are sufficiently large relative to those effects and can be detected reproducibly.
In scRNA-seq, however, relatively little information is obtained from each cell and the data are sparse, making the effects of measurement conditions such as sequencing depth relatively large. In addition, the structure ultimately observed in the data can change substantially depending on the algorithms and parameters used for normalization, dimensionality reduction, clustering, batch correction, and other steps. In other words, one important reason scRNA-seq analysis is difficult is that the effects introduced by measurement and analysis can become too large to ignore relative to the biological structure we want to observe.
With this problem in mind, recent benchmarks of scFMs begin to look somewhat different.
BioLLM: Large-scale pretraining alone is not enough
BioLLM, published in Patterns in 2025, provides a common framework for using and comparing multiple single-cell foundation models with different architectures and implementations. Comparisons across multiple tasks, including zero-shot evaluation and fine-tuning, showed substantial differences in performance among models.
scGPT showed relatively consistent performance across multiple tasks, while other models had different strengths and weaknesses. In other words, simply pretraining on massive amounts of single-cell data does not automatically produce representations that are suitable for every downstream analysis.
This does not mean that foundation models are not useful. Rather, it may suggest that even enormous models trained on tens of millions of cells cannot escape the question of what should be preserved and what should be ignored.
SCMBench: What does “good integration” mean?
SCMBench, published in Nature Communications in 2026, compared 23 methods for single-cell multi-omics integration, including 19 domain-specific models and four foundation models. In addition to foundation models such as scGPT, Geneformer, scFoundation, and UCE, the benchmark evaluated many existing methods, including Harmony, Seurat, LIGER, scVI, scJoint, and GLUE. Importantly, the study did not evaluate methods simply by asking how cleanly they could integrate datasets. In addition to integration accuracy and batch-effect correction, it evaluated the preservation of biological information through analyses such as biomarker detection and trajectory inference.
The results showed that no single method performed consistently well across all evaluation criteria. Performance depended on the dataset and on what was being evaluated. Overall, foundation models did not outperform domain-specific models.
More importantly, simulations in which the strength of batch effects was varied led the authors to identify a fundamental trade-off between batch correction and integration accuracy. Aggressively removing batch effects can also remove biological information, whereas preserving biological information can leave technical artifacts behind.
In other words, the problem cannot be reduced to simply removing as much batch effect as possible. This benchmark also highlights that the question of what to remove and what to preserve lies at the heart of single-cell data integration.
SCMBench also examined an approach in which embeddings from foundation models were used as input to the scVI framework. The choice is therefore not simply between foundation models and conventional specialized algorithms. Another direction is to combine representation learning by foundation models with task-specific analysis. Such combinations may eventually become part of standard analysis workflows.
Standard analysis procedures will still emerge
Our point here is not to decide which algorithm is best. Whether Harmony, scVI, or a future foundation model is preferable will continue to change.
The deeper problem is that researchers must choose some method in practice even though biological and technical differences cannot be completely distinguished and even though “good integration” itself involves trade-offs.
Some degree of common procedures and criteria is useful for comparing results and sharing analytical methods across research groups. As benchmarks and practical experience accumulate, a consensus is likely to emerge: use this method for this type of data, examine these metrics, and consider the data sufficiently corrected when certain conditions are met. Standardization has substantial practical value for comparing research results, facilitating collaboration, and describing analytical methods in the Methods section of a paper.
Even so, we believe it is important to keep in mind that “this procedure is widely accepted” and “this procedure produces a biologically sound result” are entirely different questions.
When standardization becomes proceduralism
Once standard analysis procedures are established, researchers no longer need to reconsider every analytical decision from the beginning. That is a major advantage of standardization. At the same time, however, the procedure itself can be carried out without asking questions such as: Why are we applying this correction? What changed as a result of the correction? Were the differences that disappeared really differences that should have been removed?
If following a standard procedure comes to be treated as though it guarantees the validity of an analysis, the purpose of analysis can shift from exploring biological phenomena in the data to correctly executing what is regarded as the “right” procedure. This can also create the risk of eliminating phenomena that might otherwise have become subjects of investigation before they can even become scientific questions.
The possibility that standardized analysis procedures could hinder new biological discoveries may be particularly important in an era when AI can perform complex analyses quickly and easily. Even if AI and robotics dramatically increase our ability to generate hypotheses and test them experimentally, the range of possibilities considered as hypotheses does not automatically expand. If existing knowledge and standard analysis procedures define a narrow search space, we may simply explore that limited space at unprecedented speed.
There is biology beyond standard analysis
This issue is not limited to batch correction in single-cell analysis. In our previous article, “Genes That Are Not Significant in DEG Analysis Can Still Have Experimentally Validated Functions,” we introduced a study in which genes that were not necessarily prominent in conventional DEG analysis became candidates when combined with other experimental data and ultimately led to functional validation.
Standard analysis methods are useful tools designed to answer particular questions. However, not being selected by a particular method does not mean that something is biologically unimportant.
As AI makes research more efficient, phenomena that are difficult to identify through standard analysis may become even less likely to be explored. Results selected by analytical methods are published and accumulated in knowledge bases, which AI can then use to generate new hypotheses. Through this feed-forward structure, in which selections made by standard analyses shape subsequent hypotheses and AI rapidly tests those hypotheses, research may become more efficient while the range of biology being explored becomes narrower. We discussed a related problem in “Is Network Analysis Necessary in RNA-Seq Analysis? | How an Unvalidated Network Diagram Can Become Scientific Knowledge.”
What skills will researchers need in the AI era?
AI and robotics will rapidly increase the efficiency not only of analysis, but also of the entire process from hypothesis generation to experimental validation. In such an era, the ability to decide what to explore may become more important than ever for researchers.
To do this, researchers need the ability not simply to accept the results they are given, but to examine what is happening in the data, understand how the analysis has changed the data, question the results, and identify new questions from what they observe. Results that fall outside standard analyses, or data that do not match expectations, may point to questions that have not yet been explored.
As AI makes it easier to produce answers, we believe that deciding what questions to ask will become even more important.
Look at the data more closely
So what should we do? Our suggestion is very simple.
Look at the data more closely.
This does not mean avoiding sophisticated algorithms. Nor does it mean avoiding AI. When useful, we should actively use the latest analytical methods and AI.
But do not look only at the final analytical result. Look at the original data. Look at differences between samples. Look at distributions. Compare the data before and after analysis. Look at data that deviate from expectations. And when something catches your attention, return to the original data. Precisely because technical and biological differences cannot be completely distinguished, it is important to examine what is happening in the data and how the analysis has changed them.
Subio Platform is an environment for exploring your data
Today, sophisticated analyses used in research papers can increasingly be performed easily through GUIs and AI. This is a major advance for researchers. However, being able to perform sophisticated analyses is not the same as being able to analyze data.
Subio Platform emphasizes repeatedly examining and thinking about the data rather than running standard analyses as a black box. Researchers can move back and forth between PCA, clustering, distributions across samples, expression patterns of individual genes, and data before and after filtering, returning from analytical results to the underlying data whenever necessary.
When needed, R, Python, and AI can then be used for more advanced analyses. Conversely, results obtained with R, Python, or AI can be brought back into Subio Platform to examine what is actually happening in the data from another perspective. We believe it is important to keep looking at the data themselves while using different tools for different purposes.
As AI makes analytical tools easier to use, the value of simply knowing how to operate those tools will likely decrease. At the same time, the value of observing data, questioning analytical results, and discovering new questions from what we see will increase.
This is the idea behind Subio Platform’s core message: Master Analysis, Not the Tool.
