The Limitations of TabPFN for High-Dimensional RNA-seq Analysis
Zhou, S.; Agarwal, V.; Gopinath, A.; Kassis, T.
10.1101/2025.08.15.670537 bioRxivShow abstract
Tabular Prior-Data Fitted Networks (TabPFN) demonstrate remarkable performance on small-to-medium tabular datasets through in-context learning, but struggle with high-dimensional genomic data such as RNA-seq with tens of thousands of features. We investigate multiple approaches to adapt TabPFN for transcriptomic analysis using two benchmark datasets: Age-ARCHS4, a regression dataset derived from the ARCHS4 dataset (57,873 samples, 10,000 genes), and an Inflammatory Bowel Disease (IBD) classification dataset encompassing Crohns Disease and Ulcerative Colitis samples (2,490 samples, 10,000 genes). Our experimental design proceeds in two phases: first evaluating existing optimization methods, then testing novel adaptations including (1) self-supervised embedding learning and (2) Bulk-Former integration. We demonstrate that when constrained to equal training conditions (500 features, 10,000 samples), TabPFN outperforms classical baselines like random forest and XGBoost. However, when classical methods utilize full feature sets while TabPFN adaptations attempt to handle higher-dimensional data, all TabPFN variants consistently underperform the naive baseline. Our findings reveal fundamental limitations in current approaches to adapting TabPFN for genomic applications, showing that architectural modifications paradoxically degrade performance, while intelligent metadata-based subgrouping emerges as the most effective strategy for deploying TabPFN on biological data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 95%
- Accessible, Reproducible, and Scalable Machine Learning for Biomedicine 94%
- Novel feature selection methods for construction of accurate epigenetic clocks 94%
Similar papers in this journal
- HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata 94%
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 94%
- DeeplyEssential: A Deep Neural Network for Predicting Essential Genes in Microbes 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.