Bridging Ancestry Gaps in Genomic Risk Prediction with Tabular Foundation Models
Das, A.; Cui, Y.
Show abstract
MotivationModels deployed for genomic prediction of diseases perform unevenly across populations, limiting clinical utility. Two factors drive this limitation: large imbalances in sample availability across ancestry groups and non-stationarity of genotype-phenotype effect sizes across the ancestry continuum. While tabular foundation models with in-context learning (ICL) have shown strong sample efficiency in other domains, their effectiveness for genotype-to-phenotype prediction and their robustness to ancestry-driven effect heterogeneity remain unclear. ResultsUsing large, ancestrally diverse biobank data, we show that ICL-capable tabular foundation models reduce performance degradation in under-sampled ancestry groups compared to conventional supervised approaches. However, we find that prevailing models trained on existing synthetic tabular tasks fail when allele effect sizes vary across ancestry space. Treating genetic ancestry as a continuous variable, we introduce an instruction-tuning framework that exposes models to synthetic tasks with ancestry-dependent non-stationary effects. Instruction-tuned models achieve improved and more stable predictive performance across the genetic ancestry continuum, including for individuals distant from in-context exemplars in ancestry space. Availability and ImplementationAll code for instruction-tuning models, synthetic task generation, data wrangling, and model evaluation, is publicly available at https://github.com/ai4pm/Bridging-Ancestry-Gaps-in-Genomic-Risk-Prediction-with-Tabular-Foundation-Models. The final instruction-tuned model (ICL-NS-G2P-proto) is also released in this repository. Detailed documentation is provided, including environment setup instructions and guidelines for running various parts. The instruction-tuning task datasets are available at https://zenodo.org/records/18309187. Contactadas23@uthsc.edu, ycui2@uthsc.edu Supplementary InformationSupplementary data are available online.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- LDAK-KVIK performs fast and powerful mixed-model association analysis of quantitative and binary phenotypes 95%
- MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data 95%
- Fast and flexible joint fine-mapping of multiple traits via the Sum of Single Effects model 95%
Similar papers in this journal
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
- Causal differential expression analysis under unmeasured confounders with causarray 94%
- Clair3-Trio: high-performance Nanopore long-read variant calling in family trios with Trio-to-Trio deep neural networks 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.