Benchmarking large language models for cell-free RNA diagnostic biomarker discovery
De Vlaminck, I.; Gaudio, H.; Bliss, A.; Loy, C.; Eweis-LaBolle, D.; Gardella, A.
Show abstract
Large-language models (LLMs) can parse vast amounts of data and generate executable code, positioning them as promising tools for the development of biomarkers and classifiers from high-throughput omics data. Here, we benchmarked six LLMs, OpenAIs o3 and GPT-4o, Anthropics Claude Opus 4 and Claude 3.7 Sonnet, and Googles Gemini 2.5 Pro and Gemini 2.0 Flash, for disease classification based on plasma cell-free RNA (cfRNA) profiles obtained by RNA sequencing. We analyzed data from cohorts of children with Kawasaki disease (KD) or multisystem inflammatory syndrome in children (MIS-C), adults with active tuberculosis (TB) or other non-TB respiratory conditions, and individuals with myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) or sedentary lifestyle. We assessed two tasks: (i) gene-panel design, where each LLM mined public knowledge to nominate diagnostic genes for use in machine learning (ML), and (ii) end-to-end modeling, where LLMs built an ML workflow directly from raw RNA-seq counts. In the first task, the LLM-derived panels captured canonical immune pathways and outperformed randomly selected genes in all cohorts. They underperformed panels chosen by differential gene expression (DGE) analysis in the KD vs. MIS-C and ME/CFS cohorts but performed comparably or better for the TB cohort. In the second task, o3 produced classifiers for KD vs. MIS-C that performed just as well as conventional statistical methods without human intervention. Performance for TB and ME/CFS cohorts was slightly lower than the conventional approach. These findings delineate current capabilities and limitations of LLMs in diagnostics and open a path for their future use in biomarker discovery.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Community assessment of methods to deconvolve cellular composition from bulk gene expression 95%
- Biochemical-free enrichment or depletion of RNA classes in real-time during direct RNA sequencing with RISER 94%
- Deep representation learning for clustering longitudinal survival data from electronic health records 94%
Similar papers in this journal
- Pan-cancer detection of driver genes at the single-patient resolution 93%
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 93%
- Diagnostic Evidence GAuge of Single cells (DEGAS): A flexible deep-transfer learning framework for prioritizing cells in relation to disease 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.