Surprisal-based large language models reveal immunologic insights in lobular breast cancer
Majumder, B. P.; Linak, J. A.; Adamson, R.; Aguilera, R. L.; Agarwal, D.; Reitz, Z.; Loiselle, S.; Devarakonda, S.; Clark, P.; Paulson, K. G.; Stanton, S.
Show abstract
In large data sets discovery is often limited to pre-conceived hypotheses and data fishing. Here we tested whether systematic exploration of AI generated hypotheses could uncover clinically meaningful signals in extensively studied data. We deployed AutoDiscovery, a newly launched large language model (LLM) framework designed to search for hypotheses based on surprisal and systematically interrogate complex datasets, on The Cancer Genome Atlas breast cancer cohort. The system did not identify clinically meaningful novel findings without human input. However, a seeded warm-start run with minimal text input from an oncologist revealed multiple interesting and surprising hypotheses. Among these was that a robust immune signature was present across all subtypes of invasive lobular carcinoma (ILC) that exceeded invasive ductal carcinoma (IDC). This observation was independently validated in independent cohorts and confirmed by high-sensitivity multi-immunofluorescence tumor tissue analyses. These results suggest immunotherapy approaches should be tested in ILC including early stage ER+HER2- ILC; these patients are currently excluded from large neoadjuvant immunotherapy trials. They further demonstrate that surprisal-based hypothesis generation frameworks can extract previously unappreciated patterns from deeply interrogated cancer datasets and imply that disease domain experts working with LLMs can derive more meaningful insights from complex data than either could achieve alone.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromatin-informed inference of transcriptional programs in gynecologic and basal breast cancers 94%
- Deactivation of ligand-receptor interactions enhancing lymphocyte infiltration drives melanoma resistance to Immune Checkpoint Blockade 94%
- Dissecting tumor cell programs through group biology estimation in clinical single-cell transcriptomics 94%
Similar papers in this journal
- Improved detection of gene fusions by applying statistical methods reveals new oncogenic RNA cancer drivers 92%
- Virtual patient analysis identifies strategies to improve the performance of predictive biomarkers for PD-1 blockade 92%
- Circulating immune cell phenotype dynamics reflect the strength of tumor-immune cell interactions in patients during immunotherapy 92%
Similar papers in this journal
- AI-Driven Predictive Biomarker Discovery with Contrastive Learning to Improve Clinical Trial Outcomes 94%
- Evolutionary states and trajectories characterized by distinct pathways stratify ovarian high-grade serous carcinoma patients 92%
- The Tumor Profiler Study: Integrated, multi-omic, functional tumor profiling for clinical decision support 92%
Similar papers in this journal
- Systematic Discovery of the Functional Impact of Somatic Genome Alterations in Individual Tumors through Tumor-specific Causal Inference 93%
- Impact of between-tissue differences on pan-cancer predictions of drug sensitivity 93%
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.