JCO Clinical Cancer Informatics
● American Society of Clinical Oncology (ASCO)
All preprints, ranked by how well they match JCO Clinical Cancer Informatics's content profile, based on 22 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Adamson, B. J.; Waskom, M.; Blarre, A.; Kelly, J.; Krismer, K.; Nemeth, S.; Gipetti, J.; Ritten, J.; Harrison, K.; Ho, G.; Linzmayer, R.; Bansal, T.; Wilkinson, S.; Amster, G.; Estola, E.; Benedum, C. M.; Fidyk, E.; Estevez, M.; Shapiro, W.; Cohen, A. B.
Show abstract
BackgroundAs artificial intelligence (AI) continues to advance with breakthroughs in natural language processing (NLP) and machine learning (ML), such as the development of models like OpenAIs ChatGPT, new opportunities are emerging for efficient curation of electronic health records (EHR) into real-world data (RWD) for evidence generation in oncology. Our objective is to describe the research and development of industry methods to promote transparency and explainability. MethodsWe applied NLP with ML techniques to train, validate, and test the extraction of information from unstructured documents (eg, clinician notes, radiology reports, lab reports, etc.) to output a set of structured variables required for RWD analysis. This research used a nationwide electronic health record (EHR)-derived database. Models were selected based on performance. Variables curated with an approach using ML extraction are those where the value is determined solely based on an ML model (ie, not confirmed by abstraction), which identifies key information from visit notes and documents. These models do not predict future events or infer missing information. ResultsWe developed an approach using NLP and ML for extraction of clinically meaningful information from unstructured EHR documents and found high performance of output variables compared with variables curated by manually abstracted data. These extraction methods resulted in research-ready variables including initial cancer diagnosis with date, advanced/metastatic diagnosis with date, disease stage, histology, smoking status, surgery status with date, biomarker test results with dates, and oral treatments with dates. ConclusionsNLP and ML enable the extraction of retrospective clinical data in EHR with speed and scalability to help researchers learn from the experience of every person with cancer.
Yoshinari, G. H.; Goulart, W. C. S.; Urbano, A. B. O.; Rabello, M. M.; Macedo, S. O.
Show abstract
Clinical decision-making generates vast unstructured data that remain underexploited for trial recruitment. We present Patient2Sentence (P2S), a framework that transforms electronic health records into language-based representations to enable automated eligibility screening for oncology trials. Using synthetic patient records derived from three completed breast cancer studies (KATHERINE, MONARCH, and OLYMPIA), we created 25 virtual patients per trial and compared eligibility classification between full records and their condensed "patient sentences." P2S achieved a mean concordance of 93.2% (95% CI 89.8-96.6%; Cohens {kappa} = 0.91) between sentence-level and full-record decisions while reducing token usage by ~67%. This compression preserved semantic fidelity and reduced computational cost approximately threefold. By encoding heterogeneous clinical data into compact natural-language form, P2S provides a reproducible and efficient approach to patient-trial matching, with potential applications across diverse clinical decision-support systems.
Coroller, T. P.; Sahiner, B.; Amatya, A.; Gossman, A.; Karagiannis, K.; Samala, R. K.; Santana-Quintero, L.; Solovieff, N.; Wang, C.; Amiri-Kordestani, L.; Cao, Q.; Cha, K. H.; Charlab, R.; Cross, F. H.; Hu, T.; Huang, R.; Kraft, J.; Krusche, P.; Li, Y.; Li, Z.; Mazo, I.; Moloney, C.; Paul, R.; Plawinski, J.; Schnakenberg, S.; Serra, P.; Smith, S.; Song, C.; Su, F.; Subramaniam, S.; Tiwari, M.; Vechery, C.; Xiong, X.; Zarate, J. P.; Ziegler, J.; Zhu, H.; Chakravartty, A.; Liu, Q.; Ohlssen, D.; Petrick, N.; Schneider, J. A.; Walderhaug, M.; Zuber, E.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWIn 2020, Novartis Pharmaceuticals Corporation and the U.S. Food and Drug Administration (FDA) started a 4-year scientific collaboration to find novel radiogenomics-based prognostic and predictive factors for HR+/HER2-metastatic breast cancer under a Research Collaboration Agreement. This manuscript aims to detail the guiding principles and methodology for this study. We include a discussion of internal and external clinical, genomics, imaging datasets, data processing workflows, and machine learning model development strategies. We also prospectively define our success criteria to ensure robust scientific outputs. DisclosureThis publication reflects the views of the authors and should not be construed to represent FDAs views or policies.
Nguyen, N. H.; Dodd-Eaton, E. B.; Peng, G.; Corredor, J. L.; Jiao, W.; Woodman-Ross, J.; Arun, B. K.; Wang, W.
Show abstract
PurposeLFSPRO is an R library that implements risk prediction models for Li-Fraumeni syndrome (LFS), a genetic disorder characterized by deleterious germline mutations in the TP53 gene. To facilitate the use of these models in clinics, we developed LFSPROShiny, an interactive R/Shiny interface of LFSPRO that allows genetic counselors (GCs) to perform risk predictions without any programming components, and further visualize the risk profiles of their patients to aid the decision-making process. MethodsLFSPROShiny implements two models that have been validated on multiple LFS patient cohorts: a competing-risk model that predicts cancer-specific risks for the first primary, and a recurrent-event model that predicts the risk of a second primary tumor. Starting with a visualization template, we keep regular contact with GCs, who ran LFSPROShiny in their counseling sessions, to collect feedback and discuss potential improvement. Upon receiving the family history as input, LFSPROShiny renders the family into a pedigree, and displays the risk estimates of the family members in a tabular format. The software offers interactive overlaid side-by-side bar charts for visualization of the patients cancer risks relative to the general population. ResultsWe walk through a detailed example to illustrate how GCs can run LFSPROShiny in clinics, from data preparation to downstream analyses and interpretation of results with an emphasis on the utilities that LFSPROShiny provides to aid decision making. ConclusionSince Dec 2021, we have applied LFSPROShiny to over 100 families from counseling sessions at MD Anderson Cancer Center. Our study suggests that software tools with easy-to-use interfaces are crucial for the dissemination of risk prediction models in clinical settings, hence serving as a guideline for future development of similar models.
McInerney, S.; Gurku, H.; Balasubramanian, R.; Vikram, P.; Bhaskaran, S.; Sekaran, K.
Show abstract
ObjectivesTo evaluate the performance of SROTAS IQ, a custom fine-tuned large language model (LLM), in automating clinical trial eligibility screening for breast cancer patients using synthetic data. MethodsTen breast cancer trials were selected across diverse treatment settings and molecular subtypes. Fifteen synthetic patient summaries per trial were generated, including realistic and enriched eligibility scenarios. Two independent oncologists assessed trial eligibility for each patient, establishing ground truth. SROTAS IQ LLM was evaluated against expert consensus using standard classification metrics. Time-to-verdict was measured to compare clinician effort with automated assessment. ResultsSROTAS IQ demonstrated strong concordance with expert assessments, achieving 90% or greater accuracy in 5 of 10 trials. Across 150 patient-trial evaluations, the model correctly classified 88% of overall eligibility decisions. Performance was highest in trials with moderate complexity and fewer nested criteria, while more intricate protocols showed reduced accuracy. The LLM consistently delivered rapid assessments (<0.5 minutes per patient), with explainable outputs that aligned with clinical reasoning. These findings underscore the models potential to support high-fidelity, scalable trial matching in oncology. ConclusionSROTAS IQ offers a promising approach to automating clinical trial matching in oncology. Further real-world validation is needed to confirm generalisability and integration into clinical practice.
Liu, T.; Wang, X.; Inkman, M.; Hong, J. C.; Waters, M. R.; Zhang, J.
Show abstract
The ability of pre-trained large language models (LLMs) to rapidly master novel natural language processing tasks holds transformative potential. However, pre-trained LLMs often struggle to achieve high performance in specialized domains such as oncology and have the tendency to deliver incorrect information confidently ("hallucinate"), limiting their utility in such contexts. Retrieval-augmented generation (RAG) addresses this limitation by dynamically incorporating authoritative, domain-specific knowledge directly into the LLMs inference process. This approach significantly enhances LLM performance without the typical requirement for extensive fine-tuning or retraining. In this study, we demonstrate the exceptional performance of a minimalist RAG pipeline (without additional model fine-tuning) on radiation oncology board-style examinations. Leveraging a meticulously curated knowledge base sourced from Gunderson & Teppers Clinical Radiation Oncology, Fifth Edition and NCCN guidelines, our model substantially surpassed the performance of contemporary OpenAI models, achieving an outstanding accuracy of 91.5% on the 2021 American College of Radiology (ACR) TXIT examination. This result markedly exceeds the performance benchmarks set by previous LLM-based approaches in this field, which attained a maximum accuracy of 74%. Crucially, our model exhibited robust self-awareness regarding its knowledge boundaries, overcoming a glaring weakness of pre-trained LLMs; questions answered incorrectly were reliably flagged with low confidence scores (mean 4.12/10 vs. 7.36/10 for correct answers), highlighting areas inadequately represented within the RAG knowledge base. This precise uncertainty estimation underscores RAGs unique strength in enhancing not just accuracy, but also the reliability and interpretability of model outputs. We demonstrate that integrating domain-specific knowledge via RAG significantly enhances large language model performance in radiation oncology, enabling reliable confidence scoring previously unattainable with pretrained LLMs. This scalable approach may be well-suited for clinical decision support and medical education. Future efforts will incorporate clinical guidelines and select primary literature to broaden applicability.
Das, R.; Maheswari, K.; Siddiqui, S.; Arora, N.; Paul, A.; Nanshi, J.; Udbalkar, V.; Sarvade, A.; Chaturvedi, H.; Shvartsman, T.; Masih, S.; Thippeswamy, R.; Patil, S.; Nirni, S. S.; Garsson, B.; Bandyopadhyay, S.; Maulik, U.; Farooq, M.; Sengupta, D.
Show abstract
The clinical adoption of Large Language Models (LLMs) in biomedical research has been limited by concerns regarding the quality, accuracy, and reliability of their outputs, particularly in precision oncology, where clinical decision-making demands high precision. Current models, often based on fine-tuned foundational LLMs, are prone to issues such as hallucinations, incoherent reasoning, and loss of context. In this work, we present GeneSilico Copilot, an advanced agent-based architecture that transforms LLMs from simple response synthesizers to clinical reasoning systems. Our approach is centred around a bespoke ReAct agent that orchestrates a suite of specialized tools for asynchronous information retrieval and synthesis. These tools access curated document vector stores containing clinical treatment guidelines, genomic insights, drug information, clinical trials, and breast cancer-specific literature. To leverage large context windows of current LLMs, we implement a hybrid search strategy that prioritizes key information and dynamically integrates summarized content, reducing context fragmentation. Incorporating additional metadata further allows for precise, transparent and evidence-backed reasoning at each step of the thought process. The system ensures that at every stage, the agent can synthesize meaningful, context-aware observations that contribute to a coherent and comprehensive final response that aligns with clinical standards. Evaluations on real-world breast cancer cases show that GeneSilico Copilot significantly improves response accuracy and personalization. This system represents a critical advancement toward making LLMs clinically deployable in precision oncology and has potential applications in broader medical domains requiring complex, data-driven decision-making.
Ramji, V.; Muralitharan, B.; Ram, G.; Ganesan, N.
Show abstract
Patients facing cancer diagnoses must navigate fragmented information spanning prognosis, clinical research opportunities, and genomics-informed therapy. CancerStop.dev is a React-based web platform that consolidates trusted public resources into a single patient-centered interface to support informed decision-making. The platforms built-in ComboReg module presents interactive relative survival estimates up to 10 years after diagnosis by modeling (combined regression) age-at-diagnosis and stage-at-diagnosis using publicly available data from the SEER Program. Users can adjust an age slider and view stage-specific curves, enabling individualized and comprehensible visualizations of survivorship trends. A Clinical Trials module links directly to ClinicalTrials.gov, providing context-aware queries and free-form keyword filtering (e.g., mutations, investigational agents) to surface ongoing studies. The Genes & More module connects to NCBI ClinVar for variant-level insights, facilitating precision medicine exploration when genetic testing results are available. An Approved Drugs module routes to the National Cancer Institute resources listing FDA-approved agents relevant to specific cancers, aiding therapy literacy. A curated search tool (PresciQure) complements these features by streamlining access to the biomedical literature, package inserts of FDA-approved drugs, and related oncology resources. The platform is non-prescriptive, emphasizes transparency regarding data provenance and limitations, and is positioned to incorporate additional cancer types, demographic stratifications, and multi-omic resources. CancerStop.dev aims to empower patients, caregivers, and clinicians with timely, integrated, and navigable information, thereby strengthening their advocacy and encouraging participation in research and precision care. SignificanceCancerStop.dev integrates prognosis, clinical trials, and genomic insights into a single, user-friendly platform, enabling patient and care teams to make faster, better-informed decisions.
Vu, T.; Tran, H. X.; Li, X.; Liu, L.; Li, J.; Du, J. T.; Le, T.
Show abstract
Survival analysis is essential in oncology for modeling time-to-event outcomes such as overall survival and disease recurrence. Traditional approaches, such as the Cox Proportional Hazards (CPH) model, have been widely used due to their interpretability but rely on restrictive assumptions of linearity and proportional hazards, which limit their ability to capture the nonlinear relationships present in high-dimensional genomic data. Recent machine learning methods, including Random Survival Forests (RSF) and deep learning models such as DeepSurv, have improved flexibility and predictive performance but require extensive training data, hyperparameter tuning, and computationally expensive optimization, which hinder their practical use. We propose TabSurv, a novel survival prediction framework that leverages a foundation model for tabular data tasks using in-context learning. TabSurv predicts survival times in a regression setting, allowing rapid adaptation to new datasets with minimal computational cost. The model is trained using only uncensored samples and evaluated with the concordance index (C-index) and stability metrics to assess both accuracy and robustness. We benchmark TabSurv against seven state-of-the-art survival models across 12 breast cancer datasets. The results demonstrate that TabSurv achieves competitive or superior performance, obtaining the best C-index on six datasets and the highest overall stability score. These findings highlight TabSurv as a powerful and efficient tool for breast cancer prognosis using high-dimensional molecular data.
Song, W.; Shbita, L.; Jang, I. J. H.; Starostina, O.; Lewis, R.; Sahli, A.; Floyd, W. R.; Mahin, M.; Rinsurongkawong, W.; Barbon, C. E. A.; Lai, S. Y.; Lee, J. J.; Shah, K.; Chen, M. M.; Hutcheson, K. A.; Fuller, C. D.; Moreno, A. C.
Show abstract
Abstract Purpose Radiologic surveillance is essential for oropharyngeal cancer (OPC) survivors, guiding recurrence detection and follow-up strategies. The Neck Imaging Reporting and Data System provides a standardized framework for post-treatment risk reporting at both the primary tumor site (pNI-RADs) and cervical lymph nodes (nNI-RADS). Comprehensive surveillance additionally requires assessment of disease status, including the primary tumor, nodal involvement, and distant metastases. These clinical results are often embedded as unstructured data within free-text radiology reports. We hypothesized that a large language model (LLM) can reliably extract NI-RADS score criteria and summarize key imaging features from unstructured radiology text, achieving high concordance with expert review. Methods Previously untreated OPC patients who received definitive cancer therapy were identified. Eligible imaging reports included post-treatment head and neck CT, MRI, or FDG PET/CT scans containing narrative and impression text. Examinations lacking narrative or impression text, containing pre-existing NI-RADS annotations, or involving non-surveillance imaging modalities were excluded. A total of 200 reports were randomly selected from 7,076 eligible examinations for manual abstraction using a three-reviewer consensus framework to establish a reference dataset. Using the Palantir Foundry Pipeline Builder, a GPT-5-based LLM was deployed to extract pNI-RADS and nNI-RADS scores, and key imaging features of disease status from these reports. Performance was evaluated using exact agreement and F1-based metrics. Results Agreement for no evidence of disease (score of 1) was 93.3% (126/135; F1 = 0.94) and 90.3% (130/144; F1 = 0.93) for pNI-RADS and nNI-RADS, respectively. For NI-RADS [≥]2, exact category agreement was 73.1% (38/52; macro-F1 = 0.75) for pNI-RADS and 64.3% (27/42; macro-F1 = 0.56) for nNI-RADS. Quadratic weighted {kappa} was 0.81 and 0.59, respectively. For post-treatment disease surveillance variables, agreement was 94.9% (149/157; F1 = 0.87) for primary tumor presence, 89.1% (164/184; F1 = 0.87) for nodal disease presence, and 94.7% (126/133; F1 = 0.70) for distant metastasis detection. Specificity was high across disease-status variables (0.95-0.99), with negative predictive values of 0.95 for primary tumor, 0.87 for nodal disease, and 0.99 for distant metastasis. Conclusions Our LLM-based information retrieval and classification approach for radiographic treatment response from unstructured, multidimensional imaging reports achieved high performance for disease exclusion and moderate performance for detecting suspected residual and/or new disease. This pipeline supports scalable and standardized surveillance data capture for longitudinal monitoring, clinical analytics, and survivorship research in head and neck oncology.
Vu, T.; Tran, H. X.; Liu, L.; Li, J.; Du, J. T.; Le, T.
Show abstract
Neoadjuvant therapy, involving treatment administered before surgery to shrink tumors, significantly impacts breast cancer management. However, current clinical approaches rely predominantly on limited clinical features, leading to suboptimal patient outcomes. To enhance therapeutic decision-making, we propose a novel foundation model-based recommendation framework (FDR) utilizing TabPFN, a deep learning model trained on extensive synthetic tabular data. Our method integrates multi-omics profiles with traditional clinical factors, enabling accurate counterfactual predictions for various drug combinations. Experimental results show that FDR markedly improves personalized treatment recommendations, resulting in a three-fold increase in recovery response rates. This study introduces the first multi-omics-informed neoadjuvant recommendation system, advancing precision oncology and demonstrating effectiveness even with limited patient data.
Cohen, S.; Shamai, G.; Sabo, E.; Cretu, A.; Barshack, I.; Goldman, T.; Bar-Sela, G.; Pearson, A. T.; Huo, D.; Howard, F. M.; Kimmel, R.; Mayer, C.
Show abstract
The OncotypeDX 21-gene assay is a widely adopted tool for estimating recurrence risk and informing chemotherapy decisions in early-stage, hormone receptor-positive, HER2-negative breast cancer. Although informative, its high cost and long turnaround time limit accessibility and delay treatment in low- and middle-income countries, creating a need for alternative solutions. This study presents a deep learning-based approach for predicting OncotypeDX recurrence scores directly from hematoxylin and eosin-stained whole slide images. Our approach leverages a deep learning foundation model pre-trained on 171,189 slides via self-supervised learning, which is fine-tuned for our task. The model was developed and validated using five independent cohorts, out of which three are external. On the two external cohorts that include OncotypeDX scores, the model achieved an AUC of 0.825 and 0.817, and identified 21.9% and 25.1% of the patients as low-risk with sensitivity of 0.97 and 0.95 and negative predictive value of 0.97 and 0.96, showing strong generalizability despite variations in staining protocols and imaging devices. Kaplan-Meier analysis demonstrated that patients classified as low-risk by the model had a significantly better prognosis than those classified as high-risk, with a hazard ratio of 4.1 (P<0.001) and 2.0 (P<0.01) on the two external cohorts that include patient outcomes. This artificial intelligence-driven solution offers a rapid, cost-effective, and scalable alternative to genomic testing, with the potential to enhance personalized treatment planning, especially in resource-constrained settings.
Davoudi, A.; Yu, S.; Doucette, A.; Gabriel, P.; Miller, M.; Williams, H.; Desai, H.; Le, A.; Stoeckert, C.; Maxwell, K.; Mowery, D.
Show abstract
Although NLP has been used to support cancer research more broadly, the development of NLP algorithms to extract evidence of progression from clinical notes to support lung cancer research is still in its infancy. In this study, we trained supervised machine learning classifiers using rich semantic features to detect and classify statements of progression status from radiology exams. Our progression status classifier achieves high F1-scores for detecting and discerning progression (0.80), stable (0.82), and not relevant (0.92) sentences, demonstrating promising performance. We are actively integrating these extractions with structured electronic health record data using ontologies to instantiate a longitudinal model of progression among non-small cell lung cancer patients.
Griffiths, M.; Kubeyev, A.; Laurie, J.; Giorni, A.; Zillmann da Silva, L. A.; Sivasubramaniam, P.; Foster, M. T.; Biankin, A. V.; Asghar, U. S.
Show abstract
Oncology therapeutic development continues to be plagued by high failure rates leading to substantial costs with only incremental improvements in overall benefit and survival. Advances in technology including the molecular characterisation of cancer and computational power provide the opportunity to better model therapeutic response and resistance. Here we use a novel approach which utilises Bayesian statistical principles used by astrophysicists to measure the mass of dark matter to predict therapeutic response. We construct "Digital Twins" of individual cancer patients and predict response for cancer treatments. We validate the approach by predicting the results of clinical trials. Better prediction of therapeutic response would improve current clinical decision-making and oncology therapeutic development.
Miao, B. Y.; Rodriguez Almaraz, E.; Ashraf Ganjouei, A.; Suresh, A.; Zack, T.; Bravo, M.; Raghavendran, S.; Oskotsky, B.; Alaa, A.; Butte, A. J.
Show abstract
BackgroundMolecular biomarkers play a pivotal role in the diagnosis and treatment of oncologic diseases but staying updated with the latest guidelines and research can be challenging for healthcare professionals and patients. Large Language Models (LLMs), such as MedPalm-2 and GPT-4, have emerged as potential tools to streamline biomedical information extraction, but their ability to summarize molecular biomarkers for oncologic disease subtyping remains unclear. Auto-generation of clinical nomograms from text guidelines could illustrate a new type of utility for LLMs. MethodsIn this cross-sectional study, two LLMs, GPT-4 and Claude-2, were assessed for their ability to generate decision trees for molecular subtyping of oncologic diseases with and without expert-curated guidelines. Clinical evaluators assessed the accuracy of biomarker and cancer subtype generation, as well as validity of molecular subtyping decision trees across five cancer types: colorectal cancer, invasive ductal carcinoma, acute myeloid leukemia, diffuse large B-cell lymphoma, and diffuse glioma. ResultsBoth GPT-4 and Claude-2 "off the shelf" successfully produced clinical decision trees that contained valid instances of biomarkers and disease subtypes. Overall, GPT-4 and Claude-2 showed limited improvement in the accuracy of decision tree generation when guideline text was added. A Streamlit dashboard was developed for interactive exploration of subtyping trees generated for other oncologic diseases. ConclusionThis study demonstrates the potential of LLMs like GPT-4 and Claude-2 in aiding the summarization of molecular diagnostic guidelines in oncology. While effective in certain aspects, their performance highlights the need for careful interpretation, especially in zero-shot settings. Future research should focus on enhancing these models for more nuanced and probabilistic interpretations in clinical decision-making. The developed tools and methodologies present a promising avenue for expanding LLM applications in various medical specialties. Key Points- Large language models, such as GPT-4 and Claude-2, can generate clinical decision trees that summarize best-practice guidelines in oncology - Providing guidelines in the prompt query improves the accuracy of oncology biomarker and cancer subtype information extraction - However, providing guidelines in zero-shot settings does not significantly improve generation of clinical decision trees for either GPT-4 or Claude-2
Jonnalagadda, P.; Obeng-Gyasi, S.; Stover, D. G.; Andersen, B. L.; Rahurkar, S.
Show abstract
BackgroundMany patients with triple-negative breast cancer (TNBC), particularly those who are older, Black, or insured by Medicaid, do not receive guideline-concordant treatment, despite its association with up to 4x higher survival. Early identification of patients at risk for rapid relapse may enable timely interventions and improve outcomes. This study applies machine learning (ML) to real-world data to predict risk of rapid relapse in TNBC. MethodsWe trained various ML models (logistic regression, decision trees, random forests, XGBoost, naive Bayes, support vector machines) using National Cancer Database (NCDB) data and fine-tuned them using electronic health record (EHR) data from a cancer registry. Class imbalance was addressed using synthetic minority oversampling technique (SMOTE). Model performance was evaluated using sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV), receiver operating characteristics area under the curve ROC AUC, accuracy, and F1 scores. Transfer learning, cross-validation, and threshold optimization were applied to enhance the ensemble models performance on clinical data. ResultsInitial models trained on NCDB data exhibited high NPV but low sensitivity and PPV. SMOTE and hyperparameter tuning produced modest improvements. External testing on EHR data from a cancer registry had similar model performance. After applying transfer learning, cross-validation, and threshold optimization using the clinical data, the ensemble model achieved higher performance. The optimized ensemble model achieved a sensitivity of 0.87, specificity of 0.99, PPV of 0.90, NPV of 0.98, ROC AUC of 0.99, accuracy of 0.98, and F1-score of 0.88. This optimized model, leveraging readily available clinical data, demonstrated superior performance compared to initial NCDB-trained models and those reported in extant literature. ConclusionsTransfer learning and threshold optimization effectively adapted ML models trained on NCDB data to an independent real-world clinical dataset from a single site, producing a high-performing model for predicting rapid relapse in TNBC. This model, potentially translatable to fast health interoperability resources (FHIR)-compatible workflows, represents a promising tool for identifying patients at high risk. Future work should include prospective external validation, evaluation of integration into clinical workflows, and implementation studies to determine whether the model improves care processes such as timely patient navigation and treatment planning. Author SummaryIn this study, we set out to understand which patients with triple-negative breast cancer might experience a rapid return of their disease. Many people with this aggressive form of cancer do not receive the treatments that are known to improve survival, especially patients who are older, Black, or insured through public programs. Being able to identify those at highest risk early in their care could help health teams provide timely support and ensure that patients receive the treatments they need. To do this, we used information from a large national cancer database to build computer-based models that learn from patterns in patient data. We then refined these models using real medical records from a cancer center to make sure they worked well in everyday clinical settings. After adjusting and improving the models, we developed a tool that can correctly identify most patients who are likely to have a rapid return of their cancer. Our hope is that this type of tool could eventually be built into routine care and help guide timely follow-up, support services, and treatment planning. More testing in real clinical environments will be important to understand how well the tool improves care and outcomes for patients.
Vallon, J. J.; Panjwani, N.; Ling, X.; Vij, S.; Srinivas, S.; Leppert, J.; Bayati, M.; Buyyounouski, M. K.
Show abstract
With rising access to electronic health record data, application of artificial intelligence to create clinical risk prediction models has grown. A key component in designing these models is feature generation. Methods used to generate features differ in the degree of clinical expertise they deploy (from minimal to population-level to patient-level), and subsequently the extent to which they can extract reliable signals and be automated. In this work, we develop a new process that defines how to systematically implement patient-level clinician feature generation (CFG), which leverages clinical expertise to define concepts relevant to the outcome variable, identify each concepts associated features, and finally extract most features on a per-patient level by manual chart review. We subsequently apply this method to identifying and extracting patient-level features predictive of cancer recurrence from progress notes for a cohort of prostate cancer patients. We evaluate the performance of the CFG process against an automated feature generation (AFG) process via natural language processing techniques. The machine learning outcome prediction model leveraging the CFG process has a mean AUC-ROC of 0.80, in comparison to the AFG model that has a mean AUC-ROC of 0.74. This relationship remains qualitatively unchanged throughout extensive sensitivity analyses. Our analyses illustrate the value of in-depth specialist reasoning in generating features from progress notes and provide a proof of concept that there is a need for new research on efficient integration of in-depth clinical expertise into feature generation for clinical risk prediction.
Chen, L.-C.; Zack, T.; Demirci, A.; Sushil, M.; Miao, B.; Kasap, C.; Butte, A. J.; Collisson, E.; Hong, J.
Show abstract
PurposeWe examined the effectiveness of proprietary and open Large Language Models (LLMs) in detecting disease presence, location, and treatment response in pancreatic cancer from radiology reports. MethodsWe analyzed 203 deidentified radiology reports, manually annotated for disease status, location, and indeterminate nodules needing follow-up. Utilizing GPT-4, GPT-3.5-turbo, and open models like Gemma-7B and Llama3-8B, we employed strategies such as ablation and prompt engineering to boost accuracy. Discrepancies between human and model interpretations were reviewed by a secondary oncologist. ResultsAmong 164 pancreatic adenocarcinoma patients, GPT-4 showed the highest accuracy in inferring disease status, achieving a 75.5% correctness (F1-micro). Open models Mistral-7B and Llama3-8B performed comparably, with accuracies of 68.6% and 61.4%, respectively. Mistral-7B excelled in deriving correct inferences from "Objective Findings" directly. Most tested models demonstrated proficiency in identifying disease containing anatomical locations from a list of choices, with GPT-4 and Llama3-8B showing near parity in precision and recall for disease site identification. However, open models struggled with differentiating benign from malignant post-surgical changes, impacting their precision in identifying findings indeterminate for cancer. A secondary review occasionally favored GPT-3.5s interpretations, indicating the variability in human judgment. ConclusionLLMs, especially GPT-4, are proficient in deriving oncological insights from radiology reports. Their performance is enhanced by effective summarization strategies, demonstrating their potential in clinical support and healthcare analytics. This study also underscores the possibility of zero-shot open model utility in environments where proprietary models are restricted. Finally, by providing a set of annotated radiology reports, this paper presents a valuable dataset for further LLM research in oncology.
De La Vega, F. M.; Pouliot, Y.; Rhead, B.
Show abstract
The passage of the US Food and Drug Administration (FDA) Omnibus Reform Act of 2022 underscores a national commitment to enhancing diversity in clinical trials. This commitment recognizes not only the ethical imperative of inclusivity but also the practical necessity to ensure the safety and efficacy of medications across all demographic groups. Particularly for Phase 3 and pivotal clinical trials, the FDA has issued draft guidance that recommends sponsors to develop diversity plans with race and ethnicity (R/E) enrollment targets informed by the epidemiological landscape of the disease in the therapys target population. For biomarker-driven oncology trials, real-world data (RWD), especially when enriched with multimodal clinico-genomic information, holds immense promise for informing these R/E enrollment goals. However, leveraging RWD comes with hurdles, including the overrepresentation of insured patients, significant non-random missingness in R/E data, and disparities between R/E distributions in RWD and disease incidence databases--often attributed to healthcare access and socioeconomic disparities. Here, we propose a robust methodology to harness clinico-genomic RWD, addressing these challenges through strategies that include accurate R/E imputation and incidence adjustment factors. Our approach then utilizes clinical data and biomarker prevalence in RWD to derive a data-driven R/E distribution for clinical trial enrollment targets. Through a case study on a hypothetical biomarker-driven clinical trial targeting prostate adenocarcinoma and leveraging a cohort from the Tempus clinco-genomic database, we demonstrate the application of our methodology. This example illustrates the potential of RWD to offer enrollment target scenarios, grounded in disease epidemiology and empirical R/E distributions adjusted for biomarker prevalence. Such data-driven targets are pivotal for the development of informed and equitable diversity plans in oncology clinical trials, paving the way for more representative and generalizable research outcomes.
Yang, D. X.; Hui, Y.; Wang, B.; Avesta, A.; Zuber, C.; Park, H. S.; Aneja, S.
Show abstract
PURPOSECancer registries are important sources of real-world data (RWD) that reveal insights into practice patterns and cancer patient outcomes, but the prevalence of missing data can be high. Machine learning (ML) imputation methods can be applied to large RWD sets, but the performance of these approaches within cancer registries is unclear. METHODSWe identified non-small cell lung cancer (NSCLC) patients within the National Cancer Database diagnosed in 2014 with complete data in 19 variables of known clinical and prognostic significance. We generated synthetic missing data for each variable, then performed imputation using substitution (control) and five different ML approaches. Imputation efficacy was measured by normalized root-mean-square error (RMSE) for continuous variables and proportion of falsely classified entries (PFC) for categorical variables. We also measured algorithm runtimes and the impact of incorporating imputed values on survival modeling. RESULTS50,790 NSCLC patients were included for this study, with 81 features for each patient after data preprocessing. Among the tested ML methods, SoftImpute had the lowest RMSE (best performance) for continuous variables ranging from 0.071 to 0.080 for 10% to 50% missing data, and MissForest had the lowest PFC (best performance) for categorical variables ranging from 0.251 to 0.311 for 10 to 50% missing data. SoftImpute had a runtime of 3.28x10-4 seconds per patient record, and MissForest averaged 2.96x10-3 seconds per patient record. Deep learning imputation using a denoising autoencoder did not achieve improved performance despite higher algorithm runtimes. Cox models incorporating ML imputed data achieved similar C-index ranging from 0.787 to 0.801 for all ML methods tested. CONCLUSIONML imputation achieved promising performance for NSCLC patients within a large national cancer registry.