Back

A ReAct Agentic AI System for Natural Language Querying and Statistical Analysis of The Cancer Genome Atlas Clinical Data

Korutla, R.; Amal, S.

2026-07-17 health informatics
10.64898/2026.07.15.26358188 medRxiv
Show abstract

The Cancer Genome Atlas (TCGA) holds clinical data for over 11,000 patients across 33 cancer types, but access is hard because of complex file structures, heterogeneous formats, and the need for programming. We present an agentic system for natural language querying and statistical analysis of TCGA clinical data. The system uses a large language model as an autonomous ReAct agent that selects from eight computational tools, including data extraction, descriptive statistics, Kaplan-Meier survival analysis with log-rank tests, hypothesis testing, and verification against the curated TCGA Pan-Cancer Clinical Data Resource (CDR). The agent reasons about intermediate results, adapts its approach, and returns clinically contextualized responses with source attribution and auditable traces. We introduce TCGA-Agent-Bench, 440 queries across five difficulty tiers with ground truth from the independently curated TCGA-CDR, evaluated with dual metrics of numerical accuracy and clinical completeness. The system achieves 93.4% overall accuracy (100% single-patient lookups, 99.1% cohort statistics, 92.8% comparative analyses), outperforming a fixed rule-based pipeline (87.1%), a single-pass LLM (81.8%), and retrieval-augmented generation (66.9% on a subset). Most of the benchmark is answerable from the CDR alone, so we locate the extraction layer's value in fields the CDR lacks (drug treatments, TNM components, biomarkers, biospecimen metadata): on 26 queries targeting these, the full system answers 100% versus 3.8% for CDR-only. Ablations show the reasoning loop is most impactful (+9.1% accuracy, +22.0 completeness points). A tool-based agentic architecture enables accurate, auditable analysis of clinical repositories, with value driven by tool design and recovered fields rather than model scale.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 3%
9.7%
2
Nature Communications
5641 papers in training set
Top 21%
7.8%
3
Nature Medicine
125 papers in training set
Top 0.1%
7.8%
4
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.1%
5.5%
5
Patterns
78 papers in training set
Top 0.2%
5.4%
6
npj Digital Medicine
118 papers in training set
Top 1%
5.4%
7
iScience
1154 papers in training set
Top 2%
5.4%
8
Scientific Reports
3612 papers in training set
Top 21%
4.8%
50% of probability mass above
9
BMC Medical Informatics and Decision Making
43 papers in training set
Top 0.5%
4.3%
10
Communications Medicine
113 papers in training set
Top 1.0%
3.2%
11
Med
39 papers in training set
Top 0.1%
3.2%
12
Journal of the American Medical Informatics Association
71 papers in training set
Top 1.0%
3.2%
13
Scientific Data
209 papers in training set
Top 0.9%
3.1%
14
PLOS ONE
5266 papers in training set
Top 43%
2.4%
15
JMIR Medical Informatics
18 papers in training set
Top 0.4%
2.1%
16
GigaScience
212 papers in training set
Top 3%
1.7%
17
Journal of Medical Internet Research
87 papers in training set
Top 2%
1.7%
18
Genome Biology
637 papers in training set
Top 6%
1.5%
19
Nature Machine Intelligence
70 papers in training set
Top 2%
1.1%
20
JAMIA Open
42 papers in training set
Top 1%
1.1%
21
Science Advances
1243 papers in training set
Top 28%
1.0%
22
npj Antimicrobials and Resistance
11 papers in training set
Top 0.2%
0.8%
23
Nature Methods
385 papers in training set
Top 6%
0.8%
24
Nature Computational Science
55 papers in training set
Top 2%
0.8%
25
Journal of Biomedical Informatics
47 papers in training set
Top 1%
0.6%
26
Frontiers in Digital Health
24 papers in training set
Top 2%
0.6%
27
BMJ Health & Care Informatics
15 papers in training set
Top 1%
0.6%
28
Advanced Science
286 papers in training set
Top 11%
0.6%
29
IEEE Journal of Biomedical and Health Informatics
37 papers in training set
Top 1%
0.6%