Back

Automating the cancer registry: An Autonomous, Resource-Efficient AI for Multi-Cancer Data Abstraction from Pathology Reports

Chou, N.-H.; Chang, H.; Chen, H.-K.; Lin, C.-Y. T.; Liu, Y.-L.; Tseng, P.-Y.; Hsu, L.-C.; Chu, Y.-W.; Chang, K.-P.

2025-10-23 health informatics
10.1101/2025.10.21.25338475 medRxiv
Show abstract

Pathology reports contain the most detailed descriptions of cancer diagnoses, yet their unstructured format has long limited large-scale reuse for cancer registries and population surveillance. Prior applications of large language models (LLMs) have therefore focused on narrow extraction tasks, reflecting a persistent implementation trilemma: comprehensive abstraction, strict data privacy, and computational feasibility could not be achieved simultaneously in real-world clinical settings. Given current LLM capabilities, this trilemma can now be resolved. We show that recent open-weight LLMs enable reliable, full-length, schema-bound abstraction of pathology reports on standard on-premise hardware. We present a model-agnostic framework implemented using DSPy, a declarative framework for structured LLM pipelines, in which deterministic, programmatic prompting co-designed with pathologists enables end-to-end structured abstraction. Across 893 real-world pathology reports spanning ten major cancer types, the system achieved a mean exact-match accuracy of 94.3% across 193 CAP-aligned registry fields, including complex variable-length structures such as surgical margins, lymph nodes, and breast biomarkers. All processing was performed locally on a single workstation-class GPU, ensuring data privacy without sacrificing completeness or feasibility. Independent external validation using TCGA pathology reports confirmed robust generalizability.

Published in Diagnostics (predicted rank #24) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.