Back

Thematic Shifts in Early-High-Impact Cancer Genomics and Diagnostics Research: A Bibliometric and Semantic Analysis

Su, Z.; Li, T.

2026-07-09 bioinformatics
10.64898/2026.07.04.736459 bioRxiv
Show abstract

Cancer genomics and diagnostics is a rapidly evolving field in which identifying which topics attract early citation prominence can inform laboratory investment, clinical translation, and research strategy. We developed a bibliometric framework to identify and characterize the most influential recent publications in this domain across two consecutive annual cohorts. Using a mathematically exact threshold-expansion algorithm, we ranked over 10,000 OpenAlex-indexed research articles per cohort by 18-month post-publication citation count. Large language model (LLM)-based topical relevance filtering yielded 50 substantively on-topic papers per cohort (100 total). LLM-based concept extraction and a two-stage, embedding-guided normalization pipeline produced 1,853 canonical concepts organized into 103 parent themes, enabling structured cross-cohort comparison of paper-level concept prevalence. The most cited papers in both cohorts were large-scale genomic infrastructure resources rather than single-disease mechanistic studies. Between consecutive cohorts, normalized frequencies increased most for whole-genome sequencing, tumor microenvironment biology, molecular biomarkers, and cancer pharmacotherapy, while liquid biopsy-related themes showed the largest declines. These findings indicate that early citation impact in cancer genomics is shifting toward integrative, population-scale, and microenvironment-aware research, and demonstrate that LLM-augmented citation ranking provides a replicable, semantically enriched lens for monitoring thematic evolution in precision oncology. A web interface for exploring the results is available at https://pri.pepkio.com/.

Matching journals

The top 12 journals account for 50% of the predicted probability mass.

1
npj Precision Oncology
53 papers in training set
Top 0.1%
6.8%
2
PLOS ONE
5266 papers in training set
Top 24%
6.8%
3
Communications Medicine
113 papers in training set
Top 0.4%
5.6%
4
GigaScience
212 papers in training set
Top 0.5%
5.5%
5
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.3%
6
Genome Medicine
183 papers in training set
Top 1%
3.5%
7
Scientific Reports
3612 papers in training set
Top 30%
3.4%
8
Journal of Translational Medicine
57 papers in training set
Top 0.2%
3.3%
9
Nature Communications
5641 papers in training set
Top 35%
3.3%
10
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
3.2%
11
BMC Medical Genomics
50 papers in training set
Top 0.2%
2.8%
12
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.8%
50% of probability mass above
13
PLOS Computational Biology
1863 papers in training set
Top 12%
2.5%
14
BMC Bioinformatics
457 papers in training set
Top 3%
2.4%
15
Nucleic Acids Research
1281 papers in training set
Top 7%
2.4%
16
Database
61 papers in training set
Top 0.4%
2.1%
17
Bioinformatics
1204 papers in training set
Top 6%
2.0%
18
Cell Systems
201 papers in training set
Top 3%
1.7%
19
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.4%
1.7%
20
Nature Cancer
39 papers in training set
Top 0.9%
1.5%
21
npj Systems Biology and Applications
125 papers in training set
Top 1%
1.3%
22
BioData Mining
22 papers in training set
Top 0.5%
1.1%
23
Genome Biology
637 papers in training set
Top 7%
1.1%
24
Scientific Data
209 papers in training set
Top 2%
1.1%
25
npj Digital Medicine
118 papers in training set
Top 3%
1.1%
26
Cancer Research
130 papers in training set
Top 2%
1.1%
27
The Journal of Molecular Diagnostics
39 papers in training set
Top 0.5%
1.0%
28
Cancer Cell
42 papers in training set
Top 1%
0.9%
29
Frontiers in Oncology
103 papers in training set
Top 3%
0.8%
30
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 41%
0.8%