Back

DentaCoPilot: An LLM-Augmented Next-Procedure Recommender for General Dentistry, Designed for Dentist Augmentation

Rodrigues, C. C.; Rebello, S. D.

2026-05-08 dentistry and oral medicine
10.64898/2026.05.07.26352635 medRxiv
Show abstract

BackgroundCommercial dental artificial intelligence in 2026 is over-whelmingly diagnostic: caries, calculus, periapical, and bone-level detection on radiographs. The clinically harder question that follows every diagno-sis -- given a patients chart and most recent procedure, what should the dentist do next -- remains unsolved at general-dentistry scale. The closest published system, MultiTP (Chen et al., 2024), is a CNN-RNN restricted to partial-edentulism cases and provides neither calibrated uncertainty, structured rationale, nor an evaluation that treats the model as decision support rather than as an autonomous classifier. MethodsWe introduce DentaCoPilot, a recommender that, given a structured chart, returns (i) a calibrated top-K probability distribution over Current Dental Terminology (CDT) codes for the next procedure, (ii) a verbalised confidence label, (iii) an explicit abstain flag when context is insufficient, and (iv) a chartgrounded rationale. We compare four classical baselines (frequency bigram, TF-IDF + logistic regression, XGBoost, MultiTP-style CNN-RNN) and six large-language-model (LLM) variants (Claude Haiku, Sonnet + chain-of-thought, Sonnet + retrieval, Opus + chain-of-thought, Sonnet + classical prior, Opus + classical prior) on a synthetic chart corpus of 500 patients (1,284 test examples). All LLM inference is routed through the local Anthropic Claude Code CLI; every call is logged for full audit. ResultsOn apples-to-apples evaluation, classical baselines reach 0.567 top-1 / 0.967 top-5; pure LLM variants trail at 0.267-0.467 top-1. Prompt-conditioning a Sonnet LLM on the classical baselines top-10 candidates (M5) closes the gap: top-5 rises from 0.733 (pure Sonnet + chain-of-thought) to 0.933, matching classical baselines, while preserving rationale and abstention. Increasing the LLM backbone from Sonnet to Opus does not improve accuracy with or without priming. Calibration via temperature scaling and coverage-risk analysis is reported for the baselines. ConclusionPrompt-conditioning a small LLM on a classical baselines top-K is the most cost-effective LLM design we tested for next-procedure recommendation, and the design preserves the augmentation features that distinguish the system from an autonomous classifier. A pre-registered clinician-in-the-loop evaluation at the KLE Vish-wanath Katti Institute of Dental Sciences (Belgaum, India) and a real-data evaluation on the multi-institutional BigMouth dental data repository are the next stage of work.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
PLOS Digital Health
106 papers in training set
Top 0.1%
23.8%
2
Journal of Dental Research
13 papers in training set
Top 0.1%
4.6%
3
European Radiology
15 papers in training set
Top 0.2%
4.3%
4
Scientific Reports
3612 papers in training set
Top 23%
4.3%
5
PLOS ONE
5266 papers in training set
Top 35%
3.7%
6
Journal of the American Medical Informatics Association
71 papers in training set
Top 0.9%
3.5%
7
IEEE Access
35 papers in training set
Top 0.4%
3.0%
8
Artificial Intelligence in Medicine
17 papers in training set
Top 0.2%
2.6%
9
JAMA Network Open
130 papers in training set
Top 1%
2.5%
50% of probability mass above
10
BioMed Research International
28 papers in training set
Top 0.5%
2.5%
11
Frontiers in Medicine
120 papers in training set
Top 1%
2.5%
12
Journal of Medical Internet Research
87 papers in training set
Top 1%
2.3%
13
Children
10 papers in training set
Top 0.3%
2.0%
14
Frontiers in Public Health
148 papers in training set
Top 3%
1.8%
15
Biology Methods and Protocols
61 papers in training set
Top 0.7%
1.8%
16
Biomolecules
100 papers in training set
Top 1%
1.6%
17
npj Digital Medicine
118 papers in training set
Top 2%
1.6%
18
JAMIA Open
42 papers in training set
Top 1%
1.4%
19
Journal of Biomedical Informatics
47 papers in training set
Top 0.9%
1.4%
20
Cureus
68 papers in training set
Top 3%
1.4%
21
Translational Vision Science & Technology
39 papers in training set
Top 0.4%
1.2%
22
Medicine
31 papers in training set
Top 1%
1.2%
23
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.4%
1.2%
24
Computers in Biology and Medicine
128 papers in training set
Top 3%
1.2%
25
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
26
Journal of Neural Engineering
221 papers in training set
Top 2%
0.9%
27
eLife
5828 papers in training set
Top 63%
0.9%
28
Bioinformatics
1204 papers in training set
Top 8%
0.9%
29
BMC Medical Informatics and Decision Making
43 papers in training set
Top 2%
0.7%
30
PLOS Global Public Health
344 papers in training set
Top 8%
0.7%