Automating Handwritten Vaccination Record Transcription with Generative Multimodal AI Models: A Proof of Concept Study from The Gambia
Burstein, R.; Sowe, A.; Danovaro-Holliday, M. C.; Koh, M.; Proctor, J. L.
Show abstract
BackgroundHandwritten home-based vaccination records (HBRs) are a vital source of immunization data, yet manual transcription in household surveys is time-consuming, error-prone, and resource-intensive. Recent advances in multimodal generative AI models offer a potential pathway to partially automate this task under certain conditions, improving efficiency while maintaining data quality. MethodsWe evaluated the performance of generative multimodal AI models, specifically OpenAIs GPT-4o, in transcribing handwritten vaccination cards from a 2022 survey targeting children aged 12 - 35 months in The Gambia. Using a curated dataset of 335 cards (6,700 vaccination entries), we developed a gold-standard benchmark from three human transcribers (all of them in The Gambia) and assessed AI model performance across transcription accuracy, vaccination coverage estimates, timeliness, and missed opportunities for simultaneous vaccination (MOSV). We also tested a confidence-based segmentation approach to identify high-confidence transcriptions suitable for automation versus low-confidence entries requiring human review. ResultsThe fine-tuned GPT-4o model achieved 79% accuracy for exact date transcription and reached human-level performance (94% accuracy) on 69% of entries classified by the AI as high-confidence. Coverage and timeliness estimates from high-confidence transcriptions were 98.0% and 91.6% accurate, respectively, compared to 98.8% and 95.6% from human transcribers. Date errors by AI and humans differ systematically, with AI showing fewer year-shift errors. ConclusionMultimodal AI models show strong potential for automating HBR transcription in immunization coverage surveys, at least in a setting like The Gambia. When paired with confidence-based filtering, these models achieve human-level performance on coverage and timeliness estimates--the key metrics used in programmatic decision-making--across a large subset of records. This enables substantial gains in efficiency while preserving data quality. Further research should evaluate generalizability across diverse card formats, languages, and contexts to support integration into real-world immunization programs and health monitoring activities.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COVID-19 Vaccination Data Management and Visualization Systems for Improved Decision-Making: Lessons Learnt from Africa CDC Saving Lives and Livelihoods Program 93%
- Use of large language models as a scalable approach to understanding public health discourse 93%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 92%
Similar papers in this journal
Similar papers in this journal
- Recording vaccine doses administered: A global analysis of tally sheet design for infant and child immunizations 92%
- Prioritizing countries for TB vaccine readiness research using a global stakeholder-centric approach 92%
- Gaps and Opportunities for Data Systems and Economics to Support Priority Setting for Climate-Sensitive Infectious Diseases in Sub-Saharan Africa: A Rapid Scoping Review 91%
Similar papers in this journal
- Elucidating user behaviours in a digital health surveillance system to correct prevalence estimates 91%
- Nowcasting and Forecasting the 2022 U.S. Mpox Outbreak: Support for Public Health Decision Making and Lessons Learned 90%
- Assessing the utility of COVID-19 case reports as a leading indicator for hospitalization forecasting in the United States 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.