Evaluating large language models for natural-language-to-code generation on aggregate Czech public health data analysis
Klempir, O.; Dusek, L.; Krupicka, R.; Donin, G.; Zigmond, J.; Storchova, R.; Tichopad, A.
Show abstract
Large language models (LLMs) are increasingly explored as tools for healthcare research and data analysis. However, their applicability to structured public health datasets, especially in non-English contexts, remains underexamined. We systematically evaluated 11 state-of-the-art LLMs on their ability to generate executable Python code for analytical queries over Czech public health datasets, focusing on incidence and prevalence data provided by the National Health Information Portal (known as NZIP). A set of representative analytical queries were designed, covering filtering, aggregation, weighted averages, and identification of primary diagnoses. Each model was prompted in Czech and assessed on code executability, correctness of results, and ability to adapt to local terminology. In the majority of cases, the models generated syntactically valid code within one minute, but performance varied. For the main objective of replicating "ground truth" queries as per dataset documentation, ChatGPT-4o achieved the highest accuracy, followed closely by GPT-4.1 mini. Claude and Gemini models frequently failed to apply critical filtering instructions, while Deepseek-R1, though accurate, defaulted to English output. Some models produced code that executed successfully but returned incorrect results, underscoring the need for systematic validation. Overall, LLMs show strong potential as coding assistants in public health analytics, even in Czech-language settings. Their integration into hybrid human-AI workflows, combined with validation mechanisms and retrieval-augmented generation, may accelerate the creation of reliable analytical pipelines.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 95%
- A Simple Electronic Medical Record System Designed for Research 95%
- Transforming Estonian health data to the Observational Medical Outcomes Partnership (OMOP) Common Data Model: lessons learned 94%
Similar papers in this journal
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 95%
- FHIR-DHP: A Standardized Clinical Data Harmonisation Pipeline for scalable AI application deployment 95%
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.