Classifying Domains, Benchmarking GPT-4, A Portuguese Dataset for Medical AI Q&A
Matsuoka, F. A.; Onaga, H. N.
Show abstract
Artificial Intelligence (AI), particularly large language models (LLMs), has demonstrated remarkable capabilities in addressing complex tasks, including professional level medical question answering. While standardized benchmarks like the USMLE have been widely used for evaluating LLM performance in English, there is a significant gap in evaluating these models in other languages, such as Portuguese. To address this, we present a curated dataset derived from the Teste de Progresso (TP), a widely adopted Brazilian progress test used to assess medical knowledge across six key domains: Basic Sciences, Internal Medicine, Surgery, Obstetrics and Gynecology, Public Health, and Pediatrics. The dataset consists of 720 multiple-choice questions spanning five years (2019, 2023). We demonstrate two primary applications of this dataset. First, we benchmark the performance of GPT 4, which achieved an overall accuracy of 90% across the six medical domains, with the highest performance in Internal Medicine (10%) and the lowest in Public Health (80%). Second, we develop a classification model based on BERTimbau, achieving an overall accuracy of 94% in categorizing questions into their respective medical domains. Our results highlight the utility of the dataset for both benchmarking AI models and automating medical question classification. This work emphasizes the importance of creating domain-specific datasets in underrepresented languages, like Portuguese, to advance AI driven medical applications, ensure equitable access to AI technologies, and address linguistic and cultural gaps in healthcare education
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 92%
- The Mastery Rubric for Bioinformatics: supporting design and evaluation of career-spanning education and training 92%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 92%
Similar papers in this journal
Similar papers in this journal
- Optimizing biomedical information retrieval with a keyword frequency-driven Prompt Enhancement Strategy 94%
- Relation extraction between bacteria and biotopes from biomedical texts with attention mechanisms and domain-specific contextual representations 92%
- Transformer-based tool recommendation system in Galaxy 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.