OQA: A question-answering dataset on orthodontic literature
Rousseau, M.; Zouaq, A.; Huynh, N.
Show abstract
BackgroundThe near-exponential increase in the number of publications in orthodontics poses a challenge for efficient literature appraisal and evidence-based practice. Language models (LM) have the potential, through their question-answering fine-tuning, to assist clinicians and researchers in critical appraisal of scientific information and thus to improve decision-making. MethodsThis paper introduces OrthodonticQA (OQA), the first question-answering dataset in the field of dentistry which is made publicly available under a permissive license. A framework is proposed which includes utilization of PICO information and templates for question formulation, demonstrating their broader applicability across various specialties within dentistry and healthcare. A selection of transformer LMs were trained on OQA to set performance baselines. ResultsThe best model achieved a mean F1 score of 77.61 (SD 0.26) and a score of 100/114 (87.72%) on human evaluation. Furthermore, when exploring performance according to grouped subtopics within the field of orthodontics, it was found that for all LMs the performance can vary considerably across topics. ConclusionOur findings highlight the importance of subtopic evaluation and superior performance of paired domain specific model and tokenizer.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 93%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 92%
- Evaluating Knowledge Fusion Models on Detecting Adverse Drug Events in Text 92%
Similar papers in this journal
Similar papers in this journal
- A fast, accurate, and generalisable heuristic-based negation detection algorithm for clinical text 92%
- Refining LLMs Outputs with Iterative Consensus Ensemble (ICE) 92%
- Fusion of Electronic Health Records and Radiographic Images for a Multimodal Deep Learning Prediction Model of Atypical Femur Fractures 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.