Sino-US-DrugQA: A Benchmark for Evaluating Large Language Models in Cross-Jurisdictional Pharmaceutical Regulation
Chen, Z.; Fu, X.; Lu, W.
Show abstract
Cross-jurisdictional pharmaceutical compliance requires comparative analysis of regulatory requirements across jurisdictions such as the US FDA and Chinas NMPA. Although large language models (LLMs) are increasingly explored for healthcare-related applications, their performance in cross-jurisdictional regulatory comparison has not been systematically characterized using dedicated benchmarks. This study introduces Sino-US-DrugQA, a bilingual benchmark dataset designed to evaluate LLM performance in cross-jurisdictional pharmaceutical regulation. We constructed a bilingual corpus of 11,871 multiple-choice question-answer pairs derived from authoritative NMPA regulations and US CFR Title 21. The dataset includes monolingual retrieval tasks and cross-jurisdictional comparative reasoning tasks. We conducted baseline evaluations of four representative LLMs (GPT-5.2, Gemini-3-flash, Qwen-3-235B, DeepSeek-V3.2) under a standardized zero-shot protocol with temperature set to 0. The dataset and evaluation scripts are released as open resources to support reproducible benchmarking. Across 11,871 questions, the evaluated models achieved overall accuracies ranging from 78.97% to 84.51% in the zero-shot setting. Across models, performance consistently decreased on comparative questions relative to monolingual questions (approximately 6-9 percentage points), a gap that persisted even for the strongest-performing system, highlighting cross-jurisdictional alignment and comparison-specific deduction as key challenges. Sino-US-DrugQA provides a practical resource for benchmarking regulatory AI in a high-stakes compliance setting. Current LLMs show utility as drafting and screening assistants for monolingual regulatory queries, but their limitations in cross-jurisdictional comparative reasoning support a conservative deployment posture requiring expert review. The dataset and evaluation scripts are available at https://github.com/DodgeLU/Sino-US-DrugQA.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 94%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 94%
- Comparative Effectiveness of Medical Concept Embedding for Feature Engineering in Phenotyping 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.