Back

Sino-US-DrugQA: A Benchmark for Evaluating Large Language Models in Cross-Jurisdictional Pharmaceutical Regulation

Chen, Z.; Fu, X.; Lu, W.

2026-02-17 health informatics
10.64898/2026.02.13.26346236 medRxiv
Show abstract

Cross-jurisdictional pharmaceutical compliance requires comparative analysis of regulatory requirements across jurisdictions such as the US FDA and Chinas NMPA. Although large language models (LLMs) are increasingly explored for healthcare-related applications, their performance in cross-jurisdictional regulatory comparison has not been systematically characterized using dedicated benchmarks. This study introduces Sino-US-DrugQA, a bilingual benchmark dataset designed to evaluate LLM performance in cross-jurisdictional pharmaceutical regulation. We constructed a bilingual corpus of 11,871 multiple-choice question-answer pairs derived from authoritative NMPA regulations and US CFR Title 21. The dataset includes monolingual retrieval tasks and cross-jurisdictional comparative reasoning tasks. We conducted baseline evaluations of four representative LLMs (GPT-5.2, Gemini-3-flash, Qwen-3-235B, DeepSeek-V3.2) under a standardized zero-shot protocol with temperature set to 0. The dataset and evaluation scripts are released as open resources to support reproducible benchmarking. Across 11,871 questions, the evaluated models achieved overall accuracies ranging from 78.97% to 84.51% in the zero-shot setting. Across models, performance consistently decreased on comparative questions relative to monolingual questions (approximately 6-9 percentage points), a gap that persisted even for the strongest-performing system, highlighting cross-jurisdictional alignment and comparison-specific deduction as key challenges. Sino-US-DrugQA provides a practical resource for benchmarking regulatory AI in a high-stakes compliance setting. Current LLMs show utility as drafting and screening assistants for monolingual regulatory queries, but their limitations in cross-jurisdictional comparative reasoning support a conservative deployment posture requiring expert review. The dataset and evaluation scripts are available at https://github.com/DodgeLU/Sino-US-DrugQA.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.