Related Experiment Video
Updated: Sep 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Sino-US-DrugQA: A benchmark for evaluating large language models in cross-jurisdictional pharmaceutical regulation
Xuejing Fu1, Zhen Chen1, Wentao Lu2
1Information Application Research Center of Shanghai Municipal Administration for Market Regulation, Shanghai, China.
Abstract:
Cross-jurisdictional pharmaceutical compliance requires comparison of regulatory requirements across administrative systems such as the US Food and Drug Administration and China's National Medical Products Administration. Although large language models (LLMs) are increasingly explored for healthcare and regulatory applications, their performance in cross-jurisdictional pharmaceutical regulation has not been systematically evaluated using a dedicated benchmark. We introduce Sino-US-DrugQA, a bilingual multiple-choice benchmark covering Monolingual, Comparative, and Parallel regulatory question-answer tasks. The 11,871 candidate items underwent deterministic structural validation and full-dataset semantic quality screening, followed by risk-stratified independent review of 1,432 items by two regulatory experts. The final release comprised 11,444 items, including 10,122 classified as pass and 1,322 as borderline. Among 500 items sampled from the semantic screen-negative population, 18 were subsequently classified as material errors, corresponding to an observed residual material-error proportion of 3.60% (Wilson 95% confidence interval, 2.29%-5.62%). Four representative LLMs-gpt-5.6-terra, gemini-3.6-flash, deepseek-v4-flash, and qwen-3.5-max-were evaluated under a standardized zero-shot protocol. Overall accuracy ranged from 83.21% to 86.43%. Comparative accuracy was consistently lower than Monolingual accuracy, with absolute differences of 4.42-8.98 percentage points across models. The two highest-scoring models, GPT and Gemini, did not differ significantly after adjustment for multiple comparisons. These results indicate that explicit comparison across non-equivalent regulatory systems remains more challenging than single-jurisdiction question answering. Sino-US-DrugQA provides a validated and reproducible resource for evaluating bilingual regulatory reasoning. The findings support further investigation of expert-supervised decision-support workflows rather than autonomous regulatory interpretation. The stable dataset release and evaluation resources are available at https://github.com/DodgeLU/Sino-US-DrugQA.
Related Concept Videos
Impact of Pharmacokinetic–Pharmacodynamic Models: Regulatory Decisions
Drug Regulation
Drug Products: Biologics, Biosimilars and Interchangeables
Drug Control Governance: Regulatory Bodies and Their Impact
Drug Nomenclature
Clinically Relevant Drug Product Specifications: Methods of Establishment
