Related Experiment Videos
Enabling doctor-centric medical AI with LLMs through workflow-aligned tasks and benchmarks
Wenya Xie1,2, Qingying Xiao3, Yu Zheng1
1School of Data Science, The Chinese University of Hong Kong, Shenzhen, Shenzhen, Guangdong, China.
Npj Health Systems
|July 29, 2026
Summary
Large language models (LLMs) can aid doctors, not patients directly. A new dataset, DoctorFLAN, enhances LLM performance for clinical assistants, improving physician workflows.
Area of Science:
- Artificial Intelligence in Medicine
- Natural Language Processing for Healthcare
Background:
- Large language models (LLMs) offer clinical guidance but pose risks if deployed directly to patients due to limited medical expertise.
- Repositioning LLMs as clinical assistants for physician collaboration is proposed to mitigate these safety concerns.
Purpose of the Study:
- To develop a doctor-centered framework for medical LLM applications.
- To create a large-scale Chinese medical dataset and evaluation benchmarks tailored for clinical assistant roles.
Main Methods:
- A two-stage inspiration-feedback survey identified clinical workflow needs.
- Construction of DoctorFLAN, a 92,000 Q&A instance dataset across 22 clinical tasks and 27 specialities.
- Introduction of DoctorFLAN-test and DotaBench for evaluating LLM performance in doctor-facing applications.
Main Results:
- DoctorFLAN significantly enhances the performance of open-source LLMs in medical contexts.
- The dataset facilitates better alignment of LLMs with physician workflows.
- LLMs trained on DoctorFLAN complement existing patient-oriented medical models.
Conclusions:
- DoctorFLAN provides a valuable resource for developing doctor-centered medical LLMs.
- The proposed framework supports the safe and effective integration of LLMs into clinical practice as physician assistants.