Related Experiment Videos
CDTD-Agent: a retrieval-augmented multi-agent digital health framework for integrated diabetes care: a single-case
Jieyi Zhang1, Luqing Wu2, Jielong Wu1
1Center of Integrated Chinese and Western Medicine, The First Affiliated Hospital of Xiamen University, Xiamen, Fujian, China.
Background:
Diabetes management increasingly requires coordination across Western medicine, traditional Chinese medicine (TCM), and lifestyle-based health management. General-purpose large language models (LLMs) can generate clinical text but may lack reliable source grounding and structured cross-disciplinary reasoning.
Methods:
We developed CDTD-Agent, a retrieval-augmented multi-agent framework that combines a domain-specific knowledge base with an Extract-Retrieve-Reason workflow. The knowledge base included 12 medical textbooks, 8 clinical guidelines, and 183 expert-reviewed diabetes cases. In a blinded single-case pilot evaluation, 30 senior experts rated anonymized recommendations from CDTD-Agent, GPT-4o, DeepSeek-V3.2-Exp, Doubao, and two attending physicians across eight dimensions. Linear mixed-effects models accounted for repeated ratings within evaluators, with Holm correction across eight omnibus tests and global Benjamini-Hochberg correction across 40 focused CDTD-Agent-versus-comparator contrasts. A post hoc backbone-only ablation compared the full framework with Qwen3-235B-A22B using the same case and external prompt without retrieval or role-structured multi-agent processing.
Results:
Significant source effects were detected in all eight dimensions after Holm correction (all adjusted p < 0.001). After global false discovery rate correction, CDTD-Agent received significantly higher ratings than all three general-purpose LLMs in five dimensions: safety, examination appropriateness, Western medicine treatment appropriateness, TCM treatment appropriateness, and evidence adherence. The largest differences occurred in Western medicine treatment appropriateness, where CDTD-Agent scored 4.07 ± 0.94 versus 1.37-1.60 for the general-purpose LLMs (all adjusted p < 0.001). Non-parametric sensitivity analyses agreed with the mixed-effects results for 39 of 40 focused contrasts. In the post hoc ablation, the full framework outperformed the backbone-only condition in seven of eight dimensions after Holm correction; health management was not significantly different (adjusted p = 0.080).
Conclusion:
In this standardized complex T2DM case, CDTD-Agent received favorable expert ratings relative to the evaluated general-purpose LLMs in several clinically important dimensions. These findings are scenario-specific and do not establish generalizable clinical performance or isolate the effects of individual framework components. Multicase, repeated-generation, factorial-ablation, external, and prospective validation is needed before clinical deployment.