Related Experiment Videos
Comparative Performance of AI Models and Clinicians in Evidence-Based Cardiovascular Disease Management for People
Tianqi Kong1, Liqin Sun2, Yinsong Luo1
1School of Public Health, Shenzhen University Medical School, Shenzhen University, No.1066 Xueyuan Avenue, Shenzhen, Guangdong, 518060, China, 86 18819026906.
Insights
Large language models (LLMs) significantly outperformed human clinicians in managing cardiovascular disease (CVD) in people with HIV. DeepSeek-R1 demonstrated superior performance, highlighting AI
Area of Science:
- Artificial Intelligence in Medicine
- Cardiovascular Disease Management
- HIV/AIDS Clinical Practice
Background:
- Antiretroviral therapy has increased life expectancy for people with HIV, but cardiovascular disease (CVD) is now a major comorbidity.
- Gaps in cross-specialty knowledge hinder guideline adherence for HIV-associated CVD management.
- Integrated, evidence-based tools are needed to address interdisciplinary barriers in managing HIV and CVD.
Purpose of the Study:
- To compare the performance of four large language models (LLMs) against human clinicians in managing guideline-based CVD in people with HIV.
- To evaluate the effectiveness of AI in addressing cross-specialty knowledge gaps in HIV-associated CVD care.
Main Methods:
- Developed a 25-question assessment based on HIV/CVD guidelines via Delphi consultation.
- Four LLMs (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (infectious disease specialists and cardiologists) responded to the assessment.
- Six multidisciplinary experts rated responses on accuracy, completeness, readability, and reliability using a 4-point scale.
Main Results:
- All AI models significantly outperformed clinicians across all evaluation dimensions (P<.001).
- DeepSeek-R1 achieved the highest performance, significantly outperforming other AI models (P<.001).
- Cardiologists showed higher accuracy in CVD risk assessment, while infectious disease specialists excelled in drug adverse effect evaluation.
Conclusions:
- LLMs, particularly DeepSeek-R1, demonstrate superior performance in HIV-associated CVD management compared to human clinicians.
- AI tools show potential for integrating complex clinical knowledge and mitigating specialty gaps in multidisciplinary care.
- Integrating AI into clinical workflows may optimize the management of complex comorbidities in people living with HIV.
Background:
Although widespread antiretroviral therapy has extended the life expectancy of people living with HIV, cardiovascular disease (CVD) has emerged as a primary comorbidity. Persistent cross-specialty knowledge gaps in routine clinical practice lead to suboptimal adherence to guidelines. Integrated, evidence-based tools are urgently needed to overcome these interdisciplinary barriers. While large language models (LLMs) have demonstrated significant capabilities in medicine, no systematic evaluation has assessed their ability to facilitate multidisciplinary CVD management for people living with HIV.
Objective:
This study compared the performance of 4 mainstream AI models (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, and ChatGPT-o4-mini) against 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for people living with HIV.
Methods:
Based on 4 authoritative domestic and international HIV/CVD guidelines, a structured 25-question assessment was developed via 2 rounds of Delphi consultation. Standard reference answers and an evaluation framework were finalized through expert consensus. LLM responses were generated using standardized prompts. Clinicians answered identical questions via one-on-one structured interviews, transcribed verbatim. Six multidisciplinary experts independently rated all responses across 4 dimensions-accuracy, completeness, readability, and reliability-using a 4-point ordinal scale (1=poor to 4=excellent). Cumulative link mixed models analyzed intergroup differences.
Results:
All AI models achieved significantly higher scores than clinicians across all dimensions (P<.001). The AI group's mean scores ranged from 3.44 to 3.68 (median 4, IQR 3.0-4.0; coefficient of variation=0.145-0.178). Conversely, clinicians' scores were lower (mean 1.78-2.05; median 2, IQR 1.0-3.0; coefficient of variation=0.428-0.473) with marked dispersion. DeepSeek-R1 delivered the optimal performance, significantly outperforming the other 3 models (all P<.001). Specialty-stratified analysis revealed no significant overall score difference between cardiologists and infectious disease specialists (odds ratio 0.92, 95% CI 0.84-1.01; P=.09). However, dimension-specific analysis indicated that cardiologists scored higher in accuracy (odds ratio 0.81, 95% CI 0.67-0.97; P=.03). Domain-specific divergence was evident: cardiologists outperformed infectious disease specialists in CVD risk assessment (2.26 vs 1.83), whereas infectious disease specialists led in drug adverse effect evaluation (2.23 vs 1.65).
Conclusions:
In this structured question-and-answer study, LLMs outperformed human clinicians across all metrics for HIV-associated CVD management, with DeepSeek-R1 achieving superior composite scores. These findings validate DeepSeek-R1's potential as a cross-disciplinary decision-support tool capable of integrating complex clinical knowledge, mitigating specialty gaps, and enhancing information precision. Integrating AI systems into multidisciplinary workflows, complemented by targeted clinical training, may optimize the management of complex comorbidities in people living with HIV.