Related Experiment Videos

Comparative Performance of AI Models and Clinicians in Evidence-Based Cardiovascular Disease Management for People

Tianqi Kong1, Liqin Sun2, Yinsong Luo1

  • 1School of Public Health, Shenzhen University Medical School, Shenzhen University, No.1066 Xueyuan Avenue, Shenzhen, Guangdong, 518060, China, 86 18819026906.

Insights

Large language models (LLMs) significantly outperformed human clinicians in managing cardiovascular disease (CVD) in people with HIV. DeepSeek-R1 demonstrated superior performance, highlighting AI

Area of Science:

  • Artificial Intelligence in Medicine
  • Cardiovascular Disease Management
  • HIV/AIDS Clinical Practice

Background:

  • Antiretroviral therapy has increased life expectancy for people with HIV, but cardiovascular disease (CVD) is now a major comorbidity.
  • Gaps in cross-specialty knowledge hinder guideline adherence for HIV-associated CVD management.
  • Integrated, evidence-based tools are needed to address interdisciplinary barriers in managing HIV and CVD.

Purpose of the Study:

  • To compare the performance of four large language models (LLMs) against human clinicians in managing guideline-based CVD in people with HIV.
  • To evaluate the effectiveness of AI in addressing cross-specialty knowledge gaps in HIV-associated CVD care.

Main Methods:

  • Developed a 25-question assessment based on HIV/CVD guidelines via Delphi consultation.
  • Four LLMs (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, ChatGPT-o4-mini) and 12 human clinicians (infectious disease specialists and cardiologists) responded to the assessment.
  • Six multidisciplinary experts rated responses on accuracy, completeness, readability, and reliability using a 4-point scale.

Main Results:

  • All AI models significantly outperformed clinicians across all evaluation dimensions (P<.001).
  • DeepSeek-R1 achieved the highest performance, significantly outperforming other AI models (P<.001).
  • Cardiologists showed higher accuracy in CVD risk assessment, while infectious disease specialists excelled in drug adverse effect evaluation.

Conclusions:

  • LLMs, particularly DeepSeek-R1, demonstrate superior performance in HIV-associated CVD management compared to human clinicians.
  • AI tools show potential for integrating complex clinical knowledge and mitigating specialty gaps in multidisciplinary care.
  • Integrating AI into clinical workflows may optimize the management of complex comorbidities in people living with HIV.
Abstract