Related Experiment Video
Updated: Sep 28, 2026

An Affordable HIV-1 Drug Resistance Monitoring Method for Resource Limited Settings
Published on: March 30, 2014
Large language model-based triage to identify antiretroviral therapy adherence barriers and risk levels in patient
Yuanchao Ma1,2,3,4, Sofiane Achiche1, David Lessard2,3,4
1Institute of Biomedical Engineering, Polytechnique Montreal, Montreal, QC H3T 1J4, Canada.
Objective:
This study aimed to develop large language models (LLMs) to automatically identify antiretroviral therapy (ART) adherence barriers and stratify nonadherence risk levels from patient-generated messages.
Materials And Methods:
With a co-construction committee of people with HIV and providers, 15 480 sentences were annotated for barrier and risk levels. General-domain LLMs (eg, Flan-T5) and clinical foundation models (eg, Clinical-T5) were fine-tuned for multiclass classification and evaluated using Macro-F1. Best-performing models were compared with GPT family and other open-source LLMs. Model fairness, error patterns, and environmental footprints were also assessed.
Results:
Flan-T5-xl achieved the best barrier detection (Macro-F1 = 0.83 test/0.71 external), and Flan-T5-large excelled in risk stratification (0.79/0.57). Fine-tuned general-domain LLMs significantly outperformed clinical foundation models (P < .001) on the test split. Error analysis identified both annotation challenges (ambiguous inputs, cross-category overlap, and inconsistencies) and inference limitations (semantic overlap, implicit meanings, and contextual misinterpretation). On external datasets, Flan-T5 models outperformed other LLMs (P < .001), and were more robust to demographic descriptors, while GPT-4.1 and Gemma3 often misclassified inputs as "None." Additionally, GPT-4.1's inference consumed ∼30× more energy than that of Flan-T5-xl.
Discussion:
Fine-tuned Flan-T5 models demonstrated strong classification performance, greater robustness to demographic attributes, and lower energy consumption, though challenges remained for subjective and underrepresented categories, reflecting both data imbalance and model limitations in implicit reasoning.
Conclusion:
LLM-based approaches show promise for real-time ART adherence monitoring, offering a scalable solution to individualized HIV care. Beyond performance, our findings highlight the importance of fairness and environmental sustainability in clinical AI development, with next steps focused on real-world validation through deployment in patient-facing digital tools.
