Related Experiment Video
Updated: Oct 3, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Harm Reduction Advice From Large Language Models
Khy Choo1, Nathan Smith2,3, Bernadette Ward1,4
1School of Rural Health, Monash University, Mildura, Australia.
Introduction:
Large language models (LLM) and generative Artificial Intelligence are increasingly providing health information. We have applied a novel method to evaluate the harm reduction information generated by popular LLMs. We focused on the filtration of oral medications for injection and the availability of local harm reduction services.
Methods:
We designed a scoring rubric to evaluate the accuracy of the information provided by each LLM and applied it to four unique, commonly used LLMs.
Results:
Gemini 2.5 Flash (87%) and ChatGPT-4o (77%) scored higher than DeepSeek R1 Distill Llama 3.3 70B (22%) and Llama 4 (18%). Overall, both Gemini and ChatGPT-4o gave comprehensive and accurate instructions for the use of micron filters for drug injection, while neither DeepSeek R1 nor Llama 4 provided instructions on their use. Gemini 2.5 and ChatGPT-4o both scored highly for harm reduction general information and situational questions, while DeepSeek R1 and Llama 4 provided limited to no responses. Gemini 2.5 and ChatGPT-4o both provided accurate, location-specific information on local harm reduction services. Specific models also generated misinformation, further affirming the need for in-person, nuanced discussions with harm reduction specialists.
Discussion And Conclusions:
In this study, we presented a novel method for evaluating LLMs' potential effectiveness for harm reduction across common techniques and situations. Most of the tested LLMs provided reasonable information in a non-judgmental manner, but they also provided misinformation, and some wouldn't answer questions due to built-in safety parameters. The developed scoring rubric can be easily adapted to other harm reduction topics for ongoing evaluation.