Related Experiment Video
Updated: May 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Building a Silver-Standard Dataset from NICE Guidelines for Clinical LLMs
Qing Ding1, Eric Hua Qing Zhang1, Felix Jozsa1
1University College London.
None:
Large language models (LLMs) are increasingly used in healthcare, yet standardised benchmarks for evaluating guideline-based clinical reasoning are missing. This study introduces a validated dataset derived from the UK's National Institute for Health and Care Excellence (NICE) guidelines covering multiple diagnoses. The dataset was initially created with the help of GPT-4o-mini and contains patient scenarios, as well as clinical questions; notably, generated contents were further assessed for clinical validity and realism by a practicing clinician to ensure alignment with NICE guidelines. We benchmark a range of recent popular LLMs to showcase the validity of our NICE guideline-based dataset. The framework supports systematic evaluation of LLMs' clinical utility and guideline adherence. The dataset and Appendix are publicly available at https://github.com/julia-ive/guidelines_qa.
Related Concept Videos
Clinical Trials: Overview
Statistical Software for Data Analysis and Clinical Trials
Clinical Trials
There are four phases in a clinical trial. A phase one...
Standards of Care II