Related Experiment Video
Updated: Jun 13, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Note-Level Phenotyping of Multiple-Sclerosis Notes by a Large Language Model Achieves near Human-Level Agreement
Daniel B Hier1, Pavankumar Y Srinivasula2, Michael D Carrithers1
1College of Medicine, University of Illinois at Chicago, Chicago, IL 60612, USA.
None:
Background/Objectives: Clinical phenotyping from narrative electronic health records (EHRs) often relies on multi-stage pipelines involving span-level extraction, ontology mapping, and aggregation. Large language models (LLMs) may enable direct document-level abstraction of clinically meaningful phenotype features from complete notes. We evaluated whether GPT-5.2 could approximate human annotation for note-level multiple sclerosis (MS) phenotyping and compared its performance with human annotators, a locally run open-source LLM, HPO-based extraction tools, and a supervised clinical transformer encoder. Methods: We analyzed 100 de-identified MS neurology progress notes from a single academic medical center. Each note was annotated for the presence or absence of 17 predefined neurological phenotype categories. Two human annotators independently labeled all notes using a multi-label note-level framework in Prodigy, and disagreements were adjudicated to create a reference annotation set. GPT-5.2 was evaluated in a zero-shot setting using structured JSON output. Comparator methods included Llama-3.1 8B, Doc2Hpo, ClinPhen, PhenoSnap, and BioClinical ModernBERT. Performance was assessed using agreement, precision, recall, F1, Matthews correlation coefficient, and false-positive and false-negative assignments per note. Results: Human-human agreement was generally high, although lower for rare or ambiguously documented features. GPT-5.2 achieved the strongest automated performance, with macro-precision 0.734, macro-recall 0.921, macro-F1 0.801, and macro-averaged MCC 0.777, approaching human annotator performance. GPT-5.2 showed the lowest false-negative count per note but more false-positive assignments than either human annotator, reflecting a sensitive but more inclusive annotation profile. Llama-3.1 8B performed competitively among automated methods, whereas HPO-based extraction tools and BioClinical ModernBERT showed lower performance on this low-resource note-level task. Secondary review of GPT-5.2 discordant assignments found no clear hallucinations and suggested that some apparent false positives reflected phenotype evidence missed in the human-derived reference set. Conclusions: GPT-5.2 achieved near-human performance for document-level recognition of MS phenotype categories from narrative neurology notes. Direct note-level abstraction may provide a scalable approach for research and population-health phenotyping of large EHR note corpora.
Related Concept Videos
Multiple Sclerosis l: Introduction
Higher Mental Functions of the Brain: Language
Language formation and comprehension take place in the dominant hemisphere. The dominant hemisphere is responsible for understanding the meaning of spoken, written, or sign language, as well as the ability to communicate. For most people, the left hemisphere is the dominant one. The right hemisphere, then, gives tone and emotional context to the...
Improving Translational Accuracy
Improving Translational Accuracy