Related Experiment Video
Updated: Feb 24, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models as medical code selectors: a benchmark using the International Classification of Primary Care
Vinicius Anjos de Almeida1, Vinicius de Camargo2, Raquel Gómez-Bravo3
1Medical School, University of São Paulo, Av. Dr. Arnaldo, 455, São Paulo, São Paulo, 01246-903, Brazil.
Large language models (LLMs) show significant potential for automating medical coding, specifically assigning International Classification of Primary Care, 2nd edition (ICPC-2) codes. Top-performing LLMs achieved high accuracy, demonstrating a promising avenue for efficient healthcare data management.
Area of Science:
- Health Informatics
- Artificial Intelligence in Medicine
- Natural Language Processing
Background:
- Medical coding is crucial for healthcare data management, research, and policy.
- Automating the assignment of codes like the International Classification of Primary Care, 2nd edition (ICPC-2) can improve efficiency.
- Large language models (LLMs) offer new possibilities for complex natural language understanding tasks in healthcare.
Purpose of the Study:
- To assess the potential of various LLMs in automatically assigning ICPC-2 codes.
- To evaluate LLM performance using a domain-specific search engine's output.
- To benchmark LLM capabilities for medical coding tasks.
Main Methods:
- A dataset of 437 Brazilian Portuguese clinical expressions with ICPC-2 codes was utilized.
- A semantic search engine retrieved candidate codes from a large concept database.
- Thirty-three LLMs were prompted to select the best ICPC-2 code from retrieved results.
- Performance was measured by F1-score, token usage, cost, response time, and format adherence.
Main Results:
- Twenty-eight LLMs achieved an F1-score greater than 0.8, with 10 exceeding 0.85.
- GPT-4.5-preview, o3, and Gemini-2.5-pro were among the top performers.
- Retriever optimization enhanced performance by up to 4 F1-score points.
- Most models produced valid codes with reduced hallucinations and adhered to expected formats.
- Smaller models (<3B parameters) faced challenges with formatting and input length.
Conclusions:
- LLMs demonstrate strong potential for automating ICPC-2 coding without requiring fine-tuning.
- This study provides a benchmark and identifies challenges in LLM-based medical coding.
- Further research with broader, multilingual, and end-to-end evaluations is necessary for clinical validation.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:50A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Related Concept Videos
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Models of Health Promotion and Illness Prevention II
The agent-host-environment model states that disease results...
Models of Health Promotion and Illness Prevention I
The health belief model (HBM) attempts to predict health-related behavior in specific belief patterns. According to the HBM, a person's...
Secondary Healthcare System
Nursing Interventions II: Selecting and Classifying the Nursing Interventions