Related Experiment Video
Updated: Aug 29, 2026

A Metadata Extraction Approach for Clinical Case Reports to Enable Advanced Understanding of Biomedical Concepts
Published on: September 20, 2018
Artificial intelligence for automated ICD-10 coding: a systematic review of multi-label text classification in
Kamonrat Tangudomkit1, Sawrawit Chairat1, Sitthichok Chaichulee1
1Department of Biomedical Sciences and Biomedical Engineering, Faculty of Medicine, Prince of Songkla University, Songkhla, Thailand.
Background:
ICD-10 coding is an essential process in healthcare systems that supports clinical management, reimbursement, and health data analytics. However, the complexity of its hierarchical structure and the large number of available codes make manual coding limited in terms of time, cost, and consistency. Despite growing research in this area, evidence remains fragmented, particularly regarding real-world implementation readiness.
Objective:
To review and synthesize existing knowledge on algorithms, datasets, evaluation methods, and real-world implementation readiness of automatic ICD-10 coding systems.
Methods:
Eligible studies were original research articles, preprints, or conference papers published in English between January 1, 2020 and December 31, 2025, and retrieved from seven academic databases: Scopus, PubMed, Web of Science, IEEE Xplore, ACM Digital Library, arXiv, and Google Scholar. Studies were included if they investigated automatic ICD-10 coding from clinical text using machine learning, deep learning, transformer-based, or large language model (LLM) approaches. Methodological quality was assessed using a research-question-driven appraisal framework. This systematic review followed PRISMA 2020 guidance and was preregistered in the Open Science Framework (OSF) at https://osf.io/cegqk.
Results:
A total of 257 records were identified, of which 24 studies met the inclusion criteria and contributed 296 experimental evaluations overall. Study quality was high in 7 studies, moderate in 8, and limited by technical or methodological concerns in 9. Hybrid deep learning (Hybrid DL) was most often used as the main automated coding approach, while machine learning (ML) and rule-based approaches were mainly used as baselines. F1-macro was consistently lower than F1-micro among studies reporting both metrics. Hybrid DL showed the most stable performance under all-code or full-code evaluation, while AI model performance varied by the documents-per-label (D/L) ratio.
Discussion:
The evidence indicates continued technical progress, particularly through Hybrid DL and transformer-based approaches, while LLM-based methods remain emerging and less consistently effective for structured multi-label coding. The observed D/L-performance relationship suggested that AI model selection should consider dataset structure and label support, in addition to algorithmic complexity.
Conclusion:
AI-based automatic ICD-10 coding is a promising approach for clinical coding support. Future research should prioritize rare-label imbalance, reproducibility, explainability, and validation across diverse clinical settings.
Systematic Review Registration:
https://osf.io/cegqk.
