Related Experiment Video
Updated: Sep 10, 2025

Structured Approach to Colonoscopy Technique Optimization: A Single-Center Experience with Novice Endoscopists
Published on: July 11, 2025
Enhancing and Not Replacing Clinical Expertise: Improving Named-Entity Recognition in Colonoscopy Reports Through
Andrei-Constantin Ioanovici1, Andrei-Marian Feier2, Marius-Ștefan Mărușteri1
1Department M2-Complementary Functional Sciences, Medical Informatics and Biostatistics, George Emil Palade University of Medicine, Pharmacy, Science, and Technology of Targu Mures, 540142 Targu Mures, Romania.
Training named-entity recognition (NER) models on a mix of real and synthetic colonoscopy reports significantly improves information extraction. This approach enhances data privacy and supports future research in colorectal cancer surveillance.
Area of Science:
- Natural Language Processing
- Medical Informatics
- Artificial Intelligence in Healthcare
Background:
- Colonoscopy findings are typically unstructured free text, hindering secondary data use for quality monitoring, personalized medicine, and research.
- Accurate named-entity recognition (NER) is crucial for extracting valuable clinical information from these free-text reports.
Purpose of the Study:
- To compare the performance of NER models trained on real, synthetic, and mixed colonoscopy report data.
- To evaluate if privacy-preserving synthetic reports can enhance clinical information extraction accuracy.
Main Methods:
- Three Spark NLP biLSTM CRF models were trained: ModelR (real data), ModelS (synthetic data), and ModelM (mixed real/synthetic data).
- Models were evaluated on unseen real and synthetic colonoscopy reports for seven entity types.
- Performance metrics included micro-averaged precision, recall, and F1-score, with McNemar tests for statistical significance.
Main Results:
- The mixed-data model (ModelM) achieved the highest performance (F1 = 0.94), significantly outperforming single-source models (ModelR F1 = 0.70, ModelS F1 = 0.64).
- ModelR performed well on real data (F1 = 0.90) but poorly on synthetic data (F1 = 0.47); ModelS showed the opposite trend (F1 = 0.99 synthetic, F1 = 0.33 real).
- Statistical analysis confirmed the significant superiority of the mixed-data approach.
Conclusions:
- Synthetic colonoscopy reports are a valuable supplement but not a replacement for real annotated data.
- Training NER models on a balanced mix of real and synthetic data yields robust, generalizable models for structuring colonoscopy reports.
- This approach supports large-scale, privacy-preserving colorectal cancer surveillance and personalized follow-up strategies.
More Related Videos
08:26Clinical Application of Single-Surgeon, Three-Port, Laparoscopic Resection for Colorectal Cancer with Natural Orifice Specimen Extraction
Published on: March 24, 2023
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Endoscopic Procedures II: Colonoscopy
Imaging Studies III: Gastrointestinal Motility Studies and Virtual Colonoscopy
Radionuclide Testing
Radionuclide testing is a sophisticated medical technique for assessing gastrointestinal motility. It focuses on gastric emptying and colonic transit time. Radioactive markers track the movement of food through the digestive system, providing insights into gastrointestinal disorders.
In gastric emptying studies, a meal's liquid and...
Improving Translational Accuracy
Endoscopic Procedures III: Video Capsule Endoscopy