Related Experiment Video
Updated: Jul 11, 2025

07:41
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
9.0K
Development of a natural language processing model for deriving breast cancer quality indicators : A cross-sectional,
Etienne Guével1, Sonia Priou2, Rémi Flicoteaux3
1Assistance Publique - Hôpitaux de Paris, Innovation and Data, IT Department, 75012 Paris, France.
Summary
Automating healthcare quality indicators using electronic health records faces challenges. Data source availability and natural language processing algorithm performance limit automated indicator calculation.
Area of Science:
- Health Informatics
- Medical Data Analysis
- Quality Improvement
Background:
- Medico-administrative data offers potential for automated Healthcare Quality and Safety Indicators (HQSI) calculation.
- However, existing data sources are often insufficient for calculating all relevant indicators.
- This study investigates the feasibility of enhancing HQSI calculation using multiple data sources and natural language processing (NLP).
Purpose of the Study:
- To assess the availability of data sources for HQSI calculation.
- To determine the availability of elementary variables required for HQSI computation.
- To apply NLP techniques for automated data extraction from medical reports.
Main Methods:
- A multicenter, cross-sectional observational study was conducted using data from Assistance Publique - Hôpitaux de Paris (AP-HP).
- Breast cancer patient data (January 2019 - June 2021) were analyzed using Programme de Médicalisation du Système d'Information (PMSI) claims data and pathology reports.
- Rule-based NLP algorithms were developed and validated to extract data from free-text pathology reports, with performance metrics including recall, precision, and F1-score.
Main Results:
- Out of 5785 breast cancer patients, 89.0% had relevant PMSI data and 72.5% had surgery.
- Nine indicators were computable using PMSI data alone; an additional six became computable using pathology reports.
- NLP algorithms achieved an average accuracy of 76.5%, precision of 77.7%, and recall of 71.6%, enabling variable extraction for 2%-88% of patients depending on the indicator.
Conclusions:
- The availability of medical reports and elementary variables within them, along with NLP algorithm performance, restricts the number of patients for whom indicators can be calculated.
- Automated calculation of quality indicators from electronic health records presents significant practical obstacles.
Keywords:
Electronic Data ProcessingHealth CareIndicateurs de qualitéNatural Language ProcessingQuality IndicatorsSoins de santéTraitement du langage naturelTraitement électronique de données
