Related Experiment Video
Updated: Jul 20, 2026

13:44
Detection of Architectural Distortion in Prior Mammograms via Analysis of Oriented Patterns
Published on: August 30, 2013
43.0K
Mammography reporting dataset with BI-RADS system for natural language processing applications: Addressing public
José Luis Vázquez Noguera1, Alejandro Torres-Hurtado2, Helena Gómez-Adorno2
1Universidad Americana, Asunción 1029, Paraguay.
Data in Brief
|July 4, 2025
Summary
This study introduces a new Spanish mammography report dataset for Natural Language Processing (NLP) research. The dataset aids in developing AI tools for automating BI-RADS classification and analyzing clinical data.
Area of Science:
- Medical Informatics
- Natural Language Processing
- Radiology
Background:
- Automating clinical data analysis using Natural Language Processing (NLP) is crucial for enhancing diagnostic accuracy and healthcare efficiency.
- Mammography reports contain vital information for breast cancer screening and diagnosis.
Purpose of the Study:
- To present a novel, anonymized dataset of Spanish mammography reports for NLP and machine learning research.
- To facilitate the development of automated systems for classifying BI-RADS (Breast Imaging Reporting and Data System) categories and analyzing mammography findings.
Main Methods:
- A dataset of 4357 Spanish mammography reports was compiled from multiple medical units.
- Reports were segmented into sections (observations, conclusions, recommendations) and included BI-RADS classifications.
- Anonymization and manual verification ensured data quality and patient confidentiality.
Main Results:
- The dataset comprises 4357 records with 15 variables, including full report text, section-specific text, BI-RADS classification, and relevant metadata.
- The data is structured for direct application in developing and validating NLP and machine learning models.
Conclusions:
- This dataset represents a valuable resource for advancing AI in medical imaging analysis, particularly for mammography reporting.
- It supports research in automated BI-RADS classification and the extraction of clinical insights from unstructured text data.

