Related Experiment Video
Updated: Sep 24, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
A step toward inclusion: A transparent dataset for automatic text simplification for screen reader users
Nelson Pérez-Rojas1, Martín Solís2, Saúl Calderón-Ramírez2
1Doctorado en Ciencias Naturales para el Desarrollo (DOCINADE), Instituto Tecnológico de Costa Rica, Universidad Nacional, Universidad Estatal a Distancia, Cartago, Costa Rica.
Abstract:
GeoSimp is a Spanish-language dataset for automatic text simplification (ATS), comprising approximately 3200 discourse segments derived from peer-reviewed geological scientific articles. The dataset was constructed to address two documented gaps in existing ATS resources: the limited availability of Spanish-language corpora for text simplification and text complexity classification, and the scarcity of domain-specific datasets focused on specialized scientific writing. Source texts were drawn exclusively from articles published in the Revista Geológica de América Central, a peer-reviewed, open-access journal, all of which are distributed under a single Creative Commons Attribution-NonCommercial-ShareAlike 3.0 (CC BY-NC-SA 3.0) license. Text extraction was performed manually to preserve structural coherence. Paragraphs were copied into Microsoft Excel spreadsheets, assigned unique identifiers, and subsequently segmented into discourse segments by one of the authors, a researcher with expertise in linguistics. Four annotators with university degrees in Spanish philology and professional experience in text editing performed attribute identification and simplification, using a predefined scheme of 21 linguistic complexity attributes. Prior to annotation, these annotators completed a structured training and calibration phase. The dataset is distributed as two Excel files. The training file (GeoSimp_train.xlsx) contains 2956 records, each comprising the original discourse segment, one simplified version, the annotator identifier, the identified complexity attributes, the paragraph identifier, and the identifier of the source article. The test file (GeoSimp_test.xlsx) contains 301 records, each comprising the original discourse segment and four independently produced simplified versions with their corresponding attribute annotations. A simplification guidelines manual used during annotation is also included in the repository. GeoSimp is suitable for training and evaluating ATS models, for developing text complexity classification models, and for research on paragraph-level simplification. The dataset is specifically oriented toward simplification for blind and low-vision users who rely on screen readers, a population underrepresented in the ATS literature. The explicit annotation of linguistic complexity attributes per segment supports reproducibility and fine-grained error analysis. The dataset is openly available in a public GitHub repository.

