Related Experiment Video
Updated: Jul 1, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Early identification of risk factors for obstructive sleep apnea hypopnea syndrome based on large language models
Lei Cheng1, Juan Bai1, Aizhu Liu1
1Department of Otolaryngology Head and Neck Surgery, Capital Medical University Affiliated Beijing Shijitan Hospital, Beijing, China.
Background:
Obstructive sleep apnea hypopnea syndrome (OSAHS) is a highly prevalent sleep-related breathing disorder, yet it is frequently missed in its early stages. Early identification of OSAHS-related risk factors and early signals from real-world patient-generated text may enable timely screening and intervention. However, existing approaches primarily rely on structured clinical data or end-to-end classification models, which are limited in handling unstructured, colloquial descriptions and highly imbalanced class distributions.
Objectives:
This study aimed to develop and evaluate a large language model-based framework for early identification of OSAHS-related risk factors and early signals from free-text descriptions, enabling robust text-level risk stratification with improved interpretability.
Methods:
We proposed OSAHSrisk-LLM, a large language model-based framework that analyzes patient-generated text using a relevance-aware and ontology-constrained reasoning strategy. The framework first determines whether a text contains OSAHS-related risk information, and then applies stepwise, evidence-based reasoning to identify early signals and risk factors under predefined clinical knowledge guidance. Linguistic variability in real-world narratives is addressed by normalizing extracted concepts into standardized clinical terms. We evaluated the framework on a corpus of real-world Chinese text describing sleep-related experiences and compared it with multiple baseline text classification models.
Results:
OSAHSrisk-LLM achieved an overall accuracy of 92.9% in the four-class text-level classification task, numerically outperforming the baseline models including CNN, Text-CNN, Transformer, and BERT on the current dataset. Notably, the proposed framework demonstrated strong robustness under highly imbalanced class distributions and improved identification of minority categories.
Conclusion:
Our findings suggest that large language models, when integrated with clinical knowledge constraints and structured reasoning strategies, can effectively extract early OSAHS-related risk information from patient narratives. The proposed framework demonstrates the potential of large language models to identify textual mentions of OSAHS-related risk factors and early signals in patient-generated narratives. Further validation against clinically confirmed OSAHS diagnoses and prospective screening cohorts is required before use in real-world clinical risk screening.
