A Research Classification for Long COVID: Symptom Frequency and Severity Improve Accuracy of Machine Learning Models
Leonard A Jason1, Lauren Ruesink1, Jacob Furst1
1DePaul University Chicago Illinois USA.
Chronic Diseases and Translational Medicine
|July 24, 2026
Summary
Defining Long COVID by symptom occurrence alone lacks specificity. Machine learning models using symptom frequency and severity offer more accurate and efficient classification for Long COVID research.
Area of Science:
- Medical research
- Data science
- Public health
Background:
- Current definitions of Long COVID rely heavily on symptom presence, potentially limiting diagnostic accuracy.
- The complexity of Long COVID symptoms necessitates more refined classification methods.
Purpose of the Study:
- To evaluate machine-learning models for Long COVID classification.
- To compare the diagnostic performance of models using symptom occurrence versus symptom burden (frequency and severity).
Main Methods:
- Development and comparison of machine-learning models.
- Models were trained using symptom occurrence data and composite scoring incorporating frequency and severity.
- Performance was evaluated based on diagnostic accuracy and the number of predictive symptoms required.
Main Results:
- Composite scoring models achieved higher accuracy (90.12%) than occurrence-only models (88.73%).
- Models incorporating symptom frequency and severity were more efficient, requiring fewer predictive symptoms.
- This indicates that measuring symptom burden enhances precision in Long COVID research classification.
Conclusions:
- Symptom burden, including frequency and severity, is crucial for accurate Long COVID classification.
- Machine-learning approaches integrating symptom burden improve diagnostic specificity and research efficiency.
- Refined classification methods are essential for advancing Long COVID research and patient care.
