Related Experiment Video
Updated: Jan 9, 2026

A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
Chronic Disease Monitoring: Methodology for Classification Error and Self-Selection Bias Correction in Clinical
Jesuan Betancourt1, Efrain Betancourt1, Abiel Roche-Lima2
1Abartys Health, San Juan, PR 00907-3913, USA.
Abstract:
Background/Objectives: Chronic diseases are among the leading causes of morbidity and healthcare costs worldwide. Diabetes mellitus is one of the most prevalent and costly chronic conditions in the United States, with a disproportionate burden in Puerto Rico. Surveillance of diabetes relies mainly on infrequent cohort studies and self-report surveys, which are limited in accuracy, segmentation, and timeliness. This study aimed to develop a generalizable methodology for monitoring chronic disease prevalence using routinely collected laboratory data, while correcting for systematic biases and diagnostic errors. Methods: We analyzed more than five years of de-identified laboratory test results (2020-2024) from a large, island-wide network of clinical laboratories in Puerto Rico. To produce unbiased prevalence estimates, we applied a mathematical correction framework that accounted for two main sources of distortion: (1) classification errors from treatment effects and test limitations, quantified through confusion matrices derived from longitudinal records; and (2) self-selection bias from differential testing rates, estimated empirically by demographic segment. Demographic reweighting ensured representativeness with respect to census data. Results: Using diabetes as a test case, corrected estimates for 2024 showed an adult prevalence of 18.0%, compared to 14.1% based on raw laboratory frequencies. The large amount of data provided high-resolution estimates by age, sex, and location, enabling fine-grained detection of demographic and geographic disparities. Conclusions: Bias-corrected laboratory surveillance provides accurate, timely, and demographically representative estimates of chronic disease prevalence. The methodology is scalable, cost-effective, and broadly applicable to other multi-stage chronic conditions, offering a foundation for next-generation public health monitoring and targeted interventions.
More Related Videos
07:15Determining the Likelihood of Variant Pathogenicity Using Amino Acid-level Signal-to-Noise Analysis of Genetic Variation
Published on: January 16, 2019
11:21Methodology for Establishing a Community-Wide Life Laboratory for Capturing Unobtrusive and Continuous Remote Activity and Health Data
Published on: July 27, 2018
Related Concept Videos
Bias in Epidemiological Studies
Errors occurring during blood pressure monitoring
Several factors...
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Strategies for Assessing and Addressing Confounding
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Data Validation
Key parameters for method validation include: