Related Experiment Video
Updated: Jan 5, 2026

Identification of Disease-related Spatial Covariance Patterns using Neuroimaging Data
Published on: June 26, 2013
Detecting the impact of subject characteristics on machine learning-based diagnostic applications
Elias Chaibub Neto1, Abhishek Pratap1,2, Thanneer M Perumal1
11Sage Bionetworks, Seattle, USA.
Digital health studies often underestimate prediction error due to "identity confounding" from repeated measures. A new method quantifies this issue, showing record-wise data splits in machine learning must be avoided.
Area of Science:
- Digital Health
- Machine Learning
- Biostatistics
Background:
- High-dimensional, longitudinal digital health data offer significant research and clinical potential.
- Developing diagnostic algorithms requires careful consideration of repeated measurements from individuals.
- Current analytical evaluations of predictive performance often overlook the impact of repeated measures.
Purpose of the Study:
- To present a method for quantifying "identity confounding" in digital health classifiers.
- To demonstrate the prevalence and impact of identity confounding in real-world digital health datasets.
- To advocate for the avoidance of record-wise data splits in machine learning applications for digital health.
Main Methods:
- Development of a novel method to calculate identity confounding.
- Application of the method to multiple real-world digital health datasets.
- Evaluation of predictive performance in machine learning models trained with different data splitting strategies.
Main Results:
- Record-wise data splits, where repeated measures from individuals are present in both training and testing sets, lead to massive underestimation of prediction error.
- Identity confounding, where models learn to identify subjects rather than diagnostic signals, is a significant issue in digital health studies.
- The proposed method effectively quantifies the extent of identity confounding present in classifiers.
Conclusions:
- Identity confounding is a critical problem in digital health research, biasing predictive performance evaluations.
- Record-wise data splits are inappropriate for developing and evaluating machine learning-based digital health applications.
- Future research must adopt data splitting strategies that mitigate identity confounding to ensure reliable diagnostic algorithms.
More Related Videos
07:15Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Receiver Operating Characteristic Plot
Classification of Illness
An illness is a response to a disease in which the person's level of functioning is changed compared with a previous level. The general classification of illness includes acute and chronic.
Acute illness is severe...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Steps in Outbreak Investigation