Related Experiment Video
Updated: Sep 14, 2026

Frailty Assessment in an Aging Mouse Model
Published on: September 23, 2025
Diagnostic Prediction Models for Oral Frailty in Older Adults: A Systematic Review and Critical Appraisal
Yuzhu Fan1, Shuang Zhang1, Yinuo Wang1
1School of Nursing and Rehabilitation, Nantong University, Nantong, Jiangsu, People's Republic of China.
Background:
Oral frailty is a clinically relevant marker of vulnerability in older adults, but the quality, performance, and applicability of multivariable models for its identification remain uncertain.
Objectives:
To identify and critically appraise multivariable prediction models for oral frailty and summarize their predictors, performance, validation, and risk of bias.
Methods:
Nine English and Chinese language databases were searched from inception to 14 July 2026, supplemented by citation and website searches. Two reviewers independently selected studies, extracted data, and assessed risk of bias and applicability using the Prediction Model Risk of Bias Assessment Tool (PROBAST). Owing to substantial heterogeneity, findings were synthesized narratively.
Results:
Twenty-three reports representing 22 independent studies and 22 models were included. Twenty-one studies were conducted in China, 21 used cross-sectional data, and 21 defined oral frailty as an Oral Frailty Index-8 score ≥4. All models were diagnostic and identified prevalent oral frailty at assessment; none predicted incident oral frailty. Logistic regression and nomograms were the most common approaches, although two studies compared multiple machine-learning algorithms. Common predictors included age, nutritional vulnerability, physical frailty or sarcopenia, chronic disease burden, smoking, and oral-function indicators. Reported area under the receiver operating characteristic curve (AUC) values ranged from 0.713 to 0.985. Internal validation was reported for 20 models, while six models underwent external validation (four temporal and two geographical); one model reported apparent performance only. Calibration was incompletely assessed, and no model had an overall low risk of bias.
Conclusions:
Existing models show generally moderate-to-high discrimination, but the evidence remains methodologically immature. Predominantly cross-sectional designs, high risk of bias, predictor-outcome overlap, incomplete calibration, and limited external validation, especially across geographical settings, preclude routine clinical use. Current models should be regarded as candidate diagnostic screening aids that complement standardized assessment; rigorous external validation and prospective prognostic modelling are priorities.
Prospero Registration:
CRD420251229868.

