Related Experiment Videos
Benchmarking Statistical, Machine Learning, and Exploratory Deep Learning Models for the Short-Term Forecasting of
Ao Li1, Ruijia Shi1, Yanlin Wei1
1Beijing Institute of Ophthalmology, Beijing Tongren Eye Center, Beijing Tongren Hospital, Capital Medical University, Beijing Ophthalmology & Visual Sciences Key Laboratory, Beijing 100730, China.
Abstract:
Background/Objectives: First NIBUT, Average NIBUT, and tear meniscus height (TMH) are routinely used to characterize tear film stability and tear volume at individual visits, whereas their longitudinal behavior across the hospital-attending population is less well characterized. Aggregating routine examinations over time may provide a continuous hospital-level view of ocular surface status and enable the short-term forecasting of expected trajectories. We therefore evaluated a benchmark-first framework for monthly aggregated adult ocular surface indicators. Methods: This retrospective time-series study used de-identified adult eye-level Keratograph 5M records from July 2018 to November 2023. After cleaning, 31,492 records from 13,749 patients and 15,334 examination occasions were aggregated across 65 calendar months (64 observed months; March 2020 had no eligible records). Seven benchmark models and two exploratory deep learning comparators were evaluated using eight rolling-origin 3-month test windows. Results: Linear trend had the lowest mean origin-level macro-normalized RMSE (0.875; 95% bootstrap CI, 0.497-1.421), followed by simple exponential smoothing (0.907) and SARIMA(1,0,0)(1,0,0,12) (0.940). The paired difference between linear trend and simple exponential smoothing was small and did not show clear superiority (mean difference, -0.032; 95% bootstrap CI, -0.217 to 0.135; p = 0.789). The model with the lowest pooled error differed by target, while patient-month and sample-size-weighted sensitivity analyses gave a similar overall benchmark pattern. The exploratory LSTM and Transformer did not show a consistent advantage over the leading simple models. Conclusions: In this short hospital-level monthly series, simple forecasting models remained competitive, while no single model showed consistent superiority across forecast origins and sensitivity analyses. By extending ocular surface assessment from isolated examinations to longitudinal hospital-level trajectories, this framework provides a methodological basis for monitoring temporal changes in tear film stability and tear volume and for future quality monitoring, clinical, and epidemiological applications.