Related Experiment Video
Updated: Feb 28, 2026

Quantification of Vascular Parameters in Whole Mount Retinas of Mice with Non-Proliferative and Proliferative Retinopathies
Published on: March 12, 2022
P-Score Variability as a Function of Retinopathy of Prematurity Severity
Sarthak V Shah1, Arthur R Brant1, Cindy S Zhao1
1Byers Eye Institute, Horngren Family Vitreoretinal Center, Department of Ophthalmology, Stanford University School of Medicine, Palo Alto, California, United States.
Purpose:
To evaluate the variability of P-score grading within the same eye during a single imaging session (P-Var) as a function of retinopathy of prematurity (ROP) severity.
Methods:
This study analyzed retinal images from 20 consecutive subjects selected from the Stanford University Network for Diagnosis of ROP (SUNDROP) database who were screened between January 2023 and December 2024 and subsequently received treatment. Nine masked graders evaluated the P-scores of 1597 images from 129 imaging sessions. Variability was measured through calculation of range, intergrader reliability, and agreement metrics. Probabilistic modeling was conducted to assess the impact of limited image sampling on disease classification.
Results:
Significant P-score variability was observed within single eye visits, increasing with worsening vascular disease (maximum eye visit P-score: β = 0.222, P < 0.001) and overall disease severity as assessed with the Telemedicine ROP Severity Score (β = 0.0073, P < 0.001). Intergrader agreement for the nine graders was good (intraclass correlation coefficient [ICC] = 0.68; 95% confidence interval [CI], 0.65-0.71). Mean ± SD weighted kappa was 0.66 ± 0.10 between grader pairs. Exact P-score agreement was 28.5% ± 5.6%, improving to 72.5% ± 8.2% when allowing a ±1 P-score unit difference. The coefficient of variation of image-level P-scores between graders decreased with worsening disease. Probabilistic modeling demonstrated that limited image sampling with fewer than five views per eye can lead to significant misclassification, particularly in plus eyes.
Conclusions:
P-score variability exists within the same imaging session and increases with disease severity. These findings suggest that multiple images per session are necessary for accurate disease classification, particularly in eyes with more severe disease.
Related Concept Videos
Variation
When independent and dependent variables are plotted on a scatter plot, the slope of a line is a value that describes the rate of change between the two...
z Scores and Area Under the Curve
Receiver Operating Characteristic Plot
P-value
P-value stands for the probability value. P-value is the probability that, if the null hypothesis is true, the results from another randomly selected sample will be as extreme or more extreme as the results obtained from the given sample.
A large P-value calculated from the data indicates to not reject the null hypothesis. But a higher P-value does not mean that the null hypothesis is true. The smaller the P-value, the more...

