A pipeline for the fully automated estimation of continuous reference intervals using real-world data.
Tatjana Ammer1,2, André Schützenmeister2, Hans-Ulrich Prokosch1
1Chair of Medical Informatics, Friedrich-Alexander-Universität Erlangen-Nürnberg, Erlangen, Germany.
This article introduces a new, fully automated computer program that calculates continuous reference ranges for medical lab tests using routine patient data. By accounting for age-related changes, this tool helps doctors interpret test results more accurately without needing samples from healthy volunteers. The system provides high-quality, standardized outputs that can be easily integrated into hospital software.
Area of Science:
- Clinical chemistry and laboratory medicine research
- Computational biology and refineR statistical modeling
Background:
Standard medical practice relies on reference intervals to interpret laboratory test results accurately. Prior research has shown that continuous reference intervals better capture physiological dynamics across the human lifespan. That uncertainty drove the need for methods that move beyond static, age-grouped thresholds. Current approaches often require samples from strictly healthy individuals, which limits their broad clinical application. This gap motivated the development of techniques that utilize routine patient data instead. No prior work had resolved the challenge of automating these complex calculations while accounting for continuous covariates like age. Existing indirect methods typically only provide one-dimensional estimates for specific populations. This study addresses the lack of automated, integrated pipelines for generating these dynamic, age-dependent clinical benchmarks.
Purpose Of The Study:
The aim of this study is to develop a fully automated pipeline for estimating continuous reference intervals using real-world laboratory data. This research addresses the limitations of current methods that rely on samples from healthy individuals. The authors seek to overcome the restriction of static, age-grouped reference ranges in clinical practice. They propose using generalized additive models for location, scale, and shape to capture age-specific physiological dynamics. The motivation is to provide a scalable, parameter-less solution that removes subjective user-defined settings. By utilizing routine measurements, the researchers intend to make high-quality reference intervals more accessible for clinical decision-making. They also aim to enable the conversion of test results into standardized z-scores for improved interpretation. Finally, the study seeks to demonstrate that this automated approach can be integrated seamlessly into existing laboratory information systems.
Main Methods:
The researchers developed an integrated computational pipeline to automate the estimation of continuous reference intervals. Their review approach involved utilizing generalized additive models for location, scale, and shape to handle complex data distributions. The team leveraged existing discrete model estimates derived from the refineR software package. They designed the workflow to operate entirely without subjective user-defined parameters or manual intervention. To validate the performance, they compared their automated outputs against established, validated reference intervals from the CALIPER and PEDREF studies. They also assessed the pipeline against various manufacturers' package inserts to ensure broad applicability. The design focuses on transforming raw laboratory test results into standardized z-scores for clinical utility. Finally, they ensured the entire system architecture supports direct integration into standard laboratory information systems.
Main Results:
The automated pipeline generates high-quality reference intervals that show strong agreement with established clinical benchmarks. Comparisons against the CALIPER and PEDREF studies confirm the reliability of the derived reference limits. The system successfully produces continuous reference intervals that capture physiological age-specific dynamics without requiring healthy donor samples. The results are entirely free from subjective user-input, ensuring consistency across different laboratory settings. The pipeline enables the conversion of raw test results into z-scores, facilitating more precise clinical interpretations. The researchers report that their method functions as a fully automated, parameter-less solution for indirect estimation. The generated percentile charts demonstrate high precision across the tested laboratory parameters. The study confirms that the proposed approach effectively integrates continuous covariates into the estimation process for routine clinical measurements.
Conclusions:
The authors propose a novel, parameter-less solution for generating continuous reference intervals using routine clinical data. Their pipeline successfully produces high-precision percentile charts for diverse laboratory parameters. Synthesis and implications suggest that this automated approach removes the need for subjective user-defined settings. The researchers demonstrate that their results align well with established benchmarks from the CALIPER and PEDREF studies. These findings imply that the tool can reliably convert raw test measurements into standardized z-scores. Integration into existing laboratory information systems appears feasible based on the reported technical performance. The study confirms that indirect estimation methods can achieve high quality without relying on healthy donor cohorts. This work provides a scalable framework for improving clinical decision-making through more precise, age-specific laboratory interpretations.
Frequently Asked Questions
The researchers propose an integrated pipeline utilizing generalized additive models for location, scale, and shape. This mechanism processes routine patient measurements to estimate continuous reference intervals, effectively replacing manual, subjective inputs with a fully automated, parameter-less computational workflow.
The authors employ the refineR tool to generate discrete model estimates. This software component serves as the foundation for the pipeline, enabling the indirect derivation of reference limits from large, non-curated clinical datasets, unlike traditional methods that require strictly healthy donor samples.
The authors state that integrating age as a continuous covariate is necessary to capture physiological dynamics throughout life. This technical requirement ensures that the resulting reference intervals are not static, providing a more precise interpretation of laboratory values across different developmental stages.
The pipeline uses routine laboratory data to estimate reference intervals. This data type allows for the inclusion of large, real-world patient populations, whereas traditional approaches rely on curated samples from healthy individuals, which are often difficult and expensive to obtain.
The researchers measure the agreement of their reference limits against established benchmarks from the CALIPER and PEDREF studies. This measurement demonstrates that their automated pipeline produces high-quality, validated results comparable to those derived from traditional, manual clinical studies.
The authors propose that their pipeline enables the conversion of test results into z-scores. They suggest this implication will improve clinical decision-making by providing standardized, age-specific interpretations that can be seamlessly integrated into existing laboratory information systems.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This...
Kaplan-Meier Approach
Confidence Intervals
A...
Confidence Interval for Estimating Population Mean
A confidence interval for the mean is a range of values that provides an estimate of the population mean. As the...
Interpretation of Confidence Intervals
Confidence intervals have confidence coefficients that are crucial for their interpretation. The most common confidence coefficients are 0.90, 0.95, and 0.99, which can be written as percentages–90%, 95%, and 99%, respectively.
Suppose a person calculates a confidence interval with a confidence coefficient of 0.95. In that case, they can...


