Related Experiment Video
Updated: Feb 7, 2026

A Data-Driven Approach to Quantifying Immune States in Sepsis
Published on: February 7, 2025
Comparing the validity of different ICD coding abstraction strategies for sepsis case identification in German claims
Carolin Fleischmann-Struzek1, Daniel O Thomas-Rüddel1,2, Anna Schettler2
1Integrated Research and Treatment Center, Center for Sepsis Control and Care (CSCC), Jena University Hospital, Jena, Germany.
Introduction:
Administrative data are used to generate estimates of sepsis epidemiology and can serve as source for quality indicators. Aim was to compare estimates on sepsis incidence and mortality based on different ICD-code abstraction strategies and to assess their validity for sepsis case identification based on a patient sample not pre-selected for presence of sepsis codes.
Materials And Methods:
We used the national DRG-statistics for assessment of population-level sepsis incidence and mortality. Cases were identified by three previously published International Statistical Classification of Diseases (ICD) coding strategies for sepsis based on primary and secondary discharge diagnoses (clinical sepsis codes (R-codes), explicit coding (all sepsis codes) and implicit coding (combined infection and organ dysfunction codes)). For the validation study, a stratified sample of 1120 adult patients admitted to a German academic medical center between 2007-2013 was selected. Administrative diagnoses were compared to a gold standard of clinical sepsis diagnoses based on manual chart review.
Results:
In the validation study, 151/937 patients had sepsis. Explicit coding strategies performed better regarding sensitivity compared to R-codes, but had lower PPV. The implicit approach was the most sensitive for severe sepsis; however, it yielded a considerable number of false positives. R-codes and explicit strategies underestimate sepsis incidence by up to 3.5-fold. Between 2007-2013, national sepsis incidence ranged between 231-1006/100,000 person-years depending on the coding strategy.
Conclusions:
In the sample of a large tertiary care hospital, ICD-coding strategies for sepsis differ in their accuracy. Estimates using R-codes are likely to underestimate the true sepsis incidence, whereas implicit coding overestimates sepsis cases. Further multi-center evaluation is needed to gain better understanding on the validity of sepsis coding in Germany.
Related Concept Videos
Data Validation
Key parameters for method validation include:
Data Validation
Nursing assessment guides are generally based on holistic models rather than medical...
lncRNA - Long Non-coding RNAs
Radical Formation: Abstraction
Even though homolysis produces radicals, it is different from radical...
Testing a Claim about Mean: Known Population SD
Estimating a population mean requires the samples to be distributed normally. The data should be collected from the randomly selected samples having no sampling bias. The sample size needed to be higher than 30, and most importantly, the population standard deviation should be already known.
In most realistic situations, the population standard deviation is often unknown, but in rare circumstances, when it...
Reliability and Validity

