Leveraging Large Data, Statistics, and Machine Learning to Predict the Emergence of Resistant E. coli Infections

Rim Hur1,2,3, Stephine Golik1,4, Yifan She1,3

  • 1Department of Inpatient Pharmacy, Kaiser Permanente, One Kaiser Plaza, Oakland, CA 94612, USA.

PubMed

Insights

Reducing cefazolin use can lower resistant E. coli infections, a key finding for antimicrobial stewardship. This study models antibiotic resistance trends to predict and manage healthcare costs.

Area of Science:

  • Infectious Diseases
  • Computational Biology
  • Health Economics

Background:

  • Drug-resistant Gram-negative bacterial infections significantly increase hospital length of stay and costs.
  • Escherichia coli (E. coli) is a common pathogen, and resistance to antibiotics poses a growing threat.
  • Understanding the relationship between antibiotic usage and resistance is crucial for effective antimicrobial stewardship.

Purpose of the Study:

  • To explore the relationship between antibiotic usage and antibiotic resistance over time using statistical and machine-learning models.
  • To predict the clinical and financial costs associated with resistant E. coli infections.
  • To provide a framework for antimicrobial stewardship programs (ASPs) to use data-driven interventions.

Main Methods:

  • Acquired data on antibiotic utilization and microbial culture resistance/sensitivity from a Kaiser Permanente facility (April 2013 - December 2019).
  • Employed time series algorithms including Autoregressive Integrated Moving Average (ARIMA), neural networks, and random forest to model antibiotic resistance trends.
  • Evaluated model performance using Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE), with the best model predicting 2020 resistance rates.

Main Results:

  • The ARIMA model demonstrated the best performance for predicting antibiotic resistance trends, particularly with cefazolin and cephalexin.
  • Reducing cefazolin usage was identified as a potential strategy to decrease the rate of resistant E. coli infections.
  • Piperacillin/tazobactam, despite not being the top performer in models, shows potential as an intervention target in ASPs due to its broad spectrum.

Conclusions:

  • Statistical and machine-learning models can effectively predict antibiotic resistance trends and inform antimicrobial stewardship interventions.
  • Region-specific data is valuable for tailoring interventions within ASPs.
  • This study provides a framework for adopting advanced analytical approaches to combat antibiotic resistance and manage associated costs.

Related Concept Videos

Steps in Outbreak Investigation01:18

Steps in Outbreak Investigation

In the ever-evolving field of public health, statistical analysis serves as a cornerstone for understanding and managing disease outbreaks. By leveraging various statistical tools, health professionals can predict potential outbreaks, analyze ongoing situations, and devise effective responses to mitigate impact. For that to happen, there are a few possible stages of the analysis:
126
Statistical Methods for Analyzing Epidemiological Data01:25

Statistical Methods for Analyzing Epidemiological Data

Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
364
Statistical Software for Data Analysis and Clinical Trials01:12

Statistical Software for Data Analysis and Clinical Trials

Statistical software is pivotal in data analysis and clinical trials by providing tools to analyze data, draw conclusions, and make predictions. These software packages range from simple data management applications to complex analytical platforms, supporting various statistical tests, models, and simulation techniques. Their significance lies in their ability to handle vast amounts of data with precision and efficiency, enabling researchers to validate hypotheses, identify trends, and make...
547
Residuals and Least-Squares Property01:11

Residuals and Least-Squares Property

The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K