Related Experiment Video
Updated: Jan 22, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Machine Learning Assisted Discovery of Novel Predictive Lab Tests Using Electronic Health Record Data
Ross Kleiman1,2, Finn Kuusisto3,2, Ian Ross1
1University of Wisconsin - Madison, Madison, WI.
This study introduces a new pipeline to efficiently generate and rank potential laboratory tests for disease diagnosis. It uses machine learning and text mining to identify promising diagnostic markers from electronic health records and scientific literature.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Epidemiology
Background:
- Identifying biological markers for disease diagnosis is crucial but challenging.
- Current methods are often time-consuming, expensive, and require significant expertise.
- Improving the quality of initial marker hypotheses is essential for efficient research.
Purpose of the Study:
- To develop a high-throughput computational pipeline for generating and ranking hypothesized laboratory tests for disease diagnosis.
- To enhance the efficiency and reduce the cost of identifying novel diagnostic markers.
- To improve the accuracy of initial hypotheses for epidemiological studies.
Main Methods:
- A pipeline combining machine learning models for hypothesis generation.
- Text mining techniques for filtering and ranking hypotheses based on novelty.
- Logistic regression analysis to corroborate final candidate hypotheses.
- Validation using electronic health record data and the PubMed corpus.
Main Results:
- The pipeline successfully generated a large number of candidate laboratory test-diagnosis hypotheses.
- Novelty assessment using text mining effectively filtered and ranked potential markers.
- Logistic regression confirmed several promising candidate hypotheses.
- The approach demonstrated feasibility on real-world datasets.
Conclusions:
- The proposed high-throughput pipeline offers an efficient and effective method for identifying high-quality hypothesized diagnostic markers.
- This approach can accelerate the discovery of new laboratory tests for disease diagnosis.
- The integration of machine learning, text mining, and statistical analysis provides a robust framework for biomarker discovery.
More Related Videos
09:34A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018
Related Concept Videos
Data Reporting and Recording
Purpose of Health Records I
Here's a breakdown of how health records serve these purposes:
Purpose of Health Records II
Predicting Molecular Geometry
Machines
A free-body diagram of the...
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...