Related Experiment Video
Updated: Aug 7, 2026

Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Mining a clinical data warehouse to discover disease-finding associations using co-occurrence statistics
Hui Cao1, Marianthi Markatou, Genevieve B Melton
1Department of Biomedical Informatics, Columbia University, New York, NY, USA.
This study used co-occurrence statistics to identify disease-finding links in clinical data. Both chi2 statistics and the proportion confidence interval (PCI) method proved effective in building knowledge bases for clinical decision support.
Area of Science:
- Medical Informatics
- Clinical Data Mining
- Computational Medicine
Background:
- Clinical data warehouses contain vast amounts of information on diseases and patient findings.
- Identifying reliable associations between diseases and clinical findings is crucial for improving patient care and medical research.
- Automated systems for summarizing patient information often require robust knowledge bases of disease-finding relationships.
Purpose of the Study:
- To apply co-occurrence statistics for discovering disease-finding associations within a clinical data warehouse.
- To evaluate the accuracy of identified associations using both chi2 statistics and the proportion confidence interval (PCI) method.
- To construct knowledge bases (KB-chi2, KB-PCI) from these associations and assess their utility in clinical decision support.
Main Methods:
- Utilized co-occurrence statistics to measure the dependence between pairs of diseases and findings.
- Employed chi2 statistics and the proportion confidence interval (PCI) method for association measurement.
- Applied heuristic cutoff values for selecting significant disease-finding associations.
- Constructed two knowledge bases: KB-chi2 and KB-PCI.
Main Results:
- An intrinsic evaluation demonstrated high accuracy: 94% for chi2 statistics and 76.8% for the PCI method identified true disease-finding associations.
- An extrinsic evaluation confirmed that both KB-chi2 and KB-PCI effectively assisted in removing clinically non-informative and redundant findings.
- The developed knowledge bases proved valuable for enhancing automated problem list summarization systems.
Conclusions:
- Co-occurrence statistics, particularly chi2 and PCI methods, are effective for discovering reliable disease-finding associations in clinical data.
- The constructed knowledge bases (KB-chi2, KB-PCI) can significantly improve the quality of automated clinical summaries.
- This approach offers a valuable tool for enhancing clinical decision support and medical informatics research.
Related Concept Videos
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Statistical Software for Data Analysis and Clinical Trials
Investigation of Disease Outbreaks
Statistical Methods for Analyzing Epidemiological Data
Steps in Outbreak Investigation