Related Experiment Videos
Information mining over heterogeneous and high-dimensional time-series data in clinical trials databases.
Fatih Altiparmak1, Hakan Ferhatosmanoglu, Selnur Erdal
1Department of Computer Science and Engineering, The Ohio State University, Columbus, OH 43210, USA.
Summary
This study introduces a new method for analyzing complex clinical trial data, focusing on heterogeneous, high-dimensional time series. The approach identifies key biological markers for health status, improving data mining in biomedical research.
Area of Science:
- Biomedical Data Analysis
- Computational Biology
- Clinical Informatics
Background:
- Clinical trial data often presents challenges like heterogeneity and high dimensionality in time series.
- Existing time series analysis methods frequently fail due to unmet assumptions regarding data length and interval consistency.
- These limitations are particularly pronounced in clinical trial datasets, necessitating revised data mining approaches.
Purpose of the Study:
- To propose a novel information mining approach tailored for heterogeneous, high-dimensional time series clinical trial data.
- To develop methods for identifying common and distinct patterns within complex biomedical datasets.
- To enhance the analysis of time series data by addressing limitations of current statistical techniques.
Main Methods:
- A two-step approach involving data mining on homogeneous subsets and pattern identification.
- Application of frequent itemset mining, clustering, and declustering techniques with novel distance metrics.
- Utilizing modified algorithms for effective declustering and feature selection in high-dimensional time series.
Main Results:
- Successful clustering identified strongly correlated analyte groups, validating known biological relationships and discovering novel ones.
- Declustering enabled effective feature selection, identifying a concise set of analytes modeling normal health.
- The proposed framework demonstrated efficacy on industry-sponsored clinical trials data.
Conclusions:
- The novel approach effectively mines heterogeneous, high-dimensional time series clinical trial data.
- The methods facilitate the discovery of significant analyte correlations and enable robust feature selection.
- This work provides a valuable framework for advancing biomedical data analysis and biomarker discovery.