Related Experiment Video
Updated: May 24, 2026

07:41
Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
Evaluation of data integration strategies based on kernel method of clinical and microarray data.
1Faculty of Computer Science, Universitas Indonesia.
Bioinformation
|February 28, 2012
Summary
Combining patient data using kernel methods did not consistently improve breast cancer classification accuracy. Feature selection is crucial for accurate patient stratification in bioinformatics.
Area of Science:
- Bioinformatics
- Cancer Research
- Computational Biology
Background:
- Accurate cancer classification is vital for effective treatment strategies.
- Distinguishing between breast cancer patients with and without distant metastases presents a significant challenge.
- Bioinformatics approaches are increasingly used to analyze complex biological datasets for improved diagnostics.
Purpose of the Study:
- To investigate the effectiveness of combining feature sets using kernel methods for classifying breast cancer patients based on distant metastases.
- To compare single data set performance against various data integration strategies, including weighted approaches.
- To evaluate the utility of Least Square Support Vector Machine (LS-SVM) for high-dimensional cancer data.
Main Methods:
- Utilized a dataset of 295 breast cancer patients from the Netherland Cancer Institute.
- Employed kernel methods for feature set combination and data integration.
- Applied Least Square Support Vector Machine (LS-SVM) as the primary classification algorithm.
- Compared classification performance across single data sets and multiple data integration strategies.
Main Results:
- Weighted late integration and the use of microarray data alone yielded similar classification performance.
- Data integration strategies did not universally outperform single data set analysis in this specific case.
- The choice and quality of features significantly impact the overall classification performance.
Conclusions:
- Data integration is not a guaranteed method for enhancing cancer classification accuracy.
- The performance of classification models is highly dependent on the representational power of the selected features.
- Further research into feature selection and optimized integration methods is warranted for improved breast cancer patient stratification.
Related Concept Videos
DNA Microarrays
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
Kaplan-Meier Approach
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...