使用监督机器学习算法进行乳腺癌数据分析
Durga H Kutal1, Beyza N Koseoglu1
1Mathematics, Augusta University, Augusta, USA.
Cureus
|November 24, 2025
概括
机器学习模型准确地对乳腺癌瘤进行分类. 随机森林和多项式SVM显示出最佳表现,突出显示瘤大小和淋巴结参与作为关键指标.
科学领域:
- 在瘤学瘤学.
- 计算机科学 计算机科学
- 生物统计学 生物统计学
背景情况:
- 乳腺癌是全球妇女死亡的主要原因.
- 准确的瘤分类对于有效的治疗和患者的治疗结果至关重要.
研究的目的:
- 评估和比较乳腺癌瘤分类的监督机器学习算法.
- 通过使用现实世界的数据集来识别最有效的算法和重要的预测特征.
主要方法:
- 利用现实世界乳腺癌数据集 (205个观察).
- 采用了后勤回归,决策树,随机森林和具有各种内核的支持向量机 (SVM).
- 使用准确性,特征重要性和曲线下的面积 (AUC) 评估模型性能.
- 通过主要组件分析 (PCA) 调查了缩小维度的影响.
主要成果:
- 所有评估的模型都实现了超过87%的准确性.
- 随机森林和多项式SVM表现出优异的性能,其AUC分别为96.3%和96.9%.
- 确定的主要预测因素包括瘤大小,参与的淋巴结,转移和年龄.
- 在保持高模型性能的同时,PCA有效地降低了维度.
结论:
- 监督机器学习算法,特别是随机森林和多项式SVM,对于乳腺癌瘤分类非常有效.
- 瘤大小,淋巴结参与,转移和年龄是乳腺癌预测的关键因素.
- 像PCA这样的尺寸缩小技术可以在不影响分类准确性的情况下被有效地应用.
更多相关视频
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
475
07:41Performing Data Mining And Integrative Analysis Of Biomarker in Breast Cancer Using Multiple Publicly Accessible Databases
Published on: May 17, 2019
9.4K
相关概念视频
Cancer Survival Analysis
633
Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
633
Comparing the Survival Analysis of Two or More Groups
541
Survival analysis is a cornerstone of medical research, used to evaluate the time until an event of interest occurs, such as death, disease recurrence, or recovery. Unlike standard statistical methods, survival analysis is particularly adept at handling censored data—instances where the event has not occurred for some participants by the end of the study or remains unobserved. To address these unique challenges, specialized techniques like the Kaplan-Meier estimator, log-rank test, and...
541
Statistical Methods for Analyzing Epidemiological Data
885
Epidemiological data primarily involves information on specific populations' occurrence, distribution, and determinants of health and diseases. This data is crucial for understanding disease patterns and impacts, aiding public health decision-making and disease prevention strategies. The analysis of epidemiological data employs various statistical methods to interpret health-related data effectively. Here are some commonly used methods:
885
Kaplan-Meier Approach
541
The Kaplan-Meier estimator is a non-parametric method used to estimate the survival function from time-to-event data. In medical research, it is frequently employed to measure the proportion of patients surviving for a certain period after treatment. This estimator is fundamental in analyzing time-to-event data, making it indispensable in clinical trials, epidemiological studies, and reliability engineering. By estimating survival probabilities, researchers can evaluate treatment effectiveness,...
541
