使用基于蜜蜂的算法对基因表达数据进行缺失值赋值,以提高分类性能
Kritanat Chungnoy1, Tanatorn Tanantong1,2, Pokpong Songmuang1,2
1Department of Computer Science, Faculty of Science and Technology, Thammasat University (Rangsit Campus), Pathum Thani, Thailand.
PloS one
|August 29, 2024
概括
本研究引入了一种新的缺失值归算方法,使用蜜蜂算法和k-最近邻居来提高分类准确性. 这种新方法显著提高了分类性能,超过了现有的技术.
科学领域:
- 计算机科学 计算机科学
- 机器学习 机器学习
- 数据科学数据科学数据科学
背景情况:
- 传统的缺失值归算方法侧重于用于机器学习的数据集完成.
- 这些方法往往旨在复制原始数据值,可能会限制下游任务性能.
研究的目的:
- 提出一种新的缺失值归算方法,专门设计用于提高分类准确性.
- 增强数据集的分辨能力,用于分类任务.
主要方法:
- 一种混合方法,将蜜蜂算法与k-最近邻居和线性回归结合起来用于归算.
- 在归算过程中使用GINI重要性得分来选择特征.
- 评估该方法与已建立的技术相比,如k-最近邻居,主要组件分析和非线性主要组件分析.
主要成果:
- 拟议的归算方法在所有测试数据集的分类任务中实现了卓越的准确性.
- 使用拟议方法计算的数据集,与原始数据集相比,分类准确度增加了15-25%.
- 归算策略明显改善了特征信息性和分类的歧视力.
结论:
- 新的归算方法通过提高数据的区分能力,有效地提高了分类准确性,而不仅仅是通过填补缺失的值.
- 这种方法比现有的归算技术提供了显著的进步,用于以分类为重点的机器学习应用程序.
相关概念视频
Genomic Imprinting and Inheritance
34.2K
Diploid organisms inherit genetic material through chromosomes from both parents. Copies of the same gene are known as alleles. In most cases, both alleles are simultaneously expressed and allow various cellular processes to function optimally. If one of the alleles is missing or mutated, the expression of the other allele can compensate; however, this is not true for all genes.
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
The expression of some genes depends on which parent passed the gene to the offspring, through a phenomenon known as...
34.2K
DNA Microarrays
17.2K
Microarrays are high-throughput and relatively inexpensive assays that can be automated to analyze large quantities of data at a time. They are used in genome-wide studies to compare gene or protein expression under two varied conditions, such as healthy and diseased states. Microarrays consist of glass or silica slides on which probe molecules are covalently attached through surface functionalization. Most commonly, the slides are prepared through the chemisorption of silanes to silica...
17.2K


