使用积极标记和未标记实例进行疾病候选基因预测
Sepideh Molaei1, Saeed Jalili2
1Computer Engineering Department, Tarbiat Modares University, Tehran, Iran.
BMC medical genomics
|April 16, 2025
概括
这项研究引入了一种新的计算方法,用于使用机器学习识别疾病基因. 它有效地产生可靠的负数据集,改善了药物发现候选疾病基因的预测和排名.
科学领域:
- 计算生物学是一种计算生物学.
- 遗传学 遗传学 是一个
- 机器学习在生物信息学中的应用.
背景情况:
- 鉴定与遗传疾病相关的基因对于药物开发至关重要.
- 传统的基因鉴定实验室方法耗时且昂贵.
- 计算方法,特别是机器学习,为疾病基因识别提供了有希望的替代方案.
研究的目的:
- 开发一种新的计算方法,用于预测和排名候选疾病基因.
- 解决缺少负数据集的挑战,用于疾病基因识别的机器学习.
- 提高遗传疾病疾病疾病基因发现的准确性和效率.
主要方法:
- 提出了一种两步的机器学习方法.
- 步骤1:一级学习和基于距离的过,从未标记的基因组中提取可靠的负基因,用于每个疾病.
- 步骤2:对已知疾病基因 (正组) 和提取的负基因进行训练的二进制分类模型,然后进行基因评分和排名.
主要成果:
- 提出的方法成功预测和排名候选疾病基因.
- 对六种疾病和一类癌症的评估表明,与现有研究相比,其表现优越.
- 该方法有效地产生了可靠的负数据集,这是以前方法面临的重大挑战.
结论:
- 开发的计算方法为疾病基因识别提供了有效的解决方案,特别是通过克服负数据集限制.
- 这种方法提高了候选疾病基因的预测和排名,促进了对遗传疾病的更有效的药物发现.
- 这些发现表明,在将机器学习应用于遗传疾病研究方面取得了重大进展.
相关概念视频
Genome-wide Association Studies-GWAS
12.1K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
12.1K
lncRNA - Long Non-coding RNAs
8.4K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.4K


