miRBench:用于预测微RNA结合位点的新基准数据集,可以缓解普遍存在的微RNA频率类偏差
Stephanie Sammut1,2, Katarina Gresova1,2,3, Dimosthenis Tzimotoudis1,2
1Centre for Molecular Medicine and Biobanking, University of Malta, Msida, MSD 2080, Malta.
Bioinformatics (Oxford, England)
|July 15, 2025
概括
这项研究引入了一种新的方法,用于产生微RNA (miRNA) 目标预测的公正数据集,从而提高模型的准确性. 一个新的Python包,miRBench,便于访问这些数据集和模型.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 基因组学就是基因组学.
背景情况:
- 微RNAs (miRNAs) 调节基因表达,但它们的向结合机制尚未完全理解.
- 对于miRNA目标预测的现有数据集缺乏公正的负面示例,阻碍了准确的模型开发.
- 负样本的in silico生成可以引入偏差,例如miRNA频率类偏差,影响模型概括.
研究的目的:
- 开发一种新的方法,用于为miRNA目标预测数据集生成无偏见的负样本.
- 使用开发的方法来策划新的,广泛的数据集.
- 在这些精心策划的数据集上对现有最先进的方法进行比较,并提供一个用户友好的Python包.
主要方法:
- 为减轻miRNA频率类偏差,开发了一种新的负样本生成方法.
- 几个新的,广泛的数据集使用开发的方法进行了策划.
- 最先进的预测方法在新策划的数据集上进行了基准测试.
主要成果:
- 这种新的方法有效地减轻了miRNA频率类偏差.
- 一个简单的卷积神经网络在无偏见的数据集上进行了重新训练,超过了现有的最先进的方法,达到0.81-0.86.8的平均精度得分.
- 该miRBench Python包是为了便于访问数据集,序列编码和模型执行而开发的.
结论:
- 公正的数据集对于提高miRNA结合位预测模型的准确性至关重要.
- 开发的方法和精心策划的数据集为研究界提供了宝贵的资源.
- 该miRBench包降低了机器学习研究人员进入miRNA目标预测领域的障碍.
相关概念视频
MicroRNAs
3.1K
MicroRNA (miRNA) are short, regulatory RNA transcribed from introns (non-coding regions of a gene) or intergenic regions (stretches of DNA present between genes). Several processing steps are required to form biologically active, mature miRNA. The initial transcript, called primary miRNA (pri-mRNA), base-pairs with itself, forming a stem-loop structure. Within the nucleus, an endonuclease enzyme, called Drosha, shortens the stem-loop structure into hairpin-shaped pre-miRNA. After the pre-miRNA...
3.1K
Conserved Binding Sites
4.4K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.4K


