rPIMS:使用基因组数据和机器学习方法精确识别和建模牲畜品种的ShinyR包
Yuhetian Zhao1, Xuexue Liu2, Benmeng Liang3
1State Key Laboratory of Animal Biotech Breeding, Institute of Animal Sciences, Chinese Academy of Agricultural Sciences (CAAS), Beijing 100193, People's Republic of China.
Bioinformatics advances
|May 7, 2025
概括
rPIMS使用基因组数据和机器学习简化了牲畜品种识别. 该工具在分类品种方面达到100%的准确性,使遗传分析可用于保护工作.
科学领域:
- 动物遗传学和动物繁殖
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 准确的牲畜品种识别对于保护遗传资源至关重要.
- 基因组数据和机器学习越来越多地用于品种识别.
- 现有的方法通常需要大量的数据和生物信息学专业知识.
研究的目的:
- 引入rPIMS,这是一个用户友好的工具,用于畜牧品种识别和遗传分析.
- 为研究人员简化复杂的模型构建过程.
- 加强对遗传多样性和进化关系的识别.
主要方法:
- 开发rPIMS,一个闪亮的R应用程序,包含模块用于数据输入,维度减小,家族遗传树构建,人口结构分析和机器学习分类.
- 利用基因组数据和机器学习算法进行品种分类.
- 在10个品种中使用860个单核酸多态 (SNP) 验证.
主要成果:
- rPIMS仅使用860个SNP,在区分10个品种时实现了100%的分类准确性.
- 该工具简化了分析过程,促进了协作和数据共享.
- rPIMS使品种分类和遗传结构可视化具有直观和可访问性.
结论:
- rPIMS显著简化了复杂的品种识别和遗传分析工作流程.
- 该工具增强了报告牲畜遗传多样性和进化关系的能力.
- rPIMS为更广泛的用户提供了先进的基因分析.
相关概念视频
Pedigree Analysis
Overview
Genomics
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
Pedigree Analysis
Overview
RACE - Rapid Amplification of cDNA Ends
Rapid Amplification of cDNA Ends, or RACE, is one of the most effective methods to obtain a full-length cDNA from an mRNA sequence between a known internal region to the unknown sequence at the 5’ or 3’ end. The unknown region is cloned in the cDNA by a gene-specific primer that binds the known end, and a hybrid primer that attaches a predefined anchor sequence to the unknown end of the cDNA. The sequence in between is amplified by PCR with an anchor primer and a gene-specific primer.
Since the...
Since the...
Protein Modifications in the RER
Modification of secretory and transmembrane proteins entering the rough ER begins in the ER lumen. These modifications aid in protein folding and stabilize the acquired tertiary structure. Protein modifications in the rough ER co-occur at different stages of protein folding.
Broadly, these modifications can be categorized into four main categories — glycosylation, formation of disulfide bonds, assembly of protein subunits, and specific proteolytic cleavages like removal of signal sequences.
Broadly, these modifications can be categorized into four main categories — glycosylation, formation of disulfide bonds, assembly of protein subunits, and specific proteolytic cleavages like removal of signal sequences.
Protein Folding Quality Check in the RER
ER is the primary site for the maturation and folding of soluble and transmembrane secretory proteins. The calnexin cycle is a specific chaperone system that folds and assesses the confirmation of N-glycosylated proteins before they can exit the ER lumen. The primary players of this quality check pipeline are the lectins, ER-resident chaperones, and a glucosyl transferase enzyme. In case the calnexin system in the lumen fails to salvage a misfolded protein, it is transported to the cytoplasm...


