布朗诺特是一种全面的解决方案,可以为任何物种生成蛋白质序列数据库
Adrien Brown1,2, Alexandre Burel1,2, Sarah Cianférani1,2
1Laboratoire de Spectrométrie de Masse Bioorganique (LSMBO), IPHC, UMR7178, Université de Strasbourg, CNS, Strasbourg, France.
Proteomics
|January 6, 2026
概括
一个新的管道,Brownotate,生成任何物种的蛋白质序列数据库,克服了蛋白质组学研究中的一个主要瓶. 这种用户友好的工具使研究人员能够有效地分析代表性不足的物种的蛋白质组数据.
科学领域:
- 蛋白质组学是指蛋白质组学.
- 生物信息学是一种生物信息学.
- 基因组学就是基因组学.
背景情况:
- 蛋白组学研究需要全面的蛋白序列数据库.
- 对许多物种缺乏数据库,阻碍了生物复杂性研究.
- 现有的计算工具往往是有限的或需要专门的专业知识.
研究的目的:
- 开发一个可访问的管道,用于生成任何物种的蛋白质序列数据库.
- 让没有生物信息学专业知识的研究人员能够创建必要的数据库.
- 解决生物研究中缺少蛋白质序列数据的瓶.
主要方法:
- 开发了Brownotate,这是一个开源的,用户友好的基因组组装和蛋白质注释管道.
- 该管道提取现有的蛋白质序列,并注释基因组组合或DNA数据集.
- 评估了跨不同物种和基因组的管道性能.
主要成果:
- 布朗诺特生成碎片化但优质的汇编和注释.
- 蛋白质预测重叠率很高,可与参考数据库进行比较.
- 使用Brownotate生成的数据库来解释蛋白质组数据,结果与NCBI数据库类似.
结论:
- 布朗诺酸有效地缓解了对未经研究的物种缺乏蛋白质序列数据库的缺陷.
- 该管道使得可用测序数据的物种能够进行蛋白质组分析.
- 棕酸是蛋白质组学研究工具箱的宝贵补充.
相关概念视频
Multi-species Conserved Sequences
4.6K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
4.6K
Evolutionary Relationships through Genome Comparisons
6.8K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.8K
Protein Families
16.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.6K
Conservation of Protein Domains Over Different Proteins
14.0K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
14.0K
Protein Networks
4.5K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.5K
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K


