深度学习有助于发现具有超过10万亿个序列的自组合
Jiaqi Wang1,2, Zihan Liu3, Shuang Zhao1,2
1Research Center for Industries of the Future, Westlake University, Hangzhou, 310030, China.
Advanced science (Weinheim, Baden-Wurttemberg, Germany)
|September 26, 2023
概括
一个新的深度学习模型准确地预测了的自我组装特性,使得新型自我组装的高效设计能够用于各种生物和医学应用.
科学领域:
- 生物化学和分子生物学
- 计算生物学 计算生物学
- 材料科学 材料科学 材料科学
背景情况:
- 的自我组装对于生物和医学应用至关重要.
- 调查的庞大的序列空间进行自我组装是计算上具有挑战性的.
研究的目的:
- 开发和验证一个深度学习模型来预测聚 propensity (AP).
- 探索系统中的聚合规律和可转移性关系.
- 通过计算预测发现新的自我组装.
主要方法:
- 利用基于变压器的深度学习模型来预测系统的AP.
- 分析了超过10万亿个可能的序列为dekapeptides和混合pentapeptides.
- 实验验证了预测的自我组装.
主要成果:
- 深度学习模型在广泛的序空间中准确地预测了AP.
- 衍生了聚合规律,并揭示了五,十和混合系统之间的可转移性.
- 成功发现并通过实验证实了新的自我组装.
结论:
- 深度学习为快速准确设计自组装提供了一种强大的方法.
- 这种方法显著推进了科学,并为生物和医学创新开辟了新的途径.
- 能够全面探索用于自组装应用的寡序列空间.
相关概念视频
Peptide Identification Using Tandem Mass Spectrometry
6.5K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.5K
Protein Folding
8.1K
Proteins are chains of amino acids linked together by peptide bonds. Upon synthesis, a protein folds into a three-dimensional conformation, critical to its biological function. Interactions between its constituent amino acids guide protein folding, and hence the protein structure is primarily dependent on its amino acid sequence.
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
Protein Structure Is Critical to Its Biological Function
Proteins perform a wide range of biological functions such as catalyzing chemical reactions, providing...
8.1K
Protein Complex Assembly
10.6K
Proteins can form homomeric complexes with another unit of the same protein or heteromeric complexes with different types. Most protein complexes self-assemble spontaneously via ordered pathways, while some proteins need assembly factors that guide their proper assembly. Despite the crowded intracellular environment, proteins usually interact with their correct partners and form functional complexes.
Many viruses self-assemble into a fully functional unit using the infected host cell to...
Many viruses self-assemble into a fully functional unit using the infected host cell to...
10.6K
Maxam-Gilbert Sequencing
11.2K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.2K
Signal Sequences and Sorting Receptors
5.4K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.4K
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K


