SeqImprove:基因电路序列信息的机器学习辅助策划.
Jeanet Mante1, Zach Sents1, Duncan Britt1
1University of Colorado Boulder, Boulder, Colorado 80309, United States.
ACS synthetic biology
|September 4, 2024
概括
合成生物学的进步被困难的文献审查和糟糕的文档所阻碍. 机器学习工具SeqImprove帮助作者创建FAIR数据,提高可复制性.
科学领域:
- 合成生物学 合成生物学
- 生物信息学是一种生物信息学.
- 数据策划数据策划
背景情况:
- 合成生物学进步受到文献审查和复制文献工作的挑战的限制.
- 手动重建设计信息容易出现错误和噪音.
- 作者参与数据策划对于准确性至关重要.
研究的目的:
- 开发一个机器学习辅助工具,SeqImprove,以简化合成生物学数据的策划过程.
- 鼓励作者参与策划,而不增加他们的工作量.
- 提高生物序列数据的可查,可访问性,互操作性和可重复使用性 (FAIR).
主要方法:
- SeqImprove使用命名实体识别 (实体规范化) 和序列匹配.
- 该工具生成机器可访问的序列数据和元数据注释.
- 作者在最终提交之前审查和编辑生成的注释.
主要成果:
- SeqImprove 简化了机器可读的序列数据和元数据的创建.
- 该工具简化了作者提交FAIR数据的过程.
- 减少与事后数据重建相关的噪音和错误.
结论:
- SeqImprove提高了合成生物学数据处理的效率和准确性.
- 该工具促进了FAIR数据的提交,加速了研究和开发.
- ML辅助疗法是克服合成生物学当前局限性的可行策略.
相关概念视频
Genetic Screens
4.9K
Genetic screens are tools used to identify genes and mutations responsible for phenotypes of interest. Genetic screens help identify individuals or a group of people at risk of developing genetic diseases and help them with early intervention, targeted therapy, and reproductive options.
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
Forward genetic screens
Forward or “classical” genetic screens involve creating random mutations in an organism’s DNA using radiation, mutagens, or insertion of additional bases, which...
4.9K
Cis-regulatory Sequences
9.8K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
9.8K
Maxam-Gilbert Sequencing
11.1K
In the same year as the discovery of the Sanger sequencing method, another group of scientists, Allan Maxam and Walter Gilbert, demonstrated their chemical-cleavage method for DNA sequencing. The Maxam-Gilbert method relies on using different chemicals that can cleave the DNA sequence at specific sites, the separation of resulting DNA fragments of variable size using electrophoresis, and deciphering the DNA sequence from the resulting gel bands.
Challenges of the Maxam-Gilbert Method
The...
Challenges of the Maxam-Gilbert Method
The...
11.1K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Next-generation Sequencing
88.5K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
88.5K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K


