米特征本体学的策划系统,通过LLM和PubAnnotation可靠的互操作
Javeed Muhammad Ahmad1, Yawen Liu1, Jin-Dong Kim2
1College of Informatics, Hubei Key Laboratory of Agricultural Bioinformatics, Huazhong Agricultural University, Wuhan, China.
Genomics & informatics
|December 2, 2025
概括
这项研究引入了一种使用大型语言模型 (LLM) 和数据库的自动化系统,用于策划米特征本体学,显著提高了农业基因组学手动方法的效率和数据准确性.
科学领域:
- 农业基因组学 农业基因组学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 本体学框架对于组织复杂的生物数据,如基因和表型至关重要.
- 用于本体论注释的手册文献审查是耗时且难以扩展的.
- 大型语言模型 (LLM) 和策划数据库为半自动化的本体学策划提供了潜力.
研究的目的:
- 为米特征本体学开发一个高效的基于文献的策划系统.
- 将LLMs与现有数据库集成为半自动化的本体学策划.
- 为了比较自动化系统与手动策划方法的性能.
主要方法:
- 开发了一个策划系统,整合了Rice-Alterome和PubAnnotation.
- 使用LLM (DeepSeek,KIMI) 通过API进行自动信息提取和注释.
- 采用快速工程来实现LLM集成,并通过使用案例评估系统.
主要成果:
- 与手动方法相比,该系统大大改善了文献证据的检索和组织.
- 综合平台简化了本体学策划和证据发现.
- 在LLM驱动的查询中,识别了在手动策划中忽略的隐含或缺失的信息.
结论:
- 将LLMs与Rice-Alterome和PubAnnotation集成,为自动化米特征本体学策划提供了一个有希望的解决方案.
- 这种方法加快了证据收集,提高了数据的一致性和可访问性.
- 未来的工作将把框架扩展到其他作物,如小麦和玉米.
更多相关视频
11:33Investigating Interactions Between Histone Modifying Enzymes and Transcription Factors in vivo by Fluorescence Resonance Energy Transfer
Published on: October 14, 2022
2.0K
09:43Author Spotlight: Streamlining Rice Breeding with CRISPR/Cas for Obtaining Optimal Phenotypic and Agronomic Traits
Published on: January 3, 2025
3.3K
相关概念视频
Light Acquisition
9.3K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
9.3K
Genome Annotation and Assembly
20.5K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
20.5K
Cis-regulatory Sequences
11.5K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
11.5K
Cis-regulatory Sequences
4.0K
4.0K
Plant Breeding and Biotechnology
21.4K
Crop cultivation has a long history in human civilization, with records showing the cultivation of cereal plants beginning at around 8000 BC. This early plant breeding was developed primarily to provide a steady supply of food.
21.4K
The Unfolded Protein Response
6.2K
The ER is the hub of protein synthesis in a cell. It has robust systems to quality control protein folding and also for degradation of terminally misfolded proteins. Under normal conditions, a small proportion of misfolded proteins that cannot be salvaged need to be transported to the cytoplasm by the ER-associated degradation or ERAD pathways. However, if the ERAD cannot handle the misfolded proteins, the cell activates the unfolded protein response or UPR to adjust the protein folding...
6.2K
