TaeC: 一个手动注释的文本数据集,用于特征和表型提取,以及在小麦育种文献中的实体链接
Claire Nédellec1, Clara Sauvion1, Robert Bossy1
1Université Paris-Saclay, INRAE, MaIAGE, Jouy-en-Josas, France.
PloS one
|June 13, 2024
概括
一个新的小麦特征集团通过标准化科学文献中的特征和表型数据来帮助植物科学家. 该资源改善了信息提取,以提高小麦品种的效率.
科学领域:
- 农业科学 农业科学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 小麦育种依赖于理解基因-表型联系,但文献中的各种术语阻碍了数据整合.
- 现有的植物表型体很少,限制了文本挖掘和机器学习用于作物知识发现.
研究的目的:
- 为了介绍Triticum aestivum特征库,小麦特征和表型的新,手动注释的资源.
- 为应对在小麦研究中交叉引用各种文本和测试数据的挑战.
主要方法:
- 从528个PubMed引用中构建一个语料库,对特征,表型和物种进行注释.
- 利用小麦特征和表型本体学和NCBI物种分类学来规范化数据.
- 对命名实体识别和链接任务的最新语言模型的评估.
主要成果:
- "Triticum aestivum特征库"是关于植物表型NLP的最全面的手动注释资源.
- 该集体证明了其适用于培训和评估自然语言处理模型的适用性.
- 规范化的特征和物种数据增强了稀疏测试数据和出版物之间的互操作性.
结论:
- 该库是小麦特征和表型研究的黄金标准,促进了先进的文本挖掘.
- 它允许改进机器学习模型培训,以便在作物科学中更准确地提取信息.
- 该资源通过更好的基因-表型相互作用发现,支持更有效的小麦育种计划.
更多相关视频
08:36Development of Targeting Induced Local Lesions IN Genomes TILLING Populations in Small Grain Crops by Ethyl Methanesulfonate Mutagenesis
Published on: July 16, 2019
11.6K
07:18Obtaining High-Quality Transcriptome Data from Cereal Seeds by a Modified Method for Gene Expression Profiling
Published on: May 21, 2020
7.4K
相关概念视频
Light Acquisition
8.4K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
8.4K
Plant Breeding and Biotechnology
18.9K
Crop cultivation has a long history in human civilization, with records showing the cultivation of cereal plants beginning at around 8000 BC. This early plant breeding was developed primarily to provide a steady supply of food.
18.9K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Pedigree Analysis
84.2K
Overview
84.2K
Genome-wide Association Studies-GWAS
13.3K
Genome-wide association studies or GWAS are used to identify whether common SNPs are associated with certain diseases. Suppose specific SNPs are more frequently observed in individuals with a particular disease than those without the disease. In that case, those SNPs are said to be associated with the disease. Chi-square analysis is performed to check the probability of the allele likely to be associated with the disease.
GWAS does not require the identification of the target gene involved in...
GWAS does not require the identification of the target gene involved in...
13.3K
Incomplete Dominance
22.5K
Gregor Mendel's work (1822 - 1884) was primarily focused on pea plants. Through his initial experiments, he determined that every gene in a diploid cell has two variants called alleles inherited from each parent. He suggested that amongst these two alleles, one allele is dominant in character and the other recessive. The combination of alleles determines the phenotype of a gene in an organism.
22.5K
