概括
我们介绍了GERBERA,这是一种使用通用域数据集来训练生物医学命名实体识别 (BioNER) 模型的新方法. 这种方法提高了BioNER的性能,特别是在有限的生物医学数据下,降低了注释成本.
科学领域:
- 计算生物学 计算生物学
- 自然语言处理自然语言处理.
背景情况:
- 生物医学命名实体识别 (BioNER) 模型培训需要广泛,昂贵的人类注释.
- 现有的多任务学习方法与多个BioNER数据集显示不一致的性能增长和潜在的标签模两可.
研究的目的:
- 开发一种具有成本效益的BioNER培训方法,使用从通用域数据集的转移学习.
- 改进BioNER模型的性能,特别是在低资源场景中.
主要方法:
- 拟议的GERBERA:一种利用通用NER数据集进行培训的方法.
- 使用预先训练的生物医学语言模型进行多任务学习,将目标BioNER和通用域数据集结合起来.
- 专门针对目标BioNER数据集的微调模型.
主要成果:
- 与使用额外的BioNER数据集训练的基线模型相比,GERBERA模型显示出更高的性能.
- 在八个实体类型中实现了0.9%的平均改善,在六个实体类型中表现优于基线.
- 在数据有限的BioNER数据集上显著提高了性能,在JNLPBA-RNA上F1得分增加了4.7%.
结论:
- 利用具有成本效益的通用NER数据集可以有效地增强BioNER模型.
- 在缺乏或昂贵的生物医学注释资源的情况下,GERBERA方法为场景提供了有价值的解决方案.
- 这种方法提高了BioNER模型的性能,并减少了对广泛的手工策划的依赖.
相关概念视频
Drug Nomenclature
1.7K
During the development of a new pharmaceutical, the manufacturer initially assigns a code name to the drug. Once approved, the drug receives a United States Adopted Name (USAN)—a generic, nonproprietary designation. Upon being listed in the United States Pharmacopeia, this nonproprietary name becomes the drug's official name. Additionally, the manufacturer assigns a proprietary name or trademark, which serves as the brand name under which the drug is marketed. It is worth noting that...
1.7K
ER Retrieval Pathway
3.8K
In the secretory pathway, vesicles transport proteins from one cellular compartment to another in forward transport to deliver the protein to its correct location. Occasionally, misfolded proteins and incorrect proteins escape their original compartments, and a retrieval pathway is used to return the escaped proteins to their original compartment.
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
3.8K
Genomics
36.2K
Genomics is the science of genomes: it is the study of all the genetic material of an organism. In humans, the genome consists of information carried in 23 pairs of chromosomes in the nucleus, as well as mitochondrial DNA. In genomics, both coding and non-coding DNA is sequenced and analyzed. Genomics allows a better understanding of all living things, their evolution, and their diversity. It has a myriad of uses: for example, to build phylogenetic trees, to improve productivity and...
36.2K
Transducer Mechanism: Enzyme-Linked Receptors
2.4K
Enzyme-linked receptors are cell-surface receptors acting as an enzyme or associating with an enzyme intracellularly. They make excellent drug targets. Drugs can bind to the extracellular ligand-binding domain or directly affect their enzymatic domain and alter their activity.
Major types that are helpful drug targets include:
Major types that are helpful drug targets include:
2.4K
Genome Annotation and Assembly
18.8K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.8K
Tagging and Fusion Proteins
6.6K
Proteins are involved in several cellular processes and biochemical reactions. Analyzing a specific protein of interest requires it to be isolated from the other proteins in the cell. This is achieved by overexpressing the specific gene in a suitable host to produce large quantities of the target protein. A tag or label is recombined with the gene to produce a fusion protein containing the target protein and the tag. The tags on these fusion proteins can then be used for easy detection and...
6.6K


