从文本到洞察力:化学数据提取的大型语言模型.
Mara Schilling-Wilhelmi1, Martiño Ríos-García1,2, Sherjeel Shabih3
1Laboratory of Organic and Macromolecular Chemistry (IOMC), Friedrich Schiller University Jena, Humboldtstrasse 10, 07743 Jena, Germany. mail@kjablonka.com.
大型语言模型 (LLM) 现在可以有效地从非结构化化学文本中提取结构化数据,帮助材料设计. 将LLM成果与领域专业知识相结合,是可靠,数据驱动的化学研究的关键.
科学领域:
- 化学和材料科学 化学和材料科学
- 计算化学计算化学
- 数据科学数据科学数据科学
背景情况:
- 庞大的化学知识以非结构化的文本形式存在,阻碍了系统的材料设计.
- 传统的数据提取方法是手动或部分自动化的,限制了效率.
- 大型语言模型 (LLM) 为访问结构化化学信息提供了一个范式转变.
研究的目的:
- 为提供基于LLM的化学结构化数据提取的全面概述.
- 综合当前的知识,并概述化学数据LLM应用的未来方向.
- 为了解决缺乏用于化学LLM使用的标准化指导方针.
主要方法:
- 审查目前关于化学数据提取的LLM的文献.
- 开发用于将LLM能力与领域专业知识相结合的框架.
- 合成化学和材料科学中的LLM应用和挑战.
主要成果:
- 通过LLM,可以显著提高从非结构化化学文本中提取结构化数据的效率.
- 领域知识对于在科学背景下指导和验证LLM成果至关重要.
- 法律法学和化学专业知识之间的协同方法增强了数据驱动的研究.
结论:
- 基于LLM的数据提取具有加速开发新型化合物和材料的巨大潜力.
- 为了在化学中有效实施LLM,需要标准化的指导方针和验证的框架.
- 这项工作是研究人员利用化学发现中的LLMs的基础资源.
更多相关视频
09:04Identifying Per- and Polyfluorinated Chemical Species with a Combined Targeted and Non-Targeted-Screening High-Resolution Mass Spectrometry Workflow
Published on: April 18, 2019
11:00Untargeted Metabolomics from Biological Sources Using Ultraperformance Liquid Chromatography-High Resolution Mass Spectrometry UPLC-HRMS
Published on: May 20, 2013
相关概念视频
Molecular Models
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Spectroscopy of Carboxylic Acid Derivatives
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...
Extraction: Advanced Methods
Drug Discovery: Overview
