Fine-tuning large language models for chemical text mining

Wei Zhang1,2, Qinggong Wang3, Xiangtai Kong1,2

  • 1Drug Discovery and Design Center, State Key Laboratory of Drug Research, Shanghai Institute of Materia Medica, Chinese Academy of Sciences 555 Zuchongzhi Road Shanghai 201203 China myzheng@simm.ac.cn fuzunyun@simm.ac.cn.

Chemical Science
|July 12, 2024
PubMed
Summary

Fine-tuned large language models (LLMs) significantly improve chemical text mining accuracy across five complex tasks. These advanced LLMs reduce the need for extensive prompt engineering, offering a powerful new tool for automated data acquisition in chemistry.

Related Concept Videos

Molecular Models02:00

Molecular Models

Physical models representing molecular architectures of chemical compounds play essential roles in understanding chemistry. The use of molecular models makes it easier to visualize the structures and shapes of atoms and molecules.
38.2K
Ligand Binding Sites02:40

Ligand Binding Sites

Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
12.8K
Chemical Shift: Internal References and Solvent Effects01:17

Chemical Shift: Internal References and Solvent Effects

In an NMR sample, precise measurement of the absolute absorption frequencies of nuclei is difficult. A standard internal reference compound is added, and the frequency difference between the reference signal and sample signals is measured.
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...
628
Improving Translational Accuracy02:07

Improving Translational Accuracy

Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.9K