Related Experiment Video
Updated: Jun 21, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Fine-tuning large language models for chemical text mining
Wei Zhang1,2, Qinggong Wang3, Xiangtai Kong1,2
1Drug Discovery and Design Center, State Key Laboratory of Drug Research, Shanghai Institute of Materia Medica, Chinese Academy of Sciences 555 Zuchongzhi Road Shanghai 201203 China myzheng@simm.ac.cn fuzunyun@simm.ac.cn.
Fine-tuned large language models (LLMs) significantly improve chemical text mining accuracy across five complex tasks. These advanced LLMs reduce the need for extensive prompt engineering, offering a powerful new tool for automated data acquisition in chemistry.
Area of Science:
- Computational chemistry
- Natural Language Processing
- Chemical Informatics
Background:
- Extracting knowledge from chemical literature is challenging due to complex language.
- Automated data acquisition is crucial for both experimental and computational chemists.
- Large Language Models (LLMs) show potential for chemical text mining.
Purpose of the Study:
- To explore the effectiveness of fine-tuned LLMs on intricate chemical text mining tasks.
- To compare fine-tuned LLMs against prompt-engineered models (ChatGPT, GPT-4) and other open-source LLMs.
- To assess the performance of LLMs with minimal annotated data.
Main Methods:
- Fine-tuning of various LLMs including ChatGPT (GPT-3.5-turbo), GPT-4, Mistral, Llama3, Llama2, T5, and BART.
- Evaluation on five chemical text mining tasks: compound entity recognition, reaction role labelling, MOF synthesis extraction, NMR data extraction, and reaction-to-action sequence conversion.
- Comparison of fine-tuned models against prompt-engineered models using limited annotated data.
Main Results:
- Fine-tuned ChatGPT models achieved high accuracy (69%-95%) across all evaluated tasks.
- Fine-tuned LLMs outperformed models using task-adaptive pre-training with larger in-domain datasets.
- Fine-tuned Mistral and Llama3 demonstrated competitive performance.
- Significant reduction in prompt engineering effort was observed.
Conclusions:
- Fine-tuned LLMs are highly effective for chemical knowledge extraction, even with minimal data.
- These models offer a versatile, robust, and low-code solution for automated data acquisition.
- Leveraging fine-tuned LLMs can revolutionize the field of chemical informatics.
Related Concept Videos
Molecular Models
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Chemical Shift: Internal References and Solvent Effects
The internal reference compound generally used in NMR spectroscopy is tetramethylsilane (TMS). TMS is preferred because it is chemically inert, soluble in NMR solvents, and easily removable. Also, the highly shielded methyl protons in TMS yield an intense...
Improving Translational Accuracy

