Related Experiment Video
Updated: Aug 14, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
MolPACL: Molecular Property Prediction Based on Prompt Augmentation and Contrastive Learning
IEEE Transactions on Computational Biology and Bioinformatics
|August 12, 2026
Summary
This study introduces MolPACL, a novel framework for learning molecular representations using large language models (LLMs) and prompt augmentation. It enhances chemical understanding while preserving molecular structure, outperforming existing methods.
Area of Science:
- Computational chemistry
- Machine learning
- Artificial intelligence
Background:
- Large language models (LLMs) offer potential for molecular representation learning from text.
- Current contrastive learning methods often use graph or SMILES augmentations that can distort chemical structures.
- There is a need for methods that leverage high-level chemical semantics while maintaining molecular integrity.
Purpose of the Study:
- To develop MolPACL, a prompt-augmentation-based supervised contrastive learning framework.
- To incorporate high-level chemical semantics into molecular representations.
- To preserve molecular identity during the learning process.
Main Methods:
- MolPACL generates semantically consistent prompt views from SMILES strings and physicochemical descriptors.
- It utilizes diverse templates and lightweight lexical perturbations for prompt generation.
- A supervised objective based on Soft Nearest Neighbor loss is employed for training.
Main Results:
- MolPACL achieves strong performance on MoleculeNet benchmarks for classification and regression tasks.
- The framework reduces training costs.
- It requires no additional molecular pretraining and uses a relatively small pretrained LLM.
Conclusions:
- MolPACL effectively learns molecular representations by incorporating high-level chemical semantics.
- The prompt-augmentation approach preserves molecular identity.
- This method offers an efficient and effective alternative for molecular representation learning.