Related Experiment Video
Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models on multimodal chemistry olympiad exams
Yiming Cui1,2, Xin Yao3,4, Yuxuan Qin3,4
1State Key Laboratory of Cognitive Intelligence, Hefei, China. ymcui@ieee.org.
Multimodal large language models (LLMs) struggle with chemistry reasoning, often failing to integrate visual and textual data effectively. Chain-of-Thought prompting improves performance, highlighting a need for better multimodal AI in scientific domains.
Area of Science:
- Artificial Intelligence
- Chemistry Education
Background:
- Multimodal large language models (LLMs) face challenges in scientific reasoning, especially in chemistry, which involves complex visual and symbolic data.
- Effective integration of visual and textual information is crucial for advanced problem-solving in chemistry.
Purpose of the Study:
- To systematically evaluate the performance of 40 multimodal LLMs on chemistry-specific reasoning tasks.
- To identify limitations in current models' ability to fuse information from diverse modalities.
- To assess the impact of Chain-of-Thought prompting on multimodal scientific reasoning.
Main Methods:
- Curated benchmark of U.S. National Chemistry Olympiad (USNCO) questions spanning over two decades.
- Systematic evaluation of 40 proprietary and open-source multimodal LLMs.
- Ablation studies and occlusion-based interpretability to analyze model behavior and visual grounding.
- Comparison of performance with and without Chain-of-Thought prompting.
Main Results:
- Many evaluated multimodal LLMs exhibit significant difficulties with modality fusion, sometimes performing worse with visual input.
- Chain-of-Thought prompting demonstrably improves both accuracy and visual grounding in multimodal chemistry problem-solving.
- Performance variations highlight critical limitations in current multimodal LLMs' scientific reasoning capabilities.
Conclusions:
- Current multimodal LLMs require substantial improvements for robust scientific reasoning in chemistry.
- Chain-of-Thought prompting offers a promising strategy for enhancing multimodal AI performance and interpretability.
- The developed benchmark serves as a critical tool for measuring progress in domain-specific multimodal AI research.
More Related Videos
Related Concept Videos
Molecular Models
The Small x Assumption
Chemical and Solubility Equilibria
Predicting Molecular Geometry
Classification of Elements and Compounds
Compounds are pure substances composed of two or more elements in fixed, definite proportions. Compounds are classified as ionic or molecular (covalent) based on the bonds...
Experimental Determination of Chemical Formula

