Related Experiment Video
Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Performance of Large Language Models for One Atmosphere Using Automated Extracted Datasets
Shiqin Dai1,2, Qingru Wu1,2, Haowen Zhang1,2
1School of Environment, State Key Laboratory of Regional Environment and Sustainability, School of Environment, Tsinghua University, Beijing 100084, P. R. China.
Abstract:
Large language models (LLMs) have great potential to improve the efficiency of atmospheric science research and air pollution control decision-making. However, due to the absence of a systematic evaluation benchmark, it remains unclear whether LLMs can be trusted to support core air pollution control tasks such as pollution alarming and mitigation recommendation, which limits their real-world deployment. This study developed a benchmark dataset aligned with the One Atmosphere framework (OneAtmos-Bench), by using a two-stage automated extraction pipeline. Leveraging this dataset, 11 LLMs were evaluated across three key metrics: accuracy, instruction-following (IF), and hallucination rate (HR). Expert validation confirms that the OneAtmos-Bench dataset achieves 95.88% accuracy. The evaluation results indicate that while IF remains consistently strong, both accuracy and hallucination suppression are still constrained, especially for long-tail tasks in the atmospheric domain. Notably, the marginal improvements from scaling up the model size for atmospheric tasks highlight sparse domain adaptation, and common training strategies to enhance reasoning ability, particularly distillation-based supervised fine-tuning, may inadvertently increase the HR. These findings reveal that general-purpose LLMs present reliability challenges in delivering the trustworthy, low-hallucination guidance required for efficient air pollution control decision-making, highlighting the need for environment-specific adaptation.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
