Related Experiment Video
Updated: Jan 7, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
983
Evaluating the Performance of Large Language Models for One Atmosphere Using Automated Extracted Datasets
Shiqin Dai1,2, Qingru Wu1,2, Haowen Zhang1,2
1School of Environment, State Key Laboratory of Regional Environment and Sustainability, School of Environment, Tsinghua University, Beijing 100084, P. R. China.
Environmental Science & Technology
|December 31, 2025
Summary
Large language models (LLMs) show promise for atmospheric science but struggle with accuracy and hallucination rates in air pollution control tasks. Domain-specific adaptation is crucial for reliable LLM deployment in environmental decision-making.
Area of Science:
- Atmospheric science
- Environmental science
- Artificial intelligence
Background:
- Large language models (LLMs) offer potential for advancing atmospheric research and air pollution control.
- Lack of systematic evaluation benchmarks hinders trust and real-world deployment of LLMs for critical environmental tasks.
Purpose of the Study:
- To develop and utilize a benchmark dataset (OneAtmos-Bench) for evaluating LLMs in air pollution control.
- To assess the performance of 11 LLMs on accuracy, instruction-following, and hallucination rates within the atmospheric domain.
Main Methods:
- A two-stage automated extraction pipeline was used to create the OneAtmos-Bench dataset.
- Eleven LLMs were evaluated using the OneAtmos-Bench dataset, focusing on accuracy, instruction-following (IF), and hallucination rate (HR).
- Expert validation confirmed the dataset's high accuracy (95.88%).
Main Results:
- LLMs demonstrated strong instruction-following capabilities.
- Accuracy and hallucination suppression remain challenges, particularly for specialized atmospheric tasks.
- Model scaling showed limited gains, indicating sparse domain adaptation; fine-tuning methods may increase hallucination rates.
Conclusions:
- General-purpose LLMs face reliability issues for trustworthy, low-hallucination air pollution control guidance.
- Environment-specific adaptation is necessary to overcome current LLM limitations in atmospheric science applications.
- Further research is needed to enhance LLM trustworthiness for critical environmental decision-making.
Related Concept Videos
Improving Translational Accuracy
3.5K
3.5K
Improving Translational Accuracy
14.0K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.0K
