Related Experiment Video
Updated: Feb 17, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Unveiling impact of domain knowledge and data scale on open-source large language model specialization in anaerobic
Yi Zhang1, Fangyun Wang2, Yijing Feng2
1School of Environment, Tsinghua University, Beijing 100084, China.
None:
Exploring the integration of domain knowledge and its data scale holds critical value in enhancing the understanding of anaerobic digestion (AD) by open-source large language models (LLMs). This study develops an automated agent system to extract high-quality AD question-answer pairs from literature and fine-tunes three open-source LLMs. Expert evaluations reveal the fine-tuned Llama3.1-8B-AD (LAD) demonstrates notable professional competitiveness in specialized AD domains, exhibiting a knowledge comprehension level approaching that of GPT-4 in specific tasks (0.67 vs. 0.68). Notably, LAD demonstrates competitive or superior performance in advanced domains like additives and microbial knowledge. Additionally, fine-tuning with complete training datasets notably improves professional understanding (0.60 to 0.67), particularly in medium-to-high-difficulty questions. In contrast, training with only 50% of the data leads to unstable foundational knowledge hallucinations, underscoring the necessity of comprehensive domain data for deep understanding and reasoning in AD. Overall, this study provides a reference paradigm for the future development of large models for anaerobic digestion in the bioenergy field and offers new perspectives for the construction of intelligent bioenergy systems.
Related Concept Videos
Environmental Applications of Microorganisms
Overview of Archaea
Microbial Fermentation
Bioremediation
Metabolism of Chemolithotrophs
Microbial Nutrition

