Related Experiment Video
Updated: Oct 3, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating large language models for domain-specific scholarly metadata classification
Muhammad Haris1, Maryam Badar2
1TIB-Leibniz Information Center for Science and Technology, Leibniz University Hannover, Hannover, Germany.
Introduction:
Large language models (LLMs) have been widely applied to text classification; however, their effectiveness in domain-specific scholarly metadata classification remains insufficiently explored. This study investigates the use of open-source LLMs to classify modeling and simulation research articles across three predefined metadata dimensions: Industry, Application Area, and Modeling Method.
Methods:
Using a curated dataset of research articles from the AnyLogic research repository, we evaluated models from the Qwen2.5, Llama3.1, Mistral, Ministral, and Gemma3 families. We compared Direct, Few-shot, and Chain-of-Thought prompting using either the article title alone or the title and abstract together. Performance was evaluated using Macro-F1 across the three metadata dimensions.
Results:
The best overall configuration, Ministral-3-14B with Direct prompting and the title and abstract as input, achieved a Macro-F1 of 53.8%. Performance varied substantially across the metadata dimensions. Modeling Method was the easiest to classify, reaching a maximum Macro-F1 of 69.1% with Ministral-3-14B. The best Macro-F1 scores for Industry and Application Area were 57.6% and 39.4%, respectively. Including abstracts consistently improved classification performance across the evaluated models, whereas differences among the three prompting strategies were relatively small. Categories that were semantically overlapping, broadly defined, highly imbalanced, or not explicitly mentioned in the article text were particularly difficult to classify.
Discussion:
The findings indicate that open-source LLMs can support domain-specific scholarly metadata classification without task-specific fine-tuning. However, their moderate and dimension-dependent performance limits their suitability for fully automated fine-grained metadata enrichment. Richer contextual information, more clearly defined taxonomies, and human validation are therefore needed to improve the reliability of LLM-assisted scholarly metadata classification.