Related Experiment Video
Updated: Jul 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
A large language model pipeline for automated citation quality scoring across engineering journal quartiles with
Murat Isik1, Sultan Guleryuzlu2
1Department of Computer Engineering, Kirsehir Ahi Evran Universitesi, 40100, Kirsehir, Turkey. muratisik@ahievran.edu.tr.
A new two-stage large language model (LLM) pipeline automates citation quality scoring in academic papers. This tool assesses semantic relevance, challenging assumptions about citation quality and journal prestige, and aids academic integrity.
Area of Science:
- Bibliometrics
- Artificial Intelligence
- Scholarly Communication
Background:
- Assessing citation quality is crucial for academic integrity and understanding research impact.
- Traditional methods for citation analysis can be labor-intensive and subjective.
- Large Language Models (LLMs) offer potential for automating complex text analysis tasks.
Purpose of the Study:
- To propose and evaluate a novel two-stage LLM-based pipeline for automated citation quality scoring.
- To assess the semantic relevance of citation-reference pairs in academic manuscripts.
- To investigate the relationship between citation quality and journal prestige across different quartiles.
Main Methods:
- A two-stage LLM pipeline was developed: Stage 1 extracts and matches citations using Gemini 2.5 Flash; Stage 2 scores relevance using a second LLM with a structured rubric and skeptical persona.
- The pipeline was applied to 5,615 citation-reference pairs from 121 engineering articles across Web of Science quartiles (Q1-Q4).
- Statistical analyses (Kruskal-Wallis, Mann-Whitney U, Spearman correlation) were used to evaluate scores and inter-rater reliability.
Main Results:
- The pipeline achieved a mean relevance score of 7.76, with 74.7% of citations rated Strong or Excellent.
- Significant differences in citation scores were found across journal quartiles (p < 0.001), with Q2 articles showing the highest mean scores (8.04).
- A weak negative correlation between journal quartile rank and citation quality was observed (ρ = -0.105), with Q1 articles having the highest proportion of irrelevant citations (10.7%).
Conclusions:
- Citation quality does not monotonically improve with journal prestige, challenging existing assumptions.
- The LLM pipeline provides a scalable and content-aware tool for academic integrity, complementing existing solutions.
- The study highlights the potential of LLMs in nuanced academic content analysis and supports applications in editorial pre-screening and peer review.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Quartile
1; 1; 2; 2; 4; 6; 6.8; 7.2; 8; 8.3; 9; 10; 10; 11.5
The median or second quartile is seven. The lower half of the...
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Scaling
Quantifying and Rejecting Outliers: The Grubbs Test