Related Experiment Video
Updated: Sep 15, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Augmenting Large Language Models With Automated, Bibliometrics-Powered Literature Search for Knowledge Distillation:
David B Kurland1, Daniel A Alber1, Adhith Palla1
1Department of Neurosurgery, New York University Langone Medical Center, New York, New York, USA.
Background And Objectives:
Scholarly output is accelerating in medical domains, making it challenging to keep up with the latest neurosurgical literature. The emergence of large language models (LLMs) has facilitated rapid, high-quality text summarization. However, LLMs cannot autonomously conduct literature reviews and are prone to hallucinating source material. We devised a novel strategy that combines Reference Publication Year Spectroscopy-a bibliometric technique for identifying foundational articles within a corpus-with LLMs to automatically summarize and cite salient details from articles. We demonstrate our approach for four common spinal conditions in a proof of concept.
Methods:
Reference Publication Year Spectroscopy identified seminal articles from the corpora of literature for cervical myelopathy, lumbar radiculopathy, lumbar stenosis, and adjacent segment disease. The article text was split into 1024-token chunks. Queries from three knowledge domains (surgical management, pathophysiology, and natural history) were constructed. The most relevant article chunks for each query were retrieved from a vector database using chain-of-thought prompting. LLMs automatically summarized the literature into a comprehensive narrative with fully referenced facts and statistics. Information was verified through manual review, and spine surgery faculty were surveyed for qualitative feedback.
Results:
Our tandem approach cost less than $1 for each condition and ran within 5 minutes. Generative Pre-trained Transformer-4 was the best-performing model, with a near-perfect 97.5% citation accuracy. Surveys of spine faculty helped refine the prompting scheme to improve the cohesion and accessibility summaries. The final artificial intelligence-generated text provided high-fidelity summaries of each pathology's most clinically relevant information.
Conclusion:
We demonstrate the rapid, automated summarization of seminal articles for four common spinal pathologies, with a generalizable workflow implemented using consumer-grade hardware. Our tandem strategy fuses bibliometrics and artificial intelligence to bridge the gap toward fully automated knowledge distillation, obviating the need for manual literature review and article selection.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:35A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
Related Concept Videos
Spinal Cord: Gross Anatomy
Spinal Cord: Cross-sectional Anatomy
Gray Matter and its Components
Central to the gray matter is...
Spinal Cord: Information Processing
Sensory Information Processing
Sensory information processing begins at the sensory receptors located in the skin and other tissues, which detect somatic sensory stimuli such as touch, temperature, or pain. These receptors function as catalysts, initiating...
Spinal Cord
The Spinal Cord
Spinal Nerves: Anatomy
There are 31 bilateral pairs of spinal nerves, each emerging from the spinal cord through the intervertebral foramina—openings between adjacent vertebrae. These nerves are...