Related Experiment Video
Updated: Jan 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking large language models for personalized, biomarker-based health intervention recommendations.
Hans Jarchow1, Christoph Bobrowski2, Steffi Falk3
1Institute for Biostatistics and Informatics in Medicine and Ageing Research, Rostock University Medical Center, Rostock, Germany.
Large language models (LLMs) show limited suitability for personalized longevity recommendations. While proprietary models performed better, all LLMs struggled with medical validation, stability, and age biases in this study.
Area of Science:
- Artificial Intelligence in Medicine
- Biomedical Informatics
- Computational Biology
Background:
- Large language models (LLMs) are increasingly used in healthcare, but their effectiveness for personalized longevity interventions is unclear.
- Existing frameworks for evaluating AI in medicine often do not cover the nuances of generating personalized health recommendations.
Purpose of the Study:
- To benchmark the performance of LLMs in generating personalized longevity intervention recommendations using the extended BioChatter framework.
- To assess LLM adherence to critical medical validation requirements for health recommendations.
Main Methods:
- Developed and utilized an extended BioChatter framework for benchmarking LLMs.
- Created 1000 diverse test cases from 25 individual profiles across three age groups, covering interventions like caloric restriction, fasting, and supplements.
- Evaluated 56,000 model responses using an LLM-as-a-Judge system against clinician-validated ground truths.
Main Results:
- Proprietary LLMs generally outperformed open-source models in comprehensiveness for longevity recommendations.
- All evaluated LLMs, even with Retrieval-Augmented Generation (RAG), demonstrated limitations in meeting medical validation standards, prompt stability, and addressing age-related biases.
- The study identified significant challenges in using LLMs for unsupervised generation of longevity advice.
Conclusions:
- Current LLMs have limited suitability for unsupervised personalized longevity intervention recommendations due to issues with medical validation, stability, and bias.
- The developed open-source framework provides a foundation for future AI benchmarking in medical applications.
- Further research is needed to refine LLMs for safe and effective personalized health guidance.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:35Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018