Benchmarking large language models for personalized, biomarker-based health intervention recommendations.

Hans Jarchow1, Christoph Bobrowski2, Steffi Falk3

  • 1Institute for Biostatistics and Informatics in Medicine and Ageing Research, Rostock University Medical Center, Rostock, Germany.

NPJ Digital Medicine
|October 28, 2025
PubMed
Summary

Large language models (LLMs) show limited suitability for personalized longevity recommendations. While proprietary models performed better, all LLMs struggled with medical validation, stability, and age biases in this study.