Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Systematic benchmark of reduced-lead configurations for 12-lead ECG reconstruction: multi-model evaluation across all possible subsets.

Frontiers in cardiovascular medicine·2026
Same author

DCBM-Tri: a dual-channel bilinear mapping triplet model for early recognition of acute kidney injury in imbalanced cohorts.

Scientific reports·2026
Same author

Evaluation and application of electrocardiographic age model for children.

Scientific reports·2026
Same author

Generalized Few-Shot MM-Former For Surgical Scene Panoptic Segmentation.

Healthcare technology letters·2025
Same author

Correction: A human-LLM collaborative annotation approach for screening articles on precision oncology randomized controlled trials.

BMC medical research methodology·2025
Same author

Using Low-Intensity Focused Ultrasound to Treat Depression and Anxiety Disorders: A Review of Current Evidence.

Brain sciences·2025

Related Experiment Video

Updated: Mar 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K

Long-text caption generation for surgical image with a concept retrieval augmented large multimodal model.

Jiquan Liu1, Yichen Zhu1, Jingyi Feng2

  • 1Key Laboratory for Biomedical Engineering of Ministry of Education, College of Biomedical Engineering and Instrument Science, Zhejiang University, Hangzhou, Zhejiang, China.

Plos One
|March 17, 2026
PubMed
Summary

This study introduces a new framework for generating detailed surgical captions, overcoming limitations of existing models by using verified data and retrieval-augmented generation to improve accuracy and reduce hallucinations in medical AI.

More Related Videos

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

853
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

3.7K

Related Experiment Videos

Last Updated: Mar 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

1.3K
Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique
04:48

Swin-PSAxialNet: An Efficient Multi-Organ Segmentation Technique

Published on: July 5, 2024

853
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

3.7K

Area of Science:

  • Medical Imaging
  • Artificial Intelligence
  • Natural Language Processing

Background:

  • Automated surgical image captioning is vital for reporting and education.
  • Current Multimodal Large Language Models (MLLMs) struggle with long-text generation and hallucinate medical details due to limited datasets.

Purpose of the Study:

  • To develop a comprehensive framework for long-text surgical image captioning.
  • To address the lack of specialized datasets and mitigate hallucinations in MLLMs for surgical contexts.

Main Methods:

  • Constructed a verified long-text benchmark by enhancing the EndoVis2018 dataset with an automated pipeline and expert validation.
  • Implemented a retrieval-augmented generation (RAG) mechanism for domain-specific adaptation, injecting surgical knowledge into the visual encoder.
  • Established a robust evaluation protocol using clinically-aligned metrics for long medical text.

Main Results:

  • The proposed data-centric and retrieval-enhanced framework significantly outperforms baseline models.
  • The approach effectively mitigates domain-specific hallucinations common in generic MLLMs.
  • Generated captions are clinically accurate, coherent, and suitable for long medical text.

Conclusions:

  • The developed framework offers a significant advancement in long-text surgical image captioning.
  • Retrieval-augmented generation is a promising strategy for domain adaptation in medical AI.
  • The study highlights the importance of specialized datasets and evaluation metrics for medical NLP tasks.