Jove
Visualize
联系我们
JoVE
x logofacebook logolinkedin logoyoutube logo
关于 JoVE
概览领导团队博客JoVE 帮助中心
作者
出版流程编辑委员会范围与政策同行评审常见问题投稿
图书馆员
用户评价订阅访问资源图书馆顾问委员会常见问题
研究
JoVE JournalMethods CollectionsJoVE Encyclopedia of Experiments存档
教育
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab Manual教师资源中心教师网站
使用条款与条件
隐私政策
政策

相关概念视频

Vision01:24

Vision

59.2K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.2K

您也可能阅读

相关文章

通过共同作者、期刊和引用图与本文相关的文章。

排序
Same author

AI/ML-Assisted Detection of <i>HMGA2</i> RNA Isoforms in Prostate Cancer Patient Tissue.

International journal of molecular sciences·2026
Same author

A Multimodal Adaptive Inter-Region Attention-Guided Network for Brain Tumor Classification.

IEEE access : practical innovations, open solutions·2025
Same author

A Hybrid Learning-Architecture for Mental Disorder Detection Using Emotion Recognition.

IEEE access : practical innovations, open solutions·2024
Same author

Multi-Stage Classification of Retinal OCT Using Multi-Scale Ensemble Deep Architecture.

Bioengineering (Basel, Switzerland)·2023
Same author

Multi-Stage Classification-Based Deep Learning for Gleason System Grading Using Histopathological Images.

Cancers·2022
Same author

Predicting the Level of Respiratory Support in COVID-19 Patients Using Machine Learning.

Bioengineering (Basel, Switzerland)·2022

相关实验视频

Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

983

对用于放射学图像标题的低级调整视觉语言模型的实证评估.

Mahmudul Hoque1, Raisa Nusrat Chowdhury1, Md Rakibul Hasan2

  • 1Department of Computer Science, Morgan State University, Baltimore, MD 21251, USA.

Bioengineering (Basel, Switzerland)
|December 30, 2025
PubMed
概括

大型视觉语言模型 (VLMs) 在医学成像任务中通常优于较小的模型. 然而,有针对性的适应策略和建筑设计允许一些紧型号实现竞争性性能,帮助放射科医生的工作负载.

关键词:
低级别的适应 低级别的适应在ROCOv2数据集中,标题质量 标题质量医疗图像标题 标题 医学图像标题具有参数效率的微调.放射学报告的生成视觉语言模型 视觉语言模型

更多相关视频

Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
07:11

Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping

Published on: December 8, 2023

2.3K

相关实验视频

Last Updated: Jan 7, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

983
Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping
07:11

Author Spotlight: Insights into Visual Cortex Research Through Wide-View fMRI Mapping

Published on: December 8, 2023

2.3K

科学领域:

  • 医疗成像中的人工智能
  • 计算机视觉 计算机视觉
  • 自然语言处理自然语言处理.

背景情况:

  • 随着医疗成像数量的增加,需要自动化工具来支持放射科医生,并减少报告延迟.
  • 视觉语言模型 (VLMs) 显示了通过自动化标题生成加速医疗报告起草的潜力.
  • 用不同的参数尺度对VLM进行系统评估对于评估放射学中的临床效用至关重要.

研究的目的:

  • 评估和比较十个多式视觉语言模型 (VLMs) 的性能,这些模型在大型医学成像数据集 (ROCOv2) 上进行了微调.
  • 评估模型规模 (大VLM与小VLM) 和适应策略 (低级适应) 对医学成像中的VLM性能的影响.
  • 建立定量基准来选择VLM用于医学成像解释的临床应用.

主要方法:

  • 十个多式模式模型,包括大型VLM (LLaVA,IDEFICS-9B),小型VLM (MoonDream2,Qwen,SmolVLM) 和基线架构 (VisionGPT2,CNN-Transformer),在ROCOv2数据集上进行了微调 (116,635张图像,8种模式).
  • 应用了低级调整 (LoRA),重点是通过最小的参数更新 (<1%的总参数) 优化性能.
  • 模型使用相关性 (语义相似性) 和事实性 (概念级正确性) 的指标进行评估.

主要成果:

  • 性能按模型尺度分层清晰:大VLM (0.273-0.317),小VLM (0.188-0.279) 和基线 (0.154-0.177).
  • LLaVA-Mistral-7B实现了最高的整体性能,显著超过了VisionGPT2基线.
  • 月梦2 (小型VLM) 显示出竞争相关性得分,接近一些大型VLM的表现. 之前的模式标签显示了对小型VLM业绩的可变影响.

结论:

  • 模型尺寸是医疗成像VLM性能的一个重要因素,大型VLM通常优异.
  • 有针对性的适应策略,如优化的LoRA和特定的架构设计,使紧的小型VLM能够实现竞争性结果.
  • 这些发现为VLM选择提供了重要的基准,突出了临床放射学工作流程中高效,较小的模型的潜力.