卷积神经网络-视觉转换器架构,带有门式控制机制和多尺度融合,用于增强肺部疾病分类
Okpala Chibuike1, Xiaopeng Yang1,2
1Department of Human Ecology & Technology, Handong Global University, Pohang 37554, Republic of Korea.
Diagnostics (Basel, Switzerland)
|January 8, 2025
概括
结合视觉转换器 (ViT) 和卷积神经网络 (CNN) 的新型混合深度学习模型准确地分类肺部疾病,达到99.50%的准确性,用于改进医学图像分析和疾病诊断.
科学领域:
- 深度学习 (Deep Learning) 是一种深度学习.
- 医学成像分析 医学成像分析
- 计算机视觉 计算机视觉
背景情况:
- 视觉转换器 (ViT) 在全球范围内表现出色,但错过了对细粒度模式至关重要的高频细节.
- 卷积神经网络 (CNN) 能够有效地捕捉局部特征,但在医疗图像中难以处理远程依赖.
- 肺部疾病的准确分类需要整合本地和全球图像信息.
研究的目的:
- 开发一种混合深度学习架构,将ViT和CNN结合起来,用于增强医疗图像分类.
- 解决个人ViT和CNN模型在捕捉本地和全球图像特征方面的局限性.
- 提高从医学图像诊断肺部疾病的准确性和可靠性.
主要方法:
- 提出了一个混合架构,将ViT和CNN集成到模块化块中,以组合本地和全球特征提取.
- 实施了一个封闭的注意力机制 (通道,空间,元素智能),以选择性地强调关键特征.
- 整合了多尺度融合模块 (MSFM) 来整合不同尺度的功能,以实现全面的表示.
主要成果:
- 混合模型在分类四种不同的肺部疾病时,达到99.50%的高精度.
- 与仅依赖ViT或CNN的模型相比,表现出优异的性能.
- 废弃性研究证实了混合方法及其组件的有效性.
结论:
- 拟议的混合ViT-CNN模型显著提高了医疗图像分类性能.
- 该架构为医学成像中准确可靠的疾病诊断提供了一个有前途的框架.
- 该方法实现了良好的校准,表明了临床应用的可靠预测.
相关概念视频
Vision
52.9K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.9K
Visual System
501
Light enters the eye through the cornea, a transparent, dome-shaped surface covering the surface of the eyeball that helps to direct and focus incoming light. This light is then channeled toward the pupil, an adjustable opening whose size is controlled by the iris. The iris, a pigmented muscle, regulates the amount of light entering the eye by contracting or dilating the pupil, thereby ensuring optimal light levels for clear vision.
Once through the pupil, the light passes through the lens, a...
Once through the pupil, the light passes through the lens, a...
501


