一个集成的深度学习框架,利用NASNet和视觉变压器与混合处理来准确和精确地诊断肺部疾病
Sajjad Saleem1, Muhammad Zaheer Sajid2, Abida Sharif3
1Department of Information and Technology, Washington University of Science and Technology, Alexandria, VA 22314, USA..
SLAS technology
|January 24, 2026
概括
一个新的AI模型,NASNet-ViT,使用医学图像准确诊断包括癌症,COVID-19,肺炎和结核病在内的肺部疾病. 这种深度学习方法为临床环境提供了高精度和高效率.
科学领域:
- 医疗成像中的人工智能
- 深度学习用于疾病诊断和诊断
- 医学图像分析 医学图像分析
背景情况:
- 肺部疾病对全球健康构成重大挑战,需要及时诊断.
- 准确的诊断对于改善患者在呼吸道疾病的治疗结果至关重要.
- 现有的诊断方法可能面临速度和准确性的限制.
研究的目的:
- 引入NASNet-ViT,这是一个用于肺部疾病分类的新型深度学习框架.
- 通过多级图像预处理管道 (MixProcessing) 提高诊断精度.
- 评估NASNet-ViT与已建立的深度学习架构的性能.
主要方法:
- 通过将NASNet的卷积特征与视觉转换器 (ViT) 注意力机制集成,开发了NASNet-ViT.
- 实现了MixProcessing管道:波形变换,自适应式直方图等级和形态过.
- 在肺部图像上训练并测试模型,分为五个类别:正常,肺癌,COVID-19,肺炎和结核.
主要成果:
- NASNet-ViT实现了98.9%的准确性,0.99的灵敏度,0.989的F1得分和0.987的特异性.
- 与MixNet-LD,D-ResNet,MobileNet和ResNet50.50.相比,其表现出了更高的性能.
- 实现了轻量级模型大小 (25.6 MB) 和快速推断时间 (12.4 秒).
结论:
- 纳斯网-ViT提供了一个强大的和可扩展的AI解决方案,用于精确的肺病诊断.
- 该模型是实用的实时部署在资源有限的临床环境.
- 这项研究推进了医学成像中的AI,支持临床医生在早期和准确的诊断中.
相关概念视频
Vision
59.5K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.5K
Uncertainty in Measurement: Accuracy and Precision
100.5K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value.
100.5K
Color Vision
1.4K
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
1.4K
Lung Capacity
56.2K
The air in the lungs is measured in volumes and capacities. Lung volume measures reflect the amount of air taken in, released, or left over after a lung function, like a single inhalation. Lung capacity measures are sums of two or more lung volume measures.
56.2K
Bacterial Transformation
59.5K
In 1928, bacteriologist Frederick Griffith worked on a vaccine for pneumonia, which is caused by Streptococcus pneumoniae bacteria. Griffith studied two pneumonia strains in mice: one pathogenic and one non-pathogenic. Only the pathogenic strain killed host mice.
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
Griffith made an unexpected discovery when he killed the pathogenic strain and mixed its remains with the live, non-pathogenic strain. Not only did the mixture kill host mice, but it also contained living pathogenic bacteria that...
59.5K
Depth Perception and Spatial Vision
1.9K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.9K


