深度计算机视觉与基于人工智能的手语识别,以帮助听力和语言障碍者
Abrar Almjally1,2, Wafa Sulaiman Almukadi3
1Department of Information Technology, College of Computer and Information Sciences, Imam Mohammad Ibn Saud Islamic University (IMSIU), Riyadh, 13318, Saudi Arabia. aamjally@imamu.edu.sa.
Scientific reports
|September 1, 2025
概括
这项研究引入了用于手语识别的新深度学习模型 (SLR). 基于哈里斯优化的深度学习模型用于手语识别 (HHODLM-SLR) 达到98.95%的准确性,改善了聋人和听力障碍者的沟通.
科学领域:
- 计算机视觉
- 人工智能
- 人与计算机的交互
背景情况:
- 标语识别对于使聋和听障人士融入社区至关重要.
- 深度学习 (DL) 和计算机视觉 (CV) 已经显著提升了SLR方法.
- 现有的SLR技术包括基于手套和基于视觉的方法.
研究的目的:
- 开发一个先进的手语自动检测和分类系统.
- 增强听力和语言障碍者沟通的可访问性.
- 推出一种基于哈里斯优化的深度学习模型来识别手语 (HHODLM-SLR).
主要方法:
- 使用双边过 (BF) 进行图像预处理以减少噪声.
- 使用ResNet-152模型进行特征提取.
- 由双向长期短期记忆 (Bi-LSTM) 模型执行的手语识别.
- 使用哈里斯优化 (HHO) 算法对Bi-LSTM模型进行超参数优化.
主要成果:
- 通过HHODLM-SLR技术表现出优异的分类性能.
- 该模型在SL数据集上达到98.95%的高精度.
- 实验分析证实了该方法对现有技术的有效性.
结论:
- 拟议的HHODLM-SLR模型在手语识别方面取得了重大进展.
- 这项技术有潜力打破聋人和听力障碍者的沟通障碍.
- 集成HHO用于超参数调提高了SLR系统的可靠性和准确性.
相关概念视频
Vision
48.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
48.6K
Depth Perception and Spatial Vision
2.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.7K
Visual Agnosia
2.0K
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
2.0K
Prosopagnosia
1.3K
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
1.3K
Learning Disabilities
697
Learning disabilities are cognitive disorders caused by neurological impairments that affect cognitive functions like language and reading, without indicating overall intellectual or developmental challenges. These disabilities differ from global intellectual or developmental disabilities as they are limited to distinct cognitive functions. Common learning disabilities include dysgraphia, dyslexia, and dyscalculia, each of which impacts unique aspects of learning.
Dyslexia
Dyslexia is a...
Dyslexia
Dyslexia is a...
697


