聴覚や言語障害者を支援する,人工知能ベースの手話認識による深度コンピュータビジョン
Abrar Almjally1,2, Wafa Sulaiman Almukadi3
1Department of Information Technology, College of Computer and Information Sciences, Imam Mohammad Ibn Saud Islamic University (IMSIU), Riyadh, 13318, Saudi Arabia. aamjally@imamu.edu.sa.
Scientific reports
|September 1, 2025
まとめ
この研究は,手話認識 (SLR) の新しいディープラーニングモデルを導入しています. ハリス・ホーク・オプティマイゼーション・ベース・ディープ・ラーニング・モデル (HHODLM-SLR) は98.95%の精度を達成し,聴覚障害者のコミュニケーションを向上させます.
科学分野:
- コンピュータ・ビジョン
- 人工知能
- 人とコンピュータの相互作用
背景:
- 聴覚障害者のコミュニティへの統合には,手話認識 (SLR) が不可欠です.
- ディープラーニング (DL) とコンピュータビジョン (CV) は,SLRの方法を大幅に進歩させました.
- 既存のSLR技術には,グローブベースのアプローチとビジョンベースのアプローチが含まれています.
研究 の 目的:
- 手話 (SL) の高度な自動検知・分類システムを開発する.
- 聴覚障害や言語障害のある人のためのコミュニケーションのアクセシビリティを向上させる.
- 手話認識のための新しいハリスホーク最適化ベースのディープラーニングモデル (HHODLM-SLR) を導入する.
主な方法:
- 騒音を減らすために双方向フィルタリング (BF) を使用した画像の事前処理.
- ResNet-152モデルを使用した特徴抽出.
- 双方向の長期短期記憶 (Bi-LSTM) モデルによる手話認識
- ハリス・ホーク最適化 (HHO) アルゴリズムを用いたBi-LSTMモデルのハイパーパラメータ最適化.
主要な成果:
- HHODLM-SLR技術は優れた分類性能を示した.
- このモデルはSLデータセットで98.95%の高い精度を達成した.
- 実験分析により,既存の技術に対する方法論の有効性が確認されました.
結論:
- 提案されたHHODLM-SLRモデルは,手話認識の重要な進歩を提供します.
- この技術は 聴覚障害者のコミュニケーションの障壁を 打破する可能性を秘めています
- ハイパーパラメータチューニングのためのHHOの統合は,SLRシステムの信頼性と精度を高めます.
関連する概念動画
Vision
48.6K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
48.6K
Depth Perception and Spatial Vision
2.7K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
2.7K
Visual Agnosia
2.0K
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
2.0K
Prosopagnosia
1.3K
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
1.3K
Learning Disabilities
697
Learning disabilities are cognitive disorders caused by neurological impairments that affect cognitive functions like language and reading, without indicating overall intellectual or developmental challenges. These disabilities differ from global intellectual or developmental disabilities as they are limited to distinct cognitive functions. Common learning disabilities include dysgraphia, dyslexia, and dyscalculia, each of which impacts unique aspects of learning.
Dyslexia
Dyslexia is a...
Dyslexia
Dyslexia is a...
697


