一种基于空间和通道重建的ResNet结合多线索融合的凝视估计方法
Zhaoyu Shou1,2, Yanjun Lin1, Jianwen Mo1
1School of Information and Communication, Guilin University of Electronic Technology, Guilin 541004, China.
Journal of imaging
|April 25, 2025
概括
准确估计学生的目光是网上学习注意力的关键. 我们的RSP-MCGaze模型通过更好地提取特征和分析头部,脸部和眼部区域来改善视线估计,优于现有方法.
科学领域:
- 计算机科学 计算机科学
- 人与计算机的交互
- 机器学习 机器学习
背景情况:
- 由于复杂的影响因素,在线学习的注意力很难衡量.
- 目前基于外观的目光估计模型在特征提取和时空关系方面扎,导致高角度误差.
- 准确的凝视点估计对于评估和增强在线学习者参与度至关重要.
研究的目的:
- 提出一种基于外表的新目光估计模型 (RSP-MCGaze),以提高在线学习环境中的准确性.
- 为了增强特征提取和时空建模用于凝视估计.
- 为了减少视角点估计中的角误差.
主要方法:
- 通过集成ResNet和SCConv来开发一个特征提取骨干网络 (ResNetSC),以实现高效的特征提取和减少冗余.
- 通过联合定位头部,眼睛和面部区域来优化视频凝视估计.
- 采用基于外表的方法来估计眼神.
主要成果:
- 与公开数据集上的现有基线模型相比,RSP-MCGaze模型的性能明显更高.
- 在Gaze360数据集上实现了9.86的检测错误.
- 在Gaze360.0.的可检测面部子集上实现了7.11的检测错误.
结论:
- 拟议的RSP-MCGaze模型在视线估计任务中提供了卓越的性能.
- 有效的特征提取和时空建模对于准确的目光估计至关重要.
- 该模型显示了增强在线学习者注意力评估的强大潜力.
相关概念视频
Depth Perception and Spatial Vision
453
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
453
Deconvolution
116
Deconvolution, also known as inverse filtering, is the process of extracting the impulse response from known input and output signals. This technique is vital in scenarios where the system's characteristics are unknown, and they must be inferred from the observable signals.
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
Deconvolution involves several mathematical techniques to derive the impulse response. One common approach is polynomial division. In this method, the input and output sequences are treated as coefficients of...
116


