改进的yolov5算法与深度摄像头和盲人室内视觉辅助嵌入式系统相结合.
Kaikai Zhang1, Yanyan Wang1, Shengzhe Shi1
1School of Computer Science and Technology, Huaibei Normal University, 235000, Huaibei, China.
Scientific reports
|October 3, 2024
概括
这项研究引入了一种改进的YOLOv5对象检测算法,用于视力障碍者,增强室内物体检测设备. 新系统提供了更好的便携性和降低成本,帮助日常生活.
科学领域:
- 计算机视觉 计算机视觉
- 辅助技术 辅助技术 辅助技术
- 人工智能的人工智能
背景情况:
- 目前用于视力障碍者的室内物体寻找辅助器件的便携性不佳,成本高昂,环境易受影响.
- 需要改进,具有成本效益和便携性的解决方案来帮助视力受损者在日常物体识别中.
研究的目的:
- 开发了一种改进的YOLOv5算法,并将其集成到用于视力受损者室内物体寻找设备中.
- 通过提高可移植性,降低硬件成本和提高性能来解决当前辅助技术的局限性.
主要方法:
- 一个改进的YOLOv5算法使用GhostNet作为骨干,结合了坐标注意力机制和双向特征金字塔网络.
- 该算法与RealSense D435i深度摄像头和Raspberry Pi 4 B设备上的语音系统集成.
- 该系统处理语音命令进行对象识别,并使用RGB和深度图像进行检测和测距.
主要成果:
- 与原来的YOLOv5.5相比,改进的YOLOv5模型实现了42.4%的模型大小减少和47.9%的参数减少.
- 召回率增加了1.2%,同时保持了相同的精度.
- 集成设备成功检测和距离物体,提供语音反关于距离,以协助用户.
结论:
- 增强的YOLOv5算法显著优化了模型大小和参数,使其适用于便携式辅助设备.
- 开发的室内物体寻找设备通过提供准确的物体检测,距离和语音反,有效地帮助视力受损者.
- 这项技术提供了一个有希望的,具有成本效益的,便携式的解决方案,用于提高视力障碍者的独立性和日常生活.
相关概念视频
Depth Perception and Spatial Vision
602
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
602
Blind Procedures
10.6K
Ideally, the people who observe and record the children’s behavior are unaware of who was assigned to the experimental or control group, in order to control for experimenter bias. Experimenter bias refers to the possibility that a researcher’s expectations might skew the results of the study. Remember, conducting an experiment requires a lot of planning, and the people involved in the research project have a vested interest in supporting their hypotheses. If the observers knew which...
10.6K
Vision
53.0K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
53.0K
Light Acquisition
8.4K
In order to produce glucose, plants need to capture sufficient light energy. Many modern plants have evolved leaves specialized for light acquisition. Leaves can be only millimeters in width or tens of meters wide, depending on the environment. Due to competition for sunlight, evolution has driven the evolution of increasingly larger leaves and taller plants, to avoid shading by their neighbors with contaminant elaboration of root architecture and mechanisms to transport water and nutrients.
8.4K
Visual Agnosia
179
Visual agnosia is a condition characterized by the inability to recognize visually presented objects despite having normal vision. For instance, a person with visual agnosia can describe the shape and color of an object but cannot identify or name it. This impairment does not affect their visual field, acuity, color vision, brightness discrimination, language, or memory. An example of this condition in a social setting is someone at a dinner party asking for "that silver thing with a round...
179
Sight Distance in a Vertical Curve
37
Sight distance on vertical curves is critical in roadway design. It ensures drivers can see far enough ahead to identify and respond to hazards effectively. This directly impacts safety, driver comfort, and the overall efficiency of the transportation network.Vertical curves are classified into crest and sag curves based on their geometry. For crest curves, sight distance is determined by the line of sight between a driver's eye and a small object on the road's surface. Design parameters for...
37


