将ChatGPT视觉 (GPT-4V) 测试:在交通图像中的风险感知
Tom Driessen1, Dimitra Dodou1, Pavlo Bazilinskyy2
1Delft University of Technology, Delft, Zuid-Holland, The Netherlands.
Royal Society open science
|July 30, 2024
概括
这项研究表明,GPT-4V可以准确地预测人类在交通图像中的风险感知. 促进类似于人类问卷的策略可以改善AI.
科学领域:
- 人工智能的人工智能
- 计算机视觉 计算机视觉
- 人与计算机的交互
背景情况:
- 自动驾驶系统尽管有先进的计算机视觉,但仍难以理解上下文.
- 视觉语言模型有可能提高车辆的情境意识.
- 在交通场景中对人类风险的评估是复杂和主观的.
研究的目的:
- 评估GPT-4V在预测交通图像中人类评估的风险水平方面的有效性.
- 在风险评估任务中探索视觉语言模型的最佳提示策略.
- 确定是否整合对象检测功能可以改善基于AI的风险预测.
主要方法:
- 利用了210张静态交通图像,此前约650人对风险进行了评级.
- 应用心理测量构造理论和自我一致性提示方法.
- 制定并测试了关于快速重复,变化和特征集成的三个假设.
主要成果:
- 这三种假设都得到了证实,证明了特定提示技术的有效性增加.
- 在预测人类风险得分方面,获得了高的有效率系数 (r = 0.83).
- 基于GPT-4V的风险评级与物体检测功能相结合,显著提高了预测准确性.
结论:
- 在交通场景中,GPT-4V可以准确地预测人口层面的人类风险感知.
- 使用类似于多项人体调查问卷的方法促进GPT-4V提高了预测有效性.
- 当适当提示时,人工智能模型显示出通过理解交通环境来提高自动驾驶安全的巨大潜力.
更多相关视频
06:25Author Spotlight: Assessment of Visual Acuity in Central Vision Loss Through Motion-Based Peripheral Vision Testing
Published on: February 23, 2024
574
11:12Driving Simulation in the Clinic: Testing Visual Exploratory Behavior in Daily Life Activities in Patients with Visual Field Defects
Published on: September 18, 2012
17.4K
相关概念视频
Prosopagnosia
153
Prosopagnosia, also known as face blindness, is the inability to recognize faces. In severe cases, individuals with prosopagnosia may not recognize close family members, including parents and spouses, by their faces. For instance, someone with prosopagnosia might walk past their child in a crowd, only realizing their mistake upon noticing their child's distinctive backpack or favorite jacket. Prosopagnosia specifically impairs facial recognition, while the recognition of other objects or...
153
Color Vision
548
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
548
Depth Perception and Spatial Vision
616
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
616
