Related Experiment Video
Updated: Jun 7, 2025

Using a Virtual Reality Walking Simulator to Investigate Pedestrian Behavior
Published on: June 9, 2020
Decision-making of autonomous vehicles in interactions with jaywalkers: A risk-aware deep reinforcement learning
Ziqian Zhang1, Haojie Li1, Tiantian Chen2
1School of Transportation, Southeast University, China; Jiangsu Key Laboratory of Urban ITS, China; Jiangsu Province Collaborative Innovation Center of Modern Urban Traffic Technologies, China.
Abstract:
Jaywalking, as a hazardous crossing behavior, leaves little time for drivers to anticipate and respond promptly, resulting in high crossing risks. The prevalence of Autonomous Vehicle (AV) technologies has offered new solutions for mitigating jaywalking risks. In this study, we propose a risk-aware deep reinforcement learning (DRL) approach for AVs to make decisions safely and efficiently in jaywalker-vehicle interactions. Notably, a risk prediction module is incorporated into the traditional DRL framework, making the AV agent risk-aware. Considering the complexity of jaywalker-vehicle conflicts, an encoder-decoder model is adopted as the risk prediction module, which comprehensively integrates multi-source data and predicts probabilities of the final conflict severity levels. The risk-aware DRL approach is applied in a simulated environment established in Anylogic, where the motion features of jaywalkers and vehicles are calibrated using real-world survey data. The trained driving policies are evaluated from perspectives of safety and efficiency across three scenarios with escalading levels of jaywalker volume. Regarding safety performance, the Baseline policy performs the worst in "medium jaywalker volume" scenario and "high jaywalker volume" scenario, while our Proposed risk-aware method outperforms the other methods, with the "low TTC ratio" metric stabilizing near 0.08. Moreover, as the scenario gets more complex, the superiority of our Proposed risk-aware policy gets more evident. In terms of efficiency performance, our Proposed risk-aware policy ranks the second best, achieving an "AV delay" metric around 8.1 s in the "medium jaywalker volume" scenario and 8.5 s in the "high jaywalker volume" scenario. In practice, the proposed risk-aware DRL approach can help AV agents perceive potential risks in advance and navigate through potential jaywalking areas safely and efficiently, further enhancing pedestrian safety.
More Related Videos
06:28A Networked Desktop Virtual Reality Setup for Decision Science and Navigation Experiments with Multiple Participants
Published on: August 26, 2018
07:42An Automated T-maze Based Apparatus and Protocol for Analyzing Delay- and Effort-based Decision Making in Free Moving Rodents
Published on: August 2, 2018
Related Concept Videos
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...
Decision Making: P-value Method
First, a specific claim about the population parameter is proposed. The claim is based on the research question and is stated in a simple form. Further, an opposing statement to the claim is also stated. These statements can act as null and alternative hypotheses: a null hypothesis would be a neutral statement while the alternative hypothesis can...
Observational Learning
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example: