Related Experiment Video
Updated: May 20, 2026

Eye Tracking During A Complex Aviation Task For Insights Into Information Processing
Published on: April 4, 2025
Revealing the divergences between LLM-simulated and human takeover decision-making in ADS-equipped HGV operations
Zheng Xu1, Yinfei Xi2, Jun Hua3
1School of Traffic and Transportation Engineering, Central South University, Changsha, Hunan 410075, China; State Key Lab of Intelligent Transportation System, Research Institute of Highway Ministry of Transport, Beijing 100088, China.
None:
Global adoption of Level-3 (L3) autonomous driving systems (ADS) in commercial heavy goods vehicles (HGVs) remains limited, as concerns about reliability and safety foster trust miscalibration and reinforce experience-dependent risk heuristics. This study examines human-ADS interaction in automated HGV operations and investigates whether large language models (LLMs) can generate takeover decisions that resemble human behavior. Human-in-the-loop experiments were conducted within the fully compiled CARLA simulation platform featuring a large-scale road environment (30 × 30 km2) with realistic background traffic flow dynamics. Human reactions, particularly takeover behavior during ADS operations, were systematically examined from compliance rates, decision latency, intervention strategies, and crash outcomes across freeway merging, freeway navigating, and urban driving scenarios. The same scenarios were presented to LLM-agent-equipped HGVs, where the agents were prompted and contextualized as scalable surrogates for "truck drivers", and their performance was then recorded and compared with human counterparts. Results showed that (i) crash rates dropped substantially from human participants (12.0 %) to LLM-agent-controlled ADS (2.8 %-4.6 %); (ii) whereas LLM agents relied exclusively on deceleration and steering for safety responses, human participants also frequently used acceleration as a preferred intervention strategy; (iii) participants with 6-20 years of driving experience intervened more pre-emptively and exhibited below 50 % compliance with ADS operation. These findings suggest that human takeover decisions are shaped by experience-based risk heuristics and exhibit substantial heterogeneity across drivers. They also indicate that current LLMs are not yet reliable proxies for human supervisory behavior during the autonomous driving. The methodological framework developed in this investigation offers a replicable approach for researchers examining human-ADS interaction across transportation domains, and the empirical findings inform appropriate boundaries for LLM application in safety-critical human factors research.
