Related Experiment Video
Updated: Sep 13, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Assessing the efficacy of pre-trained large language models in analyzing autonomous vehicle field test disengagements
Melika Ansarinejad1, Sherif M Gaweesh2, Mohamed M Ahmed1
1Department of Civil and Architectural Engineering and Construction Management, University of Cincinnati (UC), Cincinnati, OH 45221, USA.
Abstract:
This study evaluates the efficacy of pre-trained large language models (LLMs) in analyzing disengagement reports of Levels 2-3 autonomous vehicle (AV) field tests, utilizing data provided from California Department of Motor Vehicles. Disengagement reports document instances where autonomous vehicles, tested under the Autonomous Vehicle Tester (AVT) and AVT Driverless Programs, transition from autonomous to manual control. These disengagements occur when human intervention is required due to incidents or limitations in the operational design domain that prevent AVs from functioning properly. Understanding factors leading to disengagements is pivotal for assessing AV performance and guiding infrastructure owners and operators (IOOs) about modifications needed. Manual approaches for analysis of the disengagement data are labor-intensive and prone to human error. Our research investigates the capability of LLMs to automate this analysis, focusing on identifying patterns, categorizing disengagement causes, and extracting meaningful insights from extensive datasets. GPT-4o as an LLM was employed to analyze the disengagement reports. The study aims to measure the accuracy, efficiency, and reliability of these models in comparison to traditional techniques. The application of LLMs demonstrated significant potential in identifying insights from the disengagement dataset, while effectively processing the textual data, achieving an accuracy of 87%. Several data limitations were encountered, including inconsistencies in disengagement descriptions from different manufacturers, which posed challenges to standardizing the analysis. Additionally, the disengagement reports offered limited details on the specific causes of disengagements and the surrounding conditions, restricting the depth of insights that could be drawn. Despite these challenges, our findings indicate that LLMs can substantially enhance the speed and precision of analyzing AV disengagement reports, offering valuable insights, while being cost-effective, that can inform further research and development in AV technology and safety protocols.

