Related Experiment Videos
AIvaluateXR: An Evaluation Framework for On-Device AI in XR With Benchmarking Results.
Summary
AIvaluateXR benchmarks large language models (LLMs) on extended reality (XR) devices. This framework helps select optimal LLM-device pairs for real-time applications, guiding future XR human-AI interaction research.
Area of Science:
- Artificial Intelligence
- Human-Computer Interaction
- Extended Reality
Background:
- Large language models (LLMs) offer significant potential for advancing human-AI interaction on extended reality (XR) devices.
- Selecting appropriate LLMs and XR devices for on-device inference presents a challenge due to performance variability.
Purpose of the Study:
- To introduce AIvaluateXR, a novel evaluation framework for benchmarking LLMs on XR devices.
- To provide a systematic method for assessing LLM performance and optimizing deployment on XR hardware.
Main Methods:
- Deployed 17 LLMs across four XR platforms (Magic Leap 2, Meta Quest 3, Vivo X100s Pro, Apple Vision Pro).
- Measured performance consistency, processing speed, memory usage, and battery consumption for 68 model-device pairs.
- Analyzed performance under varying string lengths, batch sizes, and thread counts, utilizing 3D Pareto Optimality for selection.
Main Results:
- Comprehensive benchmarking data was collected for diverse LLM-device configurations.
- Trade-offs between performance metrics were identified for real-time XR applications.
- On-device LLM efficiency was compared against client-server and cloud-based approaches, with accuracy evaluated on interactive tasks.
Conclusions:
- AIvaluateXR offers valuable insights for optimizing LLM deployment on XR devices.
- The proposed unified evaluation method can serve as standard groundwork for future research in on-device AI for XR.