Related Experiment Video
Updated: Aug 14, 2026

Simulation of a Scaled Assembly Process with Collaboration of a Robotic Arm and Monitoring through a Vision System for Quality Control
Published on: August 29, 2025
Vision-Based Digital Twin and AI Agent Framework for Low-Cost, Explainable Indoor Building Inspection and Safety
Zijian Jing1, Liyi Zhu2, Tianyi Chen1
1Key Laboratory of Urban and Architectural Heritage Conservation, School of Architecture, Southeast University, Nanjing 210096, China.
Abstract:
Aging residential buildings constructed under outdated design standards create an urgent need for scalable, evidence-based indoor safety assessment methods. Conventional manual inspections rely on subjective checklists, lack audit trails, and are impractical for widespread deployment. This study presents a vision-based digital twin and AI agent framework that converts a single continuous smartphone video into an explainable, evidence-constrained safety assessment. The pipeline employs MASt3R-SLAM to reconstruct a metric-scale 3D point cloud from monocular video, calibrated with AprilTag fiducials for absolute scale. SpatialLM parses the geometry to extract semantic entities and spatial relationships. Risk guidelines are formalized into a computable Risk Prototype structure, unified within a hierarchical SceneState data structure that binds geometric measurements, semantic labels, image observations, and regulatory knowledge. A LangGraph-based AI agent conducts a dual-pathway assessment: an initial whole-dwelling scan followed by iterative follow-up queries invoking tool calls for measurement, knowledge retrieval, or visual cross-checking. In a pilot validation across five heterogeneous residences, with detailed manual comparison in two representative cases, the framework achieved risk recall rates of 77.8-100% and precision rates of 45.0-70.0% against the single-assessor manual reference. The average judgment closure rate was 71.7%, with spatial granularity enhancement of up to 2.2× in complex environments. These results suggest that the framework can achieve risk coverage comparable to manual checklist inspection while offering enhanced granularity in complex environments and quantitative precision in well-defined spaces.
