Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

1.8K
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
1.8K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Tactile Interaction with Socially Assistive Robots for Children with Physical Disabilities.

Sensors (Basel, Switzerland)·2025
Same author

Multi-Functional Reconfigurable Intelligent Surfaces for Enhanced Sensing and Communication.

Sensors (Basel, Switzerland)·2023
See all related articles

Related Experiment Video

Updated: Jan 13, 2026

Photorealistic Learned Landscapes for Augmented Reality
06:54

Photorealistic Learned Landscapes for Augmented Reality

Published on: June 27, 2025

667

Large Language Models and 3D Vision for Intelligent Robotic Perception and Autonomy.

Vinit Mehta1, Charu Sharma1, Karthick Thiyagarajan2

  • 1Machine Learning Lab, International Institute of Information Technology (IIIT), Hyderabad 500032, Telangana, India.

Sensors (Basel, Switzerland)
|October 29, 2025
PubMed
Summary

Large Language Models (LLMs) combined with 3D vision are revolutionizing robotic sensing. This integration enhances robots

Keywords:
3D visionembodied agentshuman robot interactionlarge language modelsrobot sensingscene understandingsensor applicationsvisual sensing

More Related Videos

Simulation of a Scaled Assembly Process with Collaboration of a Robotic Arm and Monitoring through a Vision System for Quality Control
05:47

Simulation of a Scaled Assembly Process with Collaboration of a Robotic Arm and Monitoring through a Vision System for Quality Control

Published on: August 29, 2025

414
Haptic/Graphic Rehabilitation: Integrating a Robot into a Virtual Environment Library and Applying it to Stroke Therapy
13:44

Haptic/Graphic Rehabilitation: Integrating a Robot into a Virtual Environment Library and Applying it to Stroke Therapy

Published on: August 8, 2011

14.5K

Related Experiment Videos

Last Updated: Jan 13, 2026

Photorealistic Learned Landscapes for Augmented Reality
06:54

Photorealistic Learned Landscapes for Augmented Reality

Published on: June 27, 2025

667
Simulation of a Scaled Assembly Process with Collaboration of a Robotic Arm and Monitoring through a Vision System for Quality Control
05:47

Simulation of a Scaled Assembly Process with Collaboration of a Robotic Arm and Monitoring through a Vision System for Quality Control

Published on: August 29, 2025

414
Haptic/Graphic Rehabilitation: Integrating a Robot into a Virtual Environment Library and Applying it to Stroke Therapy
13:44

Haptic/Graphic Rehabilitation: Integrating a Robot into a Virtual Environment Library and Applying it to Stroke Therapy

Published on: August 8, 2011

14.5K

Area of Science:

  • Robotics and Artificial Intelligence
  • Computer Vision
  • Natural Language Processing

Background:

  • Rapid advancements in AI and robotics necessitate sophisticated sensing capabilities.
  • Integrating Large Language Models (LLMs) with 3D vision offers a novel approach to robotic perception.
  • Bridging linguistic intelligence and spatial understanding is key for advanced human-robot interaction.

Purpose of the Study:

  • To provide a comprehensive review of LLMs and 3D vision integration for robotic sensing.
  • To analyze state-of-the-art methodologies, applications, and challenges in this interdisciplinary field.
  • To identify future research directions for more intelligent and autonomous robotic systems.

Main Methods:

  • Introduction to foundational principles of LLMs and 3D data representations.
  • In-depth examination of 3D sensing technologies relevant to robotics.
  • Exploration of advancements in scene understanding, text-to-3D generation, object grounding, and embodied agents.

Main Results:

  • Highlighting cutting-edge techniques like zero-shot 3D segmentation and language-guided manipulation.
  • Discussion of multimodal LLMs integrating diverse sensory inputs (touch, auditory, thermal) for enhanced comprehension.
  • Cataloging benchmark datasets and evaluation metrics for 3D-language and vision tasks.

Conclusions:

  • The convergence of LLMs and 3D vision is crucial for next-generation robotic sensing.
  • Key challenges include adaptive architectures, cross-modal alignment, and real-time processing.
  • Future research will enable more context-aware and autonomous robotic sensing systems.