Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Depth Perception and Spatial Vision01:15

Depth Perception and Spatial Vision

487
Depth perception is the ability to perceive objects three-dimensionally. It relies on two types of cues: binocular and monocular. Binocular cues depend on the combination of images from both eyes and how the eyes work together. Since the eyes are in slightly different positions, each eye captures a slightly different image. This disparity between images, known as binocular disparity, helps the brain interpret depth. When the brain compares these images, it determines the distance to an object.
487
Design Example: Measuring Distance Between Two Points with Obstructions01:10

Design Example: Measuring Distance Between Two Points with Obstructions

19
When measuring distances in areas with physical obstructions, such as a lake in a field, surveyors must employ techniques to calculate accurate lengths without direct line measurements. One effective method is the offset technique, which allows for precise distance estimation over inaccessible stretches.In this scenario, a surveyor must measure a side of an area that crosses a lake. Since the measuring tape cannot span the lake, the surveyor begins by establishing a baseline that aligns with...
19

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Melatonin promotes recovery from ischemic stroke by modulating microglia polarization and inhibiting oligodendrocyte pyroptosis via the RORα/AMPKα/STAT1 pathway.

Journal of translational medicine·2026
Same author

LINC01929 promotes breast cancer progression through a TFRC-associated ferroptosis pathway.

Cell death discovery·2026
Same author

GHF-ACL: A novel contrastive learning framework with multi-order graph structures for herb-disease association prediction.

PLoS computational biology·2026
Same author

TCRBinder: Unified pre-trained language model with paired-chain synergy for predicting T-cell receptor binding specificity.

PLoS computational biology·2026
Same author

GMHAN: a heterogeneous graph attention framework for prioritizing coding and non-coding driver genes.

Bioinformatics (Oxford, England)·2026
Same author

SINTER3D: continuous 3D reconstruction of spatial transcriptomics via implicit neural representations.

Genome biology·2026

Related Experiment Video

Updated: May 16, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

438

Monocular depth estimation via a detail semantic collaborative network for indoor scenes.

Wen Song1, Xu Cui2, Yakun Xie3

  • 1School of Architecture, Southwest Jiaotong University, Chengdu, 611756, China.

Scientific Reports
|March 31, 2025
PubMed
Summary

A new detail-semantic collaborative network (DSCNet) improves monocular depth estimation for indoor scenes. This method enhances accuracy and robustness by effectively fusing detail and semantic features for applications like smart space design.

Keywords:
Deep learningDetail–semantic collaborative networkIndoor scenesMonocular depth estimation

More Related Videos

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.6K
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

8.9K

Related Experiment Videos

Last Updated: May 16, 2025

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
03:31

Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications

Published on: December 15, 2023

438
Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography
04:48

Application of Deep Learning-Based Medical Image Segmentation via Orbital Computed Tomography

Published on: November 30, 2022

2.6K
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

8.9K

Area of Science:

  • Computer Vision
  • Artificial Intelligence
  • 3D Reconstruction

Background:

  • Monocular depth estimation is vital for indoor scene reconstruction, impacting energy efficiency, environmental modeling, and smart design.
  • Indoor scenes present challenges due to low depth variability, complex object correlations, and diverse object types, hindering model robustness.

Purpose of the Study:

  • To propose a novel network, the detail-semantic collaborative network (DSCNet), for robust and accurate monocular depth estimation in indoor environments.
  • To address limitations in detail feature extraction and semantic correlation understanding in existing indoor depth estimation models.

Main Methods:

  • Utilized a hierarchical transformer structure to capture comprehensive contextual image features.
  • Developed a detail-semantic collaborative structure with selective attention feature maps to extract and fuse both detail and semantic information.
  • Aggregated multi-level semantic and detailed features to model complex inter-object correlations.

Main Results:

  • Achieved state-of-the-art performance on the NYU and SUN indoor depth estimation datasets, outperforming 14 recent optimal methods.
  • Demonstrated improved perception ability and model accuracy without increasing parameter count.
  • Validated the model's stability, robustness, and practical availability in indoor scenes through comprehensive analysis and ablation experiments.

Conclusions:

  • The DSCNet effectively enhances monocular depth estimation for indoor scenes by synergistically leveraging detail and semantic information.
  • The proposed approach offers a robust and accurate solution for applications requiring precise indoor 3D understanding.
  • The method provides a significant advancement in handling the complexities of indoor environments for depth estimation tasks.