Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Mar 28, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
08:25

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

Published on: May 7, 2019

9.7K

Human-Machine CRFs for Identifying Bottlenecks in Scene Understanding.

Roozbeh Mottaghi, Sanja Fidler, Alan Yuille

    IEEE Transactions on Pattern Analysis and Machine Intelligence
    |December 15, 2015
    PubMed
    Summary

    This study integrates human input into computer vision models for better scene understanding. Hybrid models reveal which tasks, like object detection, offer the most potential for future AI advancements.

    Related Concept Videos

    You might also read

    Related Articles

    Articles linked to this work by shared authors, journal, and citation graph.

    Sort by
    Same author

    ImageNet3D: Towards General-Purpose Object-Level 3D Understanding.

    Advances in neural information processing systems·2026
    Same author

    CLIP-Driven Universal Model for Organ Segmentation and Tumor Detection.

    Proceedings. IEEE International Conference on Computer Vision·2026
    Same author

    Acquiring Weak Annotations for Tumor Localization in Temporal and Volumetric Data.

    Machine intelligence research (Beijing)·2026
    Same author

    Scaling 3D Compositional Models for Robust Classification and Pose Estimation.

    Proceedings. IEEE International Conference on Computer Vision·2026
    Same author

    A comprehensive survey of AI agents in healthcare.

    Journal of biomedical informatics·2026
    Same author

    Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More.

    Proceedings of machine learning research·2026

    Area of Science:

    • Computer Vision
    • Artificial Intelligence
    • Human-Computer Interaction

    Background:

    • Modern image understanding requires models to perform multiple tasks simultaneously, including object detection, scene recognition, and shape analysis.
    • Evaluating the contribution of each task to overall scene understanding is crucial for directing research efforts.

    Purpose of the Study:

    • To investigate the individual roles of various computer vision tasks in enhancing semantic segmentation, object detection, and scene recognition.
    • To quantify the potential for improvement in scene understanding by focusing on specific sub-tasks.

    Main Methods:

    • Development of a conditional random field (CRF) model framework.
    • Integration of human subjects as components within the CRF model to "plug-in" human intelligence for specific tasks.

    Related Experiment Videos

    Last Updated: Mar 28, 2026

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
    08:25

    Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment

    Published on: May 7, 2019

    9.7K
  • Comparative analysis of different hybrid human-machine CRF configurations.
  • Main Results:

    • Hybrid human-machine CRFs demonstrated varying levels of performance improvement based on the integrated task.
    • The study identified specific tasks that, when augmented with human input, yielded significant gains in scene understanding capabilities.
    • Quantifiable insights into the "head room" for improvement in semantic segmentation, object detection, and scene recognition were obtained.

    Conclusions:

    • Human-in-the-loop approaches within CRF frameworks offer a viable method for assessing the impact of individual tasks on scene understanding.
    • Research efforts focused on tasks with higher identified "head room" are likely to yield the most substantial improvements in AI-driven scene understanding.
    • This work provides a framework for systematically evaluating and enhancing AI models for complex visual perception tasks.