Related Experiment Video
Updated: Sep 6, 2026

Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
Visual Food Object Detection via Open-Vocabulary Learning
Shoulong Liu1, Guorui Sheng1, Shaojie Zhou1
1School of Computer Science and Artificial Intelligence, Ludong University, Yantai, China.
Abstract:
The rapid diversification of food products, frequent packaging updates, and seasonal variations pose significant challenges for vision-based food inspection and dietary monitoring in real-world food systems. To address the need for scalable and flexible food analysis, this study proposes a domain-adaptive open-vocabulary food object detection (OVFD) framework that enables the identification of both existing and previously unseen food categories without category-specific box annotations for new items. By leveraging open-vocabulary representations, the framework allows dynamic expansion of detectable food categories as new items emerge, supporting continuous adaptation to evolving products, recipes, and dietary patterns while reducing manual annotation requirements. The proposed method integrates a dynamic prompt distribution network (DPDN) to improve region-text alignment, along with a density-aware multiple instance learning (DA-MIL) strategy that uses weakly supervised image-level tags to improve generalization and suppress background-driven detections. To reduce semantic leakage, Food2K labels are filtered by normalized matching, synonym auditing, and overlap-based rules before training. The framework was evaluated on multiple benchmark food datasets under open-vocabulary settings, demonstrating consistent improvements in detection performance for both Base (seen) and Novel (unseen) food categories, with Novel-category gains also observed under AP75 (average precision at intersection-over-union = 0.75) and COCO-style AP. Repeated-run reporting, error analysis, and limited external evaluation support stability and practical plausibility. By enabling dynamic category expansion with reduced annotation overhead, the proposed approach provides an assistive localization module for downstream dietary assessment, sorting review, and quality-inspection workflows, while task-specific validation and human verification remain necessary before operational food-safety or production-line use. PRACTICAL APPLICATIONS: This study provides an OVFD approach that can be adapted to new food items with reduced category-specific box annotation. Bounding-box localization can support upstream perception for portion estimation, ingredient-level dietary logging, assisted sorting review, and quality-inspection workflows by identifying where food items are located before secondary human or automated assessment. The present evidence supports assistive use only; human verification and validation under pilot-plant or real production conditions remain necessary before operational use. The detector is therefore intended to provide spatial evidence for subsequent review, not to replace validated food-inspection, quality-assessment, or production-control procedures.
Related Concept Videos
Observational Learning
Light Acquisition
Vision
Force Classification
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Visual Agnosia
Introduction to Learning
In contrast to learned behaviors, unlearned behaviors such as crying, sexual...
