Related Experiment Video
Updated: Apr 23, 2026

Deep Neural Networks for Image-Based Dietary Assessment
Published on: March 13, 2021
A comparative study of vision-language models for food ingredient recognition and nutrient estimation
Shenglong Wang1, Guorui Sheng1, Hongfei Yan1
1School of Computer Science and Artificial Intelligence, Ludong University, Yantai, 264025, China.
Abstract:
The accurate assessment of food composition is essential to understanding its nutritional and sensory properties. Traditional dietary assessment methods are often constrained by subjective input and low reproducibility. This study explores the use of Vision-Language Models (VLMs) for automated food composition analysis, focusing on two key tasks: food ingredient recognition and nutrient estimation. We evaluated state-of-the-art VLMs using the Nutrition5K dataset, which contains real-world food images with ingredient-level annotations. To improve model sensitivity to complex food structures, we introduce a progressive multi-view image recognition approach that enhances ingredient recognition. We also propose a prompting strategy using ingredient labels to guide nutrient estimation. Results show that while most VLMs effectively identify primary food components, challenges persist in quantifying nutrient contents, particularly for composite or visually ambiguous dishes. Our findings highlight the promise and limitations of AI-assisted food composition analysis and offer insights for future methods integrating chemical, visual, and computational perspectives.
More Related Videos
Related Concept Videos
Key Elements for Plant Nutrition
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.

