A comparative study of vision-language models for food ingredient recognition and nutrient estimation

Shenglong Wang1, Guorui Sheng1, Hongfei Yan1

  • 1School of Computer Science and Artificial Intelligence, Ludong University, Yantai, 264025, China.

Summary

Vision-Language Models (VLMs) show promise for automated food analysis, improving ingredient recognition and nutrient estimation from images. However, accurately quantifying nutrients in complex dishes remains a challenge for AI.