Related Experiment Video
Updated: Jan 11, 2026

Studying Food Reward and Motivation in Humans
Published on: March 19, 2014
Population-Level Analysis of Personalized Food Recommendation Using Reinforcement Learning
Yone Tellechea1, Markel Arrojo1, Ander Cejudo1,2
1Vicomtech Foundation, Basque Research and Technology Alliance (BRTA), Mikeletegi 57, 20009 Donostia-San Sebastián, Spain.
Abstract:
This paper introduces an innovative methodology for optimizing recommendation strategies across different populations within the food industry. While previous approaches to recommending courses have overlooked cultural and age-based preferences, our work demonstrates how understanding these differences can significantly enhance the attractiveness for consumers and create new opportunities for marketing. By simulating diverse populations using a fuzzy logic approach, based on individual characteristics such as age, gender, geographical area, and city size, the study evaluates how recommendation algorithms perform within a generated menu database. Results show that algorithms like State-Action-Reward-State-Action (SARSA), multi-armed bandit (MAB), and Deep-Q Network (DQN) exhibit varying levels of efficiency depending on the population. Notably, the DQN improves accumulated reward over a random recommender by 71.60% for "Foodies", 65.02% for "Veggies", 63.46% for "Spanish", and 8.89% for "Seniors", while MAB achieves similar performance with fewer resources. Statistically significant differences (p < 0.005) are found in the performance of the DQN between populations, with large effect sizes according to Cliff's delta. These findings highlight recommender systems as an opportunity to navigate market demand, optimize supply chains, and reduce food waste. A better understanding of public preferences enables more effective alignment of supply and demand across the entire food supply chain. As a conclusion, while the DQN effectively captures target group preferences, the optimum recommendation strategy should be chosen by balancing algorithmic performance, computational efficiency, and the specific requirements of the food sector.
Related Concept Videos
Optimal Foraging
Law of Effect
Edward Thorndike's foundational work involved studying learning in animals, particularly using puzzle...
Reinforcement
Positive reinforcement occurs when a behavior is followed by the presentation of a rewarding stimulus, increasing the frequency of that behavior. For example:
Observational Learning
Reinforcement Schedules
Once a behavior is learned,...
Avoidance Learning and Learned Helplessness
Avoidance learning occurs when an organism learns that a specific behavior can prevent an unpleasant outcome. For example, a student who receives a bad grade may start studying harder to avoid future poor grades. This behavior persists even when the negative outcome is no longer present. Avoidance learning is powerful because it maintains behavior in the absence of the...

