Estimating Nutrient Composition of Packaged Foods Using Natural Language Processing and Optimization Modeling
Mélina Côté1,2, Wiam Hamadi3, Catherine Laramée1
1Centre Nutrition, santé et société (NUTRISS), Institut sur la nutrition et les aliments fonctionnels (INAF), Université Laval, Québec, QC, Canada.
Background:
Food composition databases are fundamental for rigorous dietary assessment, yet they often include information only for generic foods.
Objectives:
This study aimed to estimate the full nutrient composition of packaged foods using natural language processing (NLP) and optimization modeling.
Methods:
Nutrition Facts tables (NFTs) and ingredient lists for 5371 packaged foods collected by the food quality observatory across 17 food categories available in Québec, Canada, were used. First, an NLP algorithm matched individual ingredients from packaged foods to the closest equivalents in the Canadian Nutrient File 2015, which contains full nutrient profiles for over 5690 ingredients and foods in Canada. Match quality was assessed using cosine similarity scores. Second, an optimization model estimated the proportion of all ingredients (grams per 100 g) from the packaged foods, enabling the reverse-engineering of nutrient composition data found on the NFT. Model performance was assessed using relative errors comparing estimated with known nutrient values reported on NFTs.
Results:
Over 55% of ingredients were matched to the Canadian Nutrient File with cosine similarity scores ≥0.9, indicating high-quality matches. Across all food categories combined, the median relative error for the estimates of energy and the 10 nutrients reported on NFTs was <|20%|, consistent with Health Canada's accepted variance for NFT declarations, suggesting reliable estimations. Six food categories showed strong results, with all nutrient estimates having median relative errors <|20%|. Eight food categories obtained moderate results, with all nutrient estimates having median relative errors <|20%|, but with a broader range of error values. Three food categories obtained poor results, with several nutrient estimates having relative errors beyond the |20%| threshold.
Conclusions:
A method based on NLP and optimization modeling can reliably estimate ingredient proportions of a wide variety of packaged foods, allowing for the generation of complete nutrient profiles.
Related Concept Videos
Key Elements for Plant Nutrition
Optimal Foraging
Methods of Medium Optimization
Predicting Products: Substitution vs. Elimination
The following factors can influence the mechanisms competing against each other:
Estimation of the Physical Quantities
Model-Independent Approaches for Pharmacokinetic Data: Noncompartmental Analysis
One important characteristic of noncompartmental analyses is that drug exposure increases proportionally with increasing doses. This relationship...

