Related Experiment Video
Updated: Aug 5, 2026

Caffeine Extraction, Enzymatic Activity and Gene Expression of Caffeine Synthase from Plant Cell Suspensions
Published on: October 2, 2018
Rapid Analysis of Caffeine, Protein and Trigonelline in Ugandan Arabica Coffee Using NIRS and Machine Learning
Joseph Mbihayeimaana1,2, Jimcall Pfumorodze3, Ephraim Nuwamanya4
1School of Law, Africa University, 1 Fairview Road Off Nyanga Road, Mutare P.O. Box 1320, Zimbabwe.
None:
Coffee is a major export earner for Uganda, raking in over USD 2 billion in 2025. The global price of coffee is tagged to the perceived quality in the cup which in turn is affected by the chemical composition of the green bean. Breeding for market-preferred Arabica coffee varieties is a major objective of coffee breeding programs. Determination of coffee bean chemical constituents is routinely done through expensive, slow and tedious laboratory procedures, making it unsustainable of resource-limited public sector coffee breeding programs. Here, we demonstrate the use of near-infrared spectroscopy (NIRS) and the machine learning algorithms partial least squares (PLS), random forest (RF) and support vector machine (SVM) for the prediction of caffeine, protein and trigonelline in Arabica coffee. NIRS provides a fast, accurate and reliable method of simultaneously predicting multiple sample constituents. Ripe coffee cherries were picked from 172 farmers' fields, air dried in the laboratory at room temperature and processed to green beans. NIRS spectra were taken on the milled green bean at 400-2500 nm, with a 0.5 nanometer (nm) step. Reference data for caffeine, protein and trigonelline were collected on the same sample scanned with NIRS. A set of 12 spectral pretreatments were applied prior to making calibrations with the PLS, RF and SVM algorithms and 70% of the data as a training set and 30% as a test set. Caffeine content of reference samples ranged from 1.94-3.0 g/100 g, protein content ranged from 11.16-15.94% while trigonelline ranged from 0.94-1.23 g/100 g. The best calibrations for all algorithms and analytes were obtained using raw (untreated) spectra, which gave the same results as the Savitzky-Golay (SG) pretreatment. For caffeine, the best model (R2p = 0.89, RMSEP = 0.007, RPD = 3.34) was obtained with the SVM algorithm, while for protein, the best model (R2p = 0.98, RMSEP = 0.14, RPD = 6.92) was obtained using the PLS algorithm. Finally, for trigonelline, all three models had very high prediction accuracies (R2p = 0.98-0.99, RMSEP = 0.007-0.009, RPD = 8.53-10.52). Collectively, these results demonstrate the potential of using NIRS for rapid and simultaneous prediction of coffee green bean constituents to aid selection decisions.
More Related Videos
10:13Using Capillary Electrophoresis to Quantify Organic Acids from Plant Tissue: A Test Case Examining Coffea arabica Seeds
Published on: November 12, 2016
08:43PTR-ToF-MS Coupled with an Automated Sampling System and Tailored Data Analysis for Food Studies: Bioprocess Monitoring, Screening and Nose-space Analysis
Published on: May 11, 2017