Related Experiment Video
Updated: Feb 20, 2026

Author Spotlight: Advancements in 3D Optical Imaging for Comprehensive Body Composition Assessment in Modern Research
Published on: June 7, 2024
Pose estimation based on keypoints and monocular depth estimation for predicting cattle body weight and hip height
Guilherme L Menezes1, Alyssa Seitz1, Enrico Casella2
1Department of Animal and Dairy Sciences, University of Wisconsin-Madison, Madison, WI 53703, United States.
Abstract:
Computer vision systems (CVS) have been developed using either top-down view 3D images or side-view 2D images to predict body weight (BW) and hip height (HH). However, 3D imaging systems are often costly compared with 2D imagery setups, and side-view cameras are often difficult to deploy under commercial farm conditions due to occlusion and variation in animal distance and posture. Herein, pose estimation using top-down view 2D images offers a promising approach to automatically extract keypoint-based features that describe body biometrics and could provide information correlated with BW and HH. Additionally, the same 2D images could be used to generate depth information, such as volume and height, using monocular depth estimation (MDE). Therefore, this study aimed to 1) develop predictive models for BW and HH based on features extracted from body pose keypoints in 2D infrared images and depth images generated using MDE, and 2) compare these models with those using features extracted from depth images collected by a 3D imaging system. A total of 395 top-down view videos from 94 beef-on-dairy crossbred cattle across four experimental blocks were collected using infrared and depth sensors. BW was recorded using an electronic scale, and HH was manually measured using a measuring stick. A pose estimation model identified seven anatomical landmarks (ie keypoints). The same 2D infrared images were converted into 3D images using zero-shot MDE, and a pipeline extracted features including volume, area, circularity, eccentricity, as well as back heights and widths. Depth images from a 3D imaging system were processed using the same pipeline. Random Forest (RF), Partial Least Squares Regression (PLS), and Support Vector Regression (SVM) models were evaluated using a leave-one-block-out cross-validation approach. The PLS model using Euclidean distances between the keypoints as features achieved R2 values of 0.90 with a Root Mean Square Error (RMSE) of 33.1 kg. Using MDE-derived depth features, PLS achieved an R2 of 0.95 with an RMSE of 24.2 kg. For HH, PLS using keypoints achieved an R2 of 0.77 with RMSE of 3.2 cm, and models using MDE-derived depth features showed similar performance. Our findings demonstrate that biometric features extracted from top-down 2D images or MDE-derived depth features enable comparable predictive performance for BW and HH, with models using features extracted from depth images collected using a 3D imaging system.
More Related Videos
06:52An Inertial Measurement Unit Based Method to Estimate Hip and Knee Joint Kinematics in Team Sport Athletes on the Field
Published on: May 26, 2020
14:14Quantification of Strain in a Porcine Model of Skin Expansion Using Multi-View Stereo and Isogeometric Kinematics
Published on: April 16, 2017
Related Concept Videos
Estimation of the Physical Quantities
Depth Perception and Spatial Vision
What are Estimates?
The estimate for the mean of a sample is denoted by ͞x, whereas the mean of the population is designated as μ. Further, parameters such...