Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Video

Updated: Mar 23, 2026

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
07:34

Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies

Published on: November 7, 2025

454

An Empirical Study Into Annotator Agreement, Ground Truth Estimation, and Algorithm Evaluation.

Thomas A Lampert, Andre Stumpf, Pierre Gancarski

    IEEE Transactions on Image Processing : a Publication of the IEEE Signal Processing Society
    |March 29, 2016
    PubMed
    Summary

    Annotator agreement significantly impacts segmentation algorithm evaluation. Algorithm rankings can be unreliable when using single-annotator ground truth (GT), especially with low agreement or few annotations.

    Related Concept Videos

    Common Leveling Mistakes and Errors01:17

    Common Leveling Mistakes and Errors

    580
    A survey team is tasked with determining the elevation difference between points Point A and Point B, separated by uneven terrain. They use a leveling instrument and a leveling rod.Common MistakesMisreading the Rod: During a backsight reading at Point A, the instrumentman observes the rod partially obscured by tall grass. Instead of reading 1.135 m, they mistakenly record 1.735 m due to the misalignment of the crosshair with the wrong graduation. This error adds 0.600 m to all subsequent...
    580

    You might also read

    Related Articles

    Articles linked to this work by shared authors, journal, and citation graph.

    Sort by
    Same author

    Combat or surveillance? Evaluation of the heterogeneous inflammatory breast cancer microenvironment.

    The Journal of pathology·2012
    Same author

    Discovering significant evolution patterns from satellite image time series.

    International journal of neural systems·2011
    See all related articles

    Area of Science:

    • Computer Vision
    • Image Processing
    • Machine Learning Evaluation

    Background:

    • Inter-annotator agreement in image annotation is crucial but often statistically analyzed.
    • Limited research quantifies the impact of annotator variance on foreground-background segmentation algorithm evaluation.
    • Ground truth (GT) is frequently derived from a single annotator, potentially biasing results.

    Purpose of the Study:

    • To quantify inter-annotator variance in image annotation.
    • To assess the effect of this variance on segmentation algorithm evaluation.
    • To investigate the reliability of different ground truth estimation methods.

    Main Methods:

    • A methodology was applied to four image-processing problems.
    • Inter-annotator variance was quantified.

    Related Experiment Videos

    Last Updated: Mar 23, 2026

    Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies
    07:34

    Utilizing vmTracking to Improve the Accuracy of Multi-Animal Pose Estimation in Rodent Social Behavior Studies

    Published on: November 7, 2025

    454
  • Automatic segmentation algorithms were compared against annotator agreement.
  • Various ground truth estimation methods (e.g., STAPLE, LSML, consensus voting) were analyzed.
  • Main Results:

    • Annotator agreement is notably low for detecting linear structures.
    • Algorithm performance ranking is highly dependent on the ground truth (GT) generation method.
    • Consensus voting tends to overestimate performance by accentuating obvious features.
    • GT estimation methods like STAPLE and LSML degrade with few or highly variable annotations.

    Conclusions:

    • Evaluating segmentation algorithms using a single annotator's ground truth (GT) can lead to unreliable rankings.
    • The choice of GT generation method significantly influences perceived algorithm performance.
    • Careful consideration of inter-annotator variance and GT estimation is necessary for robust algorithm evaluation.