Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Experiment Videos

Clinical Artificial Intelligence and Machine Learning Metrics 101: How to Evaluate Artificial Intelligence Tools for

Brian Hur1, Aparna Elangovan2, Laura Hardefeldt3

  • 1Veterinary Information Network, 777 W. Covell, Davis, CA, USA; Department of Veterinary Biosciences, Melbourne Veterinary School, University of Melbourne, Melbourne, Victoria, Australia.

The Veterinary Clinics of North America. Small Animal Practice
|May 7, 2026
PubMed
Summary

Related Concept Videos

Measures of Intelligence01:29

Measures of Intelligence

Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this; it...

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Antimicrobial prescribing in dogs and cats with urinary tract disease in a prospective intervention trial.

Journal of veterinary internal medicine·2026
Same author

Retrospective cohort study on the development of keratoconjunctivitis sicca in dogs treated with trimethoprim sulfonamide: a VetCompass Australia study.

Journal of veterinary internal medicine·2026
Same author

Tailoring task arithmetic to address bias in models trained on multi-institutional datasets.

Journal of biomedical informatics·2025
Same author

Using natural language processing and patient journey clustering for temporal phenotyping of antimicrobial therapies for cat bite abscesses.

Preventive veterinary medicine·2024
Same author

Cross-sectional evaluation of a large-scale antimicrobial stewardship trial in Australian companion animal practices.

The Veterinary record·2023
Same author

A Comparative Review of Pregnancy and Cancer and Their Association with Endoplasmic Reticulum Aminopeptidase 1 and 2.

International journal of molecular sciences·2023

Veterinary professionals need clear guidelines for evaluating artificial intelligence (AI) tools. Understanding key AI performance metrics beyond simple accuracy is crucial for informed adoption in veterinary medicine.

Area of Science:

  • Veterinary Medicine
  • Artificial Intelligence
  • Biomedical Informatics

Background:

  • Artificial intelligence (AI) tools are increasingly integrated into veterinary practice.
  • Clinicians currently lack standardized frameworks for assessing AI tool performance.
  • Evaluating AI requires understanding metrics beyond basic accuracy.

Purpose of the Study:

  • To provide a practical guide for veterinary professionals on evaluating AI performance.
  • To explain essential AI evaluation metrics and their significance.
  • To emphasize the need for evaluation literacy in AI adoption.

Main Methods:

  • Discussion of key AI evaluation metrics: sensitivity, specificity, precision, recall, and Area Under the Receiver Operating Characteristic Curve (AUC).
Keywords:
Artificial intelligenceClinical decision supportDiagnostic AIExternal validationMachine learningModel validationPerformance metricsVeterinary medicine

Related Experiment Videos

  • Explanation of the importance of interannotator agreement for setting performance benchmarks.
  • Highlighting the necessity of external validation and modality-specific considerations (AI scribes, digital imaging, pathology).
  • Main Results:

    • Accuracy alone is an insufficient metric for evaluating AI in veterinary medicine.
    • Interannotator agreement is vital for establishing realistic performance ceilings.
    • External validation and modality-specific metrics are critical for robust AI assessment.

    Conclusions:

    • Veterinary professionals must develop AI evaluation literacy due to the absence of regulatory oversight.
    • Informed decision-making regarding AI adoption depends on a thorough understanding of performance metrics.
    • Standardized evaluation frameworks are needed to ensure the effective and safe integration of AI in veterinary care.