Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Outliers and Influential Points01:08

Outliers and Influential Points

6.8K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
6.8K
Data Collection by Observations01:08

Data Collection by Observations

15.8K
Data collection refers to a systematic way of obtaining, observing, measuring, and analyzing accurate information. Observational studies are one of the most widely used methods of data collection. It involves collecting data by observing the behavior and physical characteristics of a sample without making any modifications to the sample.
An astronomer viewing the motion and brightness of stars in the sky and recording the data is an example of observational data collection. A botanist recording...
15.8K
Statgraphics01:10

Statgraphics

496
Statgraphics is a comprehensive statistical software suite designed for both basic and advanced data analysis. Originating in 1980 at Princeton University under Dr. Neil W. Polhemus, it was one of the pioneering tools for statistical computing on personal computers, with its public release in 1982 marking an early milestone in data science software. Over the years, it has evolved into a robust platform for data science, offering tools for regression analysis, ANOVA, multivariate statistics,...
496
Data Collection III01:05

Data Collection III

4.8K
The physical assessment examines the patient for objective data that defines the patient's condition, and aids in formulating the nursing care plan. The purpose of physical assessment is a health status appraisal, which includes identifying health problems, and establishing a database for nursing intervention.
The principles to begin the physical assessment include conducting a comprehensive or problem-related history in a quiet, well-lit room, emphasizing privacy and comfort for the...
4.8K
Data Collection II01:29

Data Collection II

10.5K
The nursing history captures and records the patient's health status, so that a care plan evolves to meet the patient's individual needs. The nursing health history is a part of the initial assessment. A comprehensive history covers all health dimensions and plays a significant role in the assessment process. A comprehensive history includes the patient's biographical information, reasons for seeking health care, expectations, present and past health history, medications, and...
10.5K
Data Collection by Experiments01:13

Data Collection by Experiments

28.2K
Data collection is a systematic method of obtaining, observing, measuring, and analyzing accurate information. An experimental study is a standard method of data collection that involves the manipulation of the samples by applying some form of treatment prior to data collection. It refers to manipulating one variable to determine its changes on another variable. The sample subjected to treatment is known as “experimental units.”
An example of the experimental method is a public...
28.2K

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

How to benchmark medical AI agents.

PLoS medicine·2026
Same author

AI-based selection of tumor regions for genomic profiling in neuropathology.

Neuro-oncology advances·2026
Same author

Mutation enrichment in targeted panels flags immunotherapy-responsive POLE-driven hypermutated microsatellite-stable colorectal cancers.

NPJ precision oncology·2026
Same author

Clinical decision support in hematological malignancies using a case-grounded AI agent.

Nature medicine·2026
Same author

A deep learning framework for efficient pathology image analysis.

Nature communications·2026
Same author

Functional Outcome Prediction in Young Adults With Mental Health Symptoms Using Machine Learning and Large Language Models: Longitudinal Observational Study.

JMIR mental health·2026

Related Experiment Video

Updated: Apr 17, 2026

A User-friendly and Powerful R Analysis of Large-scale Datasets
10:56

A User-friendly and Powerful R Analysis of Large-scale Datasets

Published on: November 4, 2025

471

HISTAI: a valuable dataset with a valuable lesson.

Katherine J Hewitt1, Nic G Reitsam1,2,3, Sebastian Foersch4

  • 1Else Kroener Fresenius Center for Digital Health, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology, Dresden, Germany.

The Journal of Pathology. Clinical Research
|April 16, 2026
PubMed
Summary

A pathologist review of the HISTAI dataset found significant issues with data accuracy and completeness. This highlights the need for careful validation of whole slide image datasets for artificial intelligence in pathology.

Keywords:
HISTAIcomputational pathologydata quality and reproducibilitydiagnostic concordanceground truth validationopen‐access datasets

More Related Videos

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
13:44

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques

Published on: December 9, 2022

4.6K
Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
09:43

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering

Published on: November 22, 2019

6.9K

Related Experiment Videos

Last Updated: Apr 17, 2026

A User-friendly and Powerful R Analysis of Large-scale Datasets
10:56

A User-friendly and Powerful R Analysis of Large-scale Datasets

Published on: November 4, 2025

471
Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques
13:44

Project-Based Learning Guidelines for Health Sciences Students: An Analysis with Data Mining and Qualitative Techniques

Published on: December 9, 2022

4.6K
Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering
09:43

Databases to Efficiently Manage Medium Sized, Low Velocity, Multidimensional Data in Tissue Engineering

Published on: November 22, 2019

6.9K

Area of Science:

  • Computational pathology
  • Artificial intelligence in medicine
  • Digital pathology

Background:

  • Artificial intelligence (AI) in pathology requires high-quality, validated whole slide image (WSI) datasets.
  • Existing WSI datasets are often scarce, lack diversity, and are not well-validated, limiting AI progress.
  • The HISTAI resource offers a large, open-source collection of WSIs with clinical metadata.

Purpose of the Study:

  • To conduct a pathologist-led evaluation of the HISTAI dataset's label accuracy, metadata completeness, and composition.
  • To identify limitations and potential challenges for using the HISTAI resource in AI development.
  • To assess the clinical reliability of the HISTAI dataset for computational pathology applications.

Main Methods:

  • Pathologist review of 328 selected cases from the HISTAI resource.
  • Analysis of label accuracy, metadata completeness (demographics, specialty), and dataset composition.
  • Focused review of diagnostic conclusion concordance, molecular annotation, and adherence to WHO CNS5 criteria for specific tumor types.

Main Results:

  • Identified fewer unique cases than reported, with incomplete demographic data (55%).
  • Uneven dataset composition (dermatopathology 47.1%, gastrointestinal 24.0%) with poorly reported specialties.
  • Significant discrepancies between diagnosis and conclusion fields (20.7% concordance, 27.1% conflict), ambiguous conclusions (30.3%), incomplete molecular data (18.9%), and non-compliance with WHO CNS5 criteria for gliomas.

Conclusions:

  • The HISTAI dataset exhibits substantial ambiguities in ground-truth labeling and incomplete molecular annotation.
  • Limited documentation of dataset provenance and ethical oversight requires attention.
  • Effective and responsible use of HISTAI necessitates rigorous clinical validation and pathologist-AI researcher collaboration.