Related Experiment Video
Updated: Jan 25, 2026

12:54
Vision Training Methods for Sports Concussion Mitigation and Management
Published on: May 5, 2015
18.0K
IPATH: A Large-Scale Pathology Image-Text Dataset from Instagram for Vision-Language Model Training
S Mirhosseini1, T Rai2,3, P Diaz-Santana4
1Centre for Vision, Speech and Signal Processing, University of Surrey, Guildford, GU2 7XH, UK. contact@erfan.uk.
Journal of Imaging Informatics in Medicine
|January 23, 2026
Summary
Researchers created the IPATH dataset from Instagram images to train AI for pathology. This AI model, IP-CLIP, shows strong diagnostic accuracy, improving medical image analysis.
Area of Science:
- Artificial Intelligence
- Computational Pathology
- Medical Informatics
Background:
- Artificial intelligence (AI) can detect subtle patterns in pathology images, enhancing diagnostic accuracy.
- A significant barrier to AI development in pathology is the scarcity of large, publicly available, annotated image datasets.
- Social media platforms offer a potential, yet unexplored, source for medical image data.
Purpose of the Study:
- To curate a novel dataset of pathology images from Instagram for AI model training.
- To develop and evaluate a multimodal AI model (IP-CLIP) using this curated dataset.
- To demonstrate the utility of social media data in advancing AI for medical image analysis.
Main Methods:
- Curated the IPATH dataset, comprising 45,609 pathology image-text pairs from Instagram, using automated classifiers, large language models, and manual filtering for quality control.
- Developed IP-CLIP by fine-tuning a pre-trained CLIP model on the IPATH dataset.
- Evaluated IP-CLIP's performance on seven external histopathology datasets using zero-shot classification and linear probing, and assessed image-text alignment via retrieval on a held-out subset.
Main Results:
- IP-CLIP consistently outperformed the original CLIP model on external datasets in zero-shot classification and linear probing tasks.
- IP-CLIP achieved performance comparable to or exceeding state-of-the-art pathology vision-language models, despite training on a smaller dataset.
- IP-CLIP demonstrated superior image-text alignment compared to CLIP and specialized models in retrieval tasks.
Conclusions:
- The IPATH dataset, sourced from Instagram, is a valuable resource for developing AI in medical image classification.
- Leveraging social media data can effectively address the scarcity of annotated medical images for AI training.
- The developed IP-CLIP model shows significant potential for enhancing diagnostic accuracy and decision support in pathology.
Related Concept Videos
Vision
59.5K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
59.5K
Language
898
Language is a unique communication system that uses words and systematic rules to organize and transmit information. Unlike other forms of communication, which may involve postures, movements, odors, or vocalizations, language relies on symbols and grammar. This makes human communication distinct from that of other species, who also communicate but do not use language in the same way humans do.
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
Corballis and Suddendorf (2007) and Tomasello and Rakoczy (2003) highlight the role of language in...
898
Color Vision
1.4K
Color perception begins in the retina, the light-sensitive layer at the back of the eye. Two main theories explain how colors are seen: the trichromatic theory and the opponent-process theory. The trichromatic theory, proposed by Thomas Young in 1802 and extended by Hermann von Helmholtz in 1852, suggests that color vision is based on three types of cone receptors in the retina. These cones are sensitive to different but overlapping ranges of wavelengths corresponding to red, blue, and green.
1.4K
Components of Language
802
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
802
Language Development
878
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
878
Language and Cognition
745
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
745

