Related Experiment Video
Updated: Jan 18, 2026

11:34
Building Up a High-throughput Screening Platform to Assess the Heterogeneity of HER2 Gene Amplification in Breast Cancers
Published on: December 5, 2017
13.0K
Transformer-Based HER2 Scoring in Breast Cancer: Comparative Performance of a Foundation and a Lightweight Model
Yeh-Han Wang1,2, Min-Hsiang Chang3, Hsin-Hsiu Tsai4
1Department of Anatomical Pathology, Fu Jen Catholic University Hospital, Fu Jen Catholic University, New Taipei City 24352, Taiwan.
Diagnostics (Basel, Switzerland)
|September 13, 2025
Summary
Two artificial intelligence models achieved human-level accuracy for automated Human Epidermal Growth Factor 2 (HER2) scoring in breast cancer whole-slide images. These AI tools show promise for improving diagnostic consistency and supporting antibody-drug conjugate therapy decisions.
Area of Science:
- Computational pathology
- Artificial intelligence in oncology
- Biomarker quantification
Background:
- Accurate Human Epidermal Growth Factor 2 (HER2) scoring is crucial for breast cancer treatment selection, particularly for emerging antibody-drug conjugate (ADC) therapies targeting HER2-low tumors.
- Current inter-observer agreement in HER2 scoring, especially for borderline cases, presents diagnostic challenges.
- Artificial intelligence (AI) offers a potential solution for enhancing diagnostic consistency and scalability in HER2 scoring.
Purpose of the Study:
- To develop and compare two transformer-based AI models for automated HER2 scoring of breast cancer whole-slide images (WSIs).
- To evaluate the performance of these AI models against pathologist assessments and assess their utility in clinical decision-making tasks.
Main Methods:
- Adaptation of a large-scale foundation model (Virchow) and a lightweight model (TinyViT) for patch-level annotation and WSI scoring.
- Training and integration of both models into a WSI scoring pipeline.
- Performance evaluation on a clinical test set (n=66), including diagnostic accuracy, agreement with pathologists, and inference efficiency.
Main Results:
- Both AI models demonstrated substantial agreement with pathologist reports (Virchow: kappa=0.860, TinyViT: kappa=0.825).
- Virchow model exhibited slightly higher WSI-level accuracy, while TinyViT achieved a 60% reduction in inference time.
- AI models showed pathologist-comparable performance in binary clinical tasks, including identifying HER2-low tumors for ADC therapy.
- A continuous scoring framework revealed strong model correlation (r=0.995) and alignment with human assessments.
Conclusions:
- Transformer-based AI models achieve human-level accuracy for automated HER2 scoring, offering interpretable outputs.
- The lightweight TinyViT model presents practical advantages for clinical deployment due to its efficiency.
- Continuous HER2 quantification may offer more granular assessment, particularly in borderline cases, aligning with evolving ADC therapy indications.

