Related Experiment Video
Updated: Jun 16, 2026

10:36
Systems Biology of Metabolic Regulation by Estrogen Receptor Signaling in Breast Cancer
Published on: March 17, 2016
A Reproducible Protocol and Framework for Large Language Model (LLM)-Assisted Estrogen and Progesterone Receptor
Svetoslav Bardarov1, Alireza Zarineh2
1Pathology and Laboratory Medicine, Richmond University Medical Center, New York City, USA.
Cureus
|June 15, 2026
Summary
This study developed a reproducible protocol for using large language models (LLMs) to improve estrogen and progesterone receptor (ER/PR) scoring in breast cancer diagnostics, significantly reducing interobserver variability.
Area of Science:
- Pathology and Laboratory Medicine
- Artificial Intelligence in Healthcare
- Computational Biology
Background:
- Estrogen and progesterone receptor (ER/PR) scoring in breast cancer is crucial for treatment decisions but suffers from significant interobserver variability.
- Large Language Models (LLMs) present a potential solution for standardizing ER/PR scoring, but a deployment framework is lacking.
Purpose of the Study:
- To develop and validate a reproducible protocol for implementing LLM-assisted ER/PR scoring in pathology laboratories.
- To assess the feasibility and accuracy of different LLMs in ER/PR scoring using a standardized framework.
Main Methods:
- A systematic framework combining recursive prompt engineering and zero-state reset (ZSR) protocols was developed using College of American Pathologists (CAP) proficiency-testing tissue microarrays.
- Three LLMs (Claude Haiku 4.5, Gemini 3.0, Gemma 3.0 12B) were evaluated under the ZSR protocol to prevent conversational bias and ensure reproducible results.
- Performance was measured by concordance rates with established CAP scoring standards.
Main Results:
- The ZSR protocol was essential, improving accuracy from a 70-80% baseline to ≥95% CAP concordance.
- Claude Haiku 4.5 achieved the highest concordance (90-98%), followed by Gemini 3.0 (85-100%) and Gemma 3.0 12B (73-93%).
- Systematic prompt refinement, rather than model capacity, was identified as the primary driver of diagnostic accuracy, with overall concordance exceeding 83% across models.
Conclusions:
- A reproducible protocol for LLM-assisted ER/PR scoring has been established, demonstrating feasibility for pathology labs without specialized infrastructure.
- The developed framework offers an actionable tool for research and quality assurance, providing a foundation for future clinical validation of AI in diagnostic pathology.
- Systematic error sources were identified, including image artifacts and compression effects, highlighting areas for future protocol refinement.
