Related Experiment Video
Updated: Jun 16, 2026

Systems Biology of Metabolic Regulation by Estrogen Receptor Signaling in Breast Cancer
Published on: March 17, 2016
A Reproducible Protocol and Framework for Large Language Model (LLM)-Assisted Estrogen and Progesterone Receptor
Svetoslav Bardarov1, Alireza Zarineh2
1Pathology and Laboratory Medicine, Richmond University Medical Center, New York City, USA.
None:
Estrogen and progesterone receptor (ER/PR) scoring in breast cancer is vulnerable to interobserver variability, particularly at diagnostic thresholds. While large language models (LLMs) offer potential as ancillary tools, no standardized deployment framework exists. This is a feasibility and protocol-development study that establishes a reproducible protocol for LLM implementation in pathology labs. Using College of American Pathologists (CAP) proficiency-testing tissue microarrays, we developed a systematic framework that combines recursive prompt engineering and zero-state reset (ZSR) protocols, requiring fresh chat sessions for each evaluation to prevent conversational bias. Three models were evaluated: Claude Haiku 4.5 (Anthropic, San Francisco, CA, USA), Gemini 3.0 (Google, Mountain View, CA, USA), and Gemma 3.0 12B (Google, Mountain View, CA, USA). The ZSR protocol proved essential, improving accuracy from a 70-80% baseline to ≥95% CAP concordance. Performance varied across models and scoring approaches: Claude achieved the highest concordance (90-98%), followed by Gemini (85-100%) and Gemma (73-93%), demonstrating that systematic prompt refinement, not model capacity, drives diagnostic accuracy. Final concordance exceeded 83% across all models. The locally hosted model approached cloud-level performance within a clinically meaningful range, achieving concordance rates that, while numerically lower, remained within acceptable feasibility thresholds for a protocol development context. We identified systematic error sources, including image artifacts, slide contamination, and compression effects. This protocol provides pathology laboratories with evidence-based guidance for implementing LLM-assisted ER/PR scoring without requiring specialized infrastructure or extensive computational resources. The framework is immediately actionable as a research and quality-assurance tool and provides a reproducible foundation for future clinical validation studies across laboratory settings.
