Related Experiment Video
Updated: Jan 17, 2026

Author Spotlight: Integrating Ultrasound Imaging with Biochemical Markers for Thyroid Disease Diagnosis
Published on: February 9, 2024
Prompt engineering and diagnostic accuracy of multimodal large language models in thyroid fine-needle aspiration
Bibhas Saha Dala1, Kaushik Mukhopadhyay2, Dwaipayan Roy3
1Department of Pathology, All India Institute of Medical Sciences (AIIMS), Kalyani, West Bengal, India.
Abstract:
Role of Large language models (LLMs) in fine-needle aspiration cytology (FNAC) image analysis remain uncertain. We evaluated two LLMs - Chat GPT-4o (OpenAI) and Claude 3.5 Sonnet (Anthropic) on 63 thyroid FNAC cases, each represented by eight microscopic images (Pap and MGG, 10x/40x), using generic and structured prompts. Structured prompts improved Bethesda concordance and near-match rates but inter-rater agreement remained poor (κ ≤ 0.09). Specificity reached 100% with structured prompts, but sensitivity dropped to ≤11.8% and misclassification persisted. LLMs show potential, but domain-specific training and validation are necessary for clinical use.

