Related Experiment Video
Updated: Oct 1, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Benchmarking five large language models in medical genetics: a bilingual comparative evaluation using published and
Özge Beyza Gündoğdu Öğütlü1, Benjamin D Solomon2, Yusuf Selman Çelik3
1Department of Medical Genetics, Ankara Etlik City Hospital, Ankara, Türkiye.
Abstract:
This study asked whether five contemporary large language models answer medical genetics multiple-choice questions with equivalent accuracy on published versus novel items and across English and Turkish, and sought to characterize the errors that persist. Five models (GPT-5.2, Gemini 3 Pro, Claude Sonnet 4.6, Grok 4, and DeepSeek-V3.2) answered 100 four-option questions (50 from a published board review; 50 novel, expert-authored items absent from any database) in English and Turkish, yielding 1,000 responses. Correctness was modeled with item-clustered generalized estimating equation and Bayesian mixed-effects logistic regression (the latter as a prespecified sensitivity analysis); question provenance and language were tested for equivalence (item-clustered two one-sided tests, ±5-percentage-point margin), and the paired language effect with the McNemar test. Inter-model agreement and error concordance were examined. Overall accuracy was 97.7%. In the item-clustered GEE, only Gemini 3 Pro nominally exceeded the lowest-performing model; this imprecise contrast did not remain significant after Holm correction for the four secondary model comparisons, whereas the Bayesian sensitivity analysis additionally yielded a credible interval excluding 1 for GPT-5.2 versus Grok 4. Accuracy was statistically equivalent within the prespecified margin for published versus novel items (difference, -1.4 percentage points; item-clustered 90% CI, -4.1 to +1.3; item-clustered TOST P = 0.014) and for English versus Turkish (difference, -0.6 points; paired item-level TOST P < 0.001; McNemar P = 0.58). Inter-model agreement on the exact option chosen was high (Fleiss κ, 0.94-0.97). All five items with two or more errors failed concordantly, every erring model selecting the same wrong option. Beyond high accuracy, distinct model families converged on identical answers and, on the hardest items, on identical errors, a pattern consistent with shared learned associations and/or overlapping training data, although this behavioral comparison cannot identify the underlying mechanism. In an exploratory analysis, residual errors showed concentrated and concordant patterns, clustering in multistep Bayesian reasoning and evolving facts, suggesting that a second model may provide limited independent protection on such items. Performance on this restricted multiple-choice benchmark does not establish readiness for clinical genetic counseling, variant interpretation, or quantitative risk assessment, and human oversight remains essential.
