Related Experiment Video
Updated: Apr 1, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Is one run enough? Reproducibility of flagship large language models across temperature and reasoning settings in
Paul Windisch1,2, Carole Koechli2, Fabio Dennstädt2
1Department of Radiation Oncology, Cantonal Hospital Winterthur, Winterthur, 8401, Switzerland.
Run-to-run reproducibility of large language models like Gemini and GPT-5.2 for biomedical classification is high. Single model runs are often sufficient for reliable results, with minimal replication serving as a practical stability check.
Area of Science:
- Artificial Intelligence
- Biomedical Informatics
- Natural Language Processing
Background:
- Assessing the run-to-run reproducibility of advanced AI models is crucial for reliable application in scientific research.
- Gemini 3 Flash Preview and GPT-5.2 were evaluated for their consistency in classifying trial success.
- The study investigated reproducibility across various operational parameters, including temperature and reasoning settings.
Purpose of the Study:
- To quantify the run-to-run reproducibility of Gemini 3 Flash Preview and GPT-5.2 for binary trial-success classification.
- To determine if single-run outputs are sufficient for accurate classification across different model configurations.
- To assess the impact of temperature and reasoning/thinking settings on model consistency.
Main Methods:
- Utilized 250 biomedical trial abstracts, pre-labeled for primary endpoint success.
- Evaluated Gemini across four thinking levels (minimal to high) and temperatures (0.0-2.0).
- Assessed GPT-5.2 across multiple reasoning-effort levels and conducted a temperature sweep with reasoning disabled; each setting was run three times.
Main Results:
- High run-to-run reproducibility was observed for both Gemini (κ = 0.942-1.000) and GPT-5.2 (κ = 0.984-0.995).
- Invalid output rates were minimal (0%-1.5% for Gemini, 0% for GPT-5.2).
- Classification performance (F1 score) remained stable (0.955-0.971), with minor improvements from majority voting.
Conclusions:
- Gemini and GPT-5.2 demonstrate high reproducibility in binary biomedical classification tasks with constrained outputs.
- Single-run model executions are frequently adequate for reliable classification.
- Minimal replication can serve as a practical method for verifying stability and ensuring result integrity.
Related Concept Videos
Improving Translational Accuracy
Improving Translational Accuracy
Leaky Scanning
Reasoning
Inductive reasoning involves deriving generalizations from specific observations. This type of reasoning helps form beliefs about the world. For example,...
Regulation of Expression Occurs at Multiple Steps
Transcription results in the generation of precursor (pre-mRNA) that consists of both exons and introns, which needs further processing before being translated to a...
Termination of Translation
