Related Experiment Video
Updated: Jan 16, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
Are answers obtained from artificial intelligence models for information purposes repeatable?
Yasemin Tunca1, Volkan Kaplan2, Murat Tunca1
1Department of Orthodontics, Faculty of Dentistry, Kutahya Health Sciences University, Kutahya, Turkey.
International Orthodontics
|October 4, 2025
Summary
Large language models show varied repeatability in orthodontic answers. While accurate, some AI models lack temporal consistency, impacting clinical use.
Area of Science:
- Artificial Intelligence in Dentistry
- Natural Language Processing
- Clinical Decision Support
Background:
- Assessing the reliability of AI-generated orthodontic information is crucial for clinical integration.
- Large language models (LLMs) are increasingly used, necessitating evaluation of their consistency.
Purpose of the Study:
- To evaluate the repeatability of orthodontic responses from multiple LLMs over time.
- To compare the temporal stability of ChatGPT-3.5, ChatGPT-4.0, Gemini, and Gemini-Advanced.
Main Methods:
- Four LLMs (ChatGPT-3.5, ChatGPT-4.0, Gemini, Gemini-Advanced) answered 40 orthodontic questions at three time points.
- Responses were independently assessed by two orthodontic experts using a 3-point accuracy scale.
- Inter-rater agreement (Cohen's Kappa) and model repeatability (ICC) were calculated; temporal differences were analyzed using Friedman and Spearman tests.
Main Results:
- Substantial inter-rater agreement was observed (Kappa: 0.624-0.749).
- Repeatability varied significantly, with ICC values ranging from 0.666 (Gemini) to 0.960 (ChatGPT-3.5).
- Significant differences in model accuracy were found over time (P<0.001), with weak positive correlations between time points (ρ=0.284-0.383).
Conclusions:
- Statistically significant differences exist in the temporal repeatability of AI orthodontic responses.
- Some LLMs demonstrate high accuracy but limited consistency over time.
- Evaluating both accuracy and temporal stability is essential for clinical AI implementation in orthodontics.
Related Concept Videos
Measures of Intelligence
8.3K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
8.3K
Statistical Analysis: Overview
14.5K
When we take repeated measurements on the same or replicated samples, we will observe inconsistencies in the magnitude. These inconsistencies are called errors. To categorize and characterize these results and their errors, the researcher can use statistical analysis to determine the quality of the measurements and/or suitability of the methods.
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
14.5K
Non-equilibrium in the Cell
5.3K
An important concept in studying metabolism and energy is that of chemical equilibrium. Most chemical reactions are reversible. They can proceed in both directions, releasing energy into their environment in one direction, and absorbing it from the environment in the other direction. The same is true for the chemical reactions involved in cell metabolism, such as the breaking down and building up of proteins into and from individual amino acids, respectively. Reactants within a closed system...
5.3K
Accuracy, limits, and approximation
1.1K
Accuracy, limits, and approximations are common in many fields, especially in engineering calculations. These concepts are imperative for ensuring that a given value is as close as possible to its true value.
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
1.1K
Accuracy and Precision
14.0K
Scientists typically make repeated measurements of a quantity to ensure the quality of their findings and to evaluate both the precision and the accuracy of their results. Measurements are said to be precise if they yield very similar results when repeated in the same manner. A measurement is considered accurate if it yields a result that is very close to the true or the accepted value. Precise values agree with each other; accurate values agree with a true value. Highly accurate...
14.0K
Stereotype Content Model
15.3K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.3K