Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Accuracy and Explanatory Quality of Large Language Models ChatGPT, Claude, DeepSeek, Gemini, Grok, and
Mukesh Shukla1, Deepshikha Pandey2, Samarjeet Kaur3
1Community Medicine, All India Institute of Medical Sciences, Raebareli, Raebareli, IND.
Abstract:
Background Large language models (LLMs) are increasingly integrated into academic and professional research workflows, yet their capability to accurately select appropriate statistical tests for hypothesis testing remains underexplored. Incorrect statistical test selection can lead to invalid conclusions and compromise scientific validity, making this evaluation critical for determining the reliability of LLMs in research applications. The study objective was to evaluate and compare the accuracy of six prominent LLMs (ChatGPT, Claude, DeepSeek, Gemini, Grok, and Le Chat) in selecting appropriate statistical tests for various hypothesis testing scenarios. Materials and methods A comparative, cross-sectional evaluation was conducted using 20 standardized statistical testing scenarios. Each scenario was designed to cover 20 different hypothesis testing situations, including comparisons of means, proportions, non-parametric alternatives, paired versus independent samples, and correlation and regression analyses. All models were prompted with identical instructions and evaluated by five independent experts with profound knowledge in biostatistics. Responses were assessed for accuracy and rated on five domains (clarity and accessibility, identification of necessary assumptions, pedagogical value, problem-solving approach, and statistical reasoning) using a five-point Likert scale. Analysis of Variance (ANOVA) was applied for between-group comparisons, and a p<0.05 was considered significant. Results All six LLMs achieved 100% accuracy in statistical test selection across all 20 hypothesis scenarios. However, significant variations emerged in explanatory quality. Claude demonstrated superior performance in clarity and accessibility (4.65 ± 0.58, p=0.05), while the problem-solving approach showed the most consistent excellence across models. Statistical reasoning exhibited variation ranging from 3.16 to 4.66, with complex regression methods receiving lower ratings than basic statistical tests. Gemini excelled in pedagogical value (4.50 ± 0.68), while ChatGPT ranked lowest in statistical reasoning despite strong problem-solving capabilities. Conclusions All LLMs demonstrate perfect accuracy in statistical test selection; however, differences exist in the quality of explanations and reasoning provided. These findings suggest that current-generation LLMs have become dependable tools for statistical consultation in hypothesis testing scenarios. However, users should consider model-specific strengths when seeking detailed explanations or educational content.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
07:31Defining the Role Of Language in Infants' Object Categorization with Eye-tracking Paradigms
Published on: February 8, 2019
Related Concept Videos
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
Statistical Hypothesis Testing
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Errors In Hypothesis Tests
Decision Making: Traditional Method
First, a specific claim about the population parameter is decided based on the research question and is stated in a simple form. Further, an opposing statement to this claim is also stated. These statements can act as null and alternative hypotheses, out of which a null hypothesis would be a...
Types of Hypothesis Testing
When the null and alternative hypotheses are stated, it is observed that the null hypothesis is a neutral statement against which the alternative hypothesis is tested. The alternative hypothesis is a claim that instead has a certain direction. If the null hypothesis claims that p = 0.5, the alternative hypothesis would be an opposing statement to this and can be put either p > 0.5, p < 0.5, or p...