Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Accuracy of Large Language Models in Selecting Appropriate Statistical Tests for Healthcare Research
Michael Paolella1, Aditya Tadinada1
1Oral and Maxillofacial Radiology, University of Connecticut Health, Farmington, USA.
Background:
Appropriate statistical test selection is essential for producing valid and reproducible findings in healthcare research. However, many investigators encounter difficulties when determining the most suitable statistical methods for their study design and dataset. Recently, artificial intelligence-based large language models (LLMs) have emerged as potential tools capable of assisting researchers with methodological decisions, including statistical analysis.
Objective:
The objective of this study was to evaluate whether artificial intelligence-based LLMs can accurately and reliably recommend appropriate statistical tests for healthcare research studies.
Methods:
Four LLM platforms (ChatGPT (OpenAI, San Francisco, California, USA), Google Gemini (Google LLC, Mountain View, California, USA), Microsoft Copilot (Microsoft Corporation, Redmond, Washington, USA), and Grok (xAI, Palo Alto, California, USA)) and two traditional search engines (Google (Google LLC, Mountain View, California, USA) and Bing (Microsoft Corporation, Redmond, Washington, USA)) were evaluated. Forty published research articles were selected across four common study designs: systematic reviews, randomized controlled trials, cohort studies, and case-control studies (10 articles per category). Each article's primary objective and statistical analysis were used to generate a standardized prompt describing the study scenario. These prompts were submitted to each LLM and search engine. Model responses were recorded and compared with the statistical tests used in the original studies. Accuracy was defined as agreement between the model's recommendation and the statistical test used in the published article.
Results:
LLMs demonstrated moderate-to-high accuracy in recommending appropriate statistical tests. Grok and Microsoft Copilot achieved the highest accuracy (34/40 correct recommendations), followed by Google Gemini (32/40) and ChatGPT (30/40). In contrast, traditional search engines (Google and Bing) did not provide statistical recommendations that matched the statistical tests used in the selected studies.
Conclusion:
LLMs demonstrated promising capability in identifying appropriate statistical tests for healthcare research scenarios and outperformed traditional search engines in this task. While these findings suggest that LLMs may serve as useful decision-support tools for researchers, their recommendations should be interpreted cautiously and verified by individuals with statistical expertise. Continued evaluation of AI-based tools will be necessary to ensure their responsible integration into scientific research workflows.
Related Concept Videos
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Comparing the Survival Analysis of Two or More Groups
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Statistical Methods for Analyzing Epidemiological Data
