Related Experiment Video
Updated: Aug 5, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating the Accuracy of Large Language Models in Selecting Appropriate Statistical Tests for Healthcare Research
Michael Paolella1, Aditya Tadinada1
1Oral and Maxillofacial Radiology, University of Connecticut Health, Farmington, USA.
Artificial intelligence (AI) large language models (LLMs) show promise in recommending statistical tests for healthcare research, outperforming traditional search engines. Researchers should use LLM recommendations cautiously and verify them with statistical experts.
Area of Science:
- Healthcare Research Methodology
- Artificial Intelligence in Science
- Statistical Analysis
Background:
- Accurate statistical test selection is crucial for valid and reproducible healthcare research findings.
- Researchers often face challenges in choosing appropriate statistical methods for their study designs and data.
- Large language models (LLMs) are emerging as potential tools to aid researchers in methodological decisions, including statistical analysis.
Purpose of the Study:
- To evaluate the accuracy and reliability of AI-based LLMs in recommending statistical tests for healthcare research.
- To compare the performance of LLMs against traditional search engines in this task.
Main Methods:
- Four LLM platforms (ChatGPT, Google Gemini, Microsoft Copilot, Grok) and two search engines (Google, Bing) were assessed.
- Forty published research articles across four study designs (systematic reviews, RCTs, cohort, case-control) were analyzed.
- Standardized prompts based on study objectives and statistical analyses were used to query the AI models and search engines.
Main Results:
- LLMs demonstrated moderate-to-high accuracy in recommending statistical tests.
- Grok and Microsoft Copilot showed the highest accuracy (34/40), followed by Google Gemini (32/40) and ChatGPT (30/40).
- Traditional search engines (Google, Bing) failed to provide accurate statistical test recommendations matching the studies.
Conclusions:
- LLMs show significant potential as decision-support tools for selecting appropriate statistical tests in healthcare research.
- LLMs outperformed traditional search engines in accurately recommending statistical tests.
- LLM recommendations require cautious interpretation and verification by statistical experts for responsible integration into research workflows.
Related Concept Videos
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5% chance...
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Comparing the Survival Analysis of Two or More Groups
Testing a Claim about Standard Deviation
The hypothesis testing for the claim of population standard deviation (or variance) requires the data and samples to be random and unbiased. The population distribution also must be normal. There is no specific requirement on the sample size as the estimation is based on the chi-square distribution.
As a first step, the hypothesis (null and alternative) concerning the claim about...
Statistical Methods for Analyzing Epidemiological Data
