Related Experiment Video
Updated: Sep 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
ChatGPT's performance in sample size estimation: a preliminary study on the capabilities of artificial intelligence
1University Institute for Primary Care (IuMFE), University of Geneva, 1211 Geneva, Switzerland.
Large language models like ChatGPT offer accurate sample size estimations for research, with ChatGPT-4o showing improved precision over ChatGPT-4.0. Expert validation remains crucial due to observed inconsistencies.
Area of Science:
- Clinical Research Methodology
- Statistical Computing
- Artificial Intelligence in Healthcare
Background:
- Artificial intelligence (AI) and large language models (LLMs) like ChatGPT are increasingly used in clinical research.
- The utility of these AI tools for specialized statistical tasks, such as sample size estimation, is not well-understood.
Purpose of the Study:
- To evaluate the accuracy and reproducibility of ChatGPT-4.0 and ChatGPT-4o in performing sample size estimations.
- To compare the performance of different versions of ChatGPT for statistical calculations in research.
Main Methods:
- Sample size calculations were performed for 24 standard statistical scenarios using ChatGPT-4.0 and ChatGPT-4o.
- Accuracy was measured by absolute percentage error against reference values.
- Reproducibility was assessed by comparing results from independent chat sessions.
Main Results:
- Both ChatGPT-4.0 and ChatGPT-4o provided reasonably accurate sample size estimates, with most errors below 5%.
- ChatGPT-4o demonstrated superior accuracy and slightly better reproducibility compared to ChatGPT-4.0.
- Observed inconsistencies highlight the need for careful review of AI-generated estimates.
Conclusions:
- ChatGPT-4.0 and ChatGPT-4o can be useful tools for preliminary sample size estimation in standard research scenarios.
- Users must exercise caution and seek expert validation for AI-generated statistical results.
- Further investigation is needed for complex statistical tasks and a wider array of AI models.
More Related Videos
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Related Concept Videos
Sample Size Calculation
The sample size for the given experiment or sampling effort is fundamental to any study design. Sample size decides the number of...
Testing a Claim about Population Proportion
There are two methods of testing a claim about a population proportion: (1) Using the sample proportion from the data where a binomial distribution is approximated to the normal distribution and (2) Using the binomial probabilities calculated from the data.
The first method uses normal distribution as an approximation to the binomial distribution. The requirements are as follows: sample size is large...
Quantifying and Rejecting Outliers: The Grubbs Test
Sampling Plans
Random sampling is a method where each member of the population has an equal chance of being selected for the sample. It involves selecting individuals randomly, often using random number generators or lottery-type methods. For example, when analyzing the properties of a...
Sample Proportion and Population Proportion
Expected Frequencies in Goodness-of-Fit Tests