Related Experiment Video
Updated: Aug 5, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Statistical Consistency in Artificial Intelligence-Assisted Computations: A Comparison of SPSS 31.0 and GPT-5.5
Frederick F Strale1, Rachael M German1, VeraLucia Mendes-Kramer2
1Eugene Applebaum College of Pharmacy and Health Sciences/Applied Health Sciences, Wayne State University, Detroit, USA.
Large language models (LLMs) like GPT-5.5 show statistical consistency with SPSS for common analyses, but advanced procedures require caution and verification.
Area of Science:
- * Computational statistics and artificial intelligence applications in scientific research.
- * Evaluation of large language models (LLMs) for statistical analysis.
- * Reproducibility and reliability in data analysis.
Background:
- * Artificial intelligence (AI), including large language models (LLMs), is increasingly used for statistical interpretation and research support.
- * Traditional software like IBM SPSS is the standard for transparent and reproducible analyses.
- * Concerns exist regarding the accuracy and consistency of LLM-generated statistical outputs, necessitating replication studies.
Purpose of the Study:
- * To evaluate the statistical reliability and alignment of newer LLMs with established analytic standards.
- * To compare the statistical outputs of GPT-5.5 with IBM SPSS 31.0.
- * To assess the consistency of LLM-generated statistical results across various procedures.
Main Methods:
- * 14 statistical procedures were applied to real datasets from peer-reviewed articles (2012-2023).
- * Analyses included descriptive statistics, correlations, t-tests, ANOVA, and MANOVA.
- * GPT-5.5 and SPSS 31.0 were used, with prompts and variable names copied directly.
Main Results:
- * GPT-5.5 showed concordance with SPSS 31.0 for descriptive statistics, correlations, t-tests, regression, and one-way ANOVA.
- * Discrepancies were noted in nonparametric tests, factorial ANOVA, MANOVA, and repeated-measures ANOVA.
- * Substantive interpretations were largely consistent, but not perfectly identical, between the two platforms.
Conclusions:
- * GPT-5.5 demonstrated considerable agreement with SPSS 31.0 for many common statistical procedures.
- * More complex analyses like factorial ANOVA, MANOVA, and repeated-measures ANOVA showed notable discrepancies.
- * LLM outputs can be a useful adjunct for exploratory research but require caution and verification with validated software for confirmatory analyses.
Related Concept Videos
Statistical Package for the Social Sciences (SPSS)
SPSS streamlines the process from data preparation to analysis and reporting. It is characterized by its user-friendly interface, which conceals...
Statgraphics
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance, comparing...
Statistical Hypothesis Testing
Statistical significance measures the probability that an observed result occurred by chance. If this probability, known as...
Statistical Analysis: Overview
One of the most commonly used statistical quantifiers is the mean, which is the ratio between the sum of the numerical values of all results and the...
Statistical Methods to Analyze Parametric Data: Student t-Test and Goodness-of-Fit Test
The Student's t-test is a statistical test that examines if there is a statistically significant difference between the means of two groups. This test is instrumental when dealing with data...