Related Experiment Video
Updated: Sep 14, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Can large language models accurately compute descriptive statistics from structured datasets? A comparative
1University Institute for primary care (IuMFE), University of Geneva, Geneva, Switzerland.
Background:
Large language models (LLMs) are increasingly used to support statistical analyses in biomedical research. However, their ability to accurately and reproducibly compute descriptive statistics directly from datasets has received limited evaluation.
Objective:
To compare the accuracy, within-modality repeatability, and between-modality consistency of ChatGPT and Claude in generating descriptive statistics from structured datasets provided through different input modalities.
Methods:
Two publicly available Stata datasets were evaluated: 'auto.dta' (74 observations, 12 variables) and 'citytemp.dta' (956 observations, 6 variables). Original variable names were replaced with generic labels. Each dataset was analyzed using three input modalities (copy-paste, Word, and Excel), with two independent repetitions per modality. Stata served as the reference standard. Accuracy was assessed for counts of missing and non-missing observations, minima, maxima, means, standard deviations, medians, quartiles, frequencies, and percentages.
Results:
ChatGPT and Claude produced identical results across all analyses. For each model, 624 categories of descriptive statistics were evaluated. Exact agreement with the Stata reference standard was observed for 576 of 624 categories (92.3%). The only deviations involved first and third quartiles (Q1-Q3) for eight variables. Post hoc analyses demonstrated that these differences were entirely attributable to the use of a different, but mathematically valid, quartile definition based on linear interpolation rather than computational errors. Within-modality repeatability and between-modality consistency were complete across the evaluated analyses.
Conclusions:
ChatGPT and Claude demonstrated excellent accuracy and consistent results across repeated analyses and input modalities for the datasets and descriptive statistics evaluated. After accounting for differences in quartile definitions, no calculation errors were identified across the 624 evaluated categories per model.
