The Accuracy and Replicability of Data Analysis Performed with Large Language Models
Abstract
This peer reviewed letter documents a series of experiments with OpenAI proprietary large language models assessing automated data analysis. We simulate data from known distributions and share only summary statistics and stem-leaf plots [this was started before graphics processing was widespread] and ask the model to identify the distribution that generated it. Initially, the models were of very little value in recovering the ground truth distributions. That said, through generations of models, the ability to discriminate drastically improves to the point that recent models likely perform at least as well as human experts.
Citation
BibTeX citation:
@misc{eaglesmith2026,
author = {Eaglesmith, Justus and Johnson, Tim and W. Walker, Robert},
publisher = {Taylor and Francis for the American Statistical
Association},
title = {The {Accuracy} and {Replicability} of {Data} {Analysis}
{Performed} with {Large} {Language} {Models}},
date = {2026-06-29},
url = {https://robertwwalker.work/quarto_publications/AmStatLetter/},
langid = {en}
}
For attribution, please cite this work as:
Eaglesmith, Justus, Tim Johnson, and Robert W. Walker. 2026. “The
Accuracy and Replicability of Data Analysis Performed with Large
Language Models.” American Statistician. Taylor and
Francis for the American Statistical Association. https://robertwwalker.work/quarto_publications/AmStatLetter/.