The Accuracy and Replicability of Data Analysis Performed with Large Language Models

Authors

Justus Eaglesmith

Tim Johnson

Robert W. Walker

Published

June 29, 2026

Abstract

This peer reviewed letter documents a series of experiments with OpenAI proprietary large language models assessing automated data analysis. We simulate data from known distributions and share only summary statistics and stem-leaf plots [this was started before graphics processing was widespread] and ask the model to identify the distribution that generated it. Initially, the models were of very little value in recovering the ground truth distributions. That said, through generations of models, the ability to discriminate drastically improves to the point that recent models likely perform at least as well as human experts.

Citation

BibTeX citation:
@misc{eaglesmith2026,
  author = {Eaglesmith, Justus and Johnson, Tim and W. Walker, Robert},
  publisher = {Taylor and Francis for the American Statistical
    Association},
  title = {The {Accuracy} and {Replicability} of {Data} {Analysis}
    {Performed} with {Large} {Language} {Models}},
  date = {2026-06-29},
  url = {https://robertwwalker.work/quarto_publications/AmStatLetter/},
  langid = {en}
}
For attribution, please cite this work as:
Eaglesmith, Justus, Tim Johnson, and Robert W. Walker. 2026. “The Accuracy and Replicability of Data Analysis Performed with Large Language Models.” American Statistician. Taylor and Francis for the American Statistical Association. https://robertwwalker.work/quarto_publications/AmStatLetter/.