Many AI analysts, one dataset: Navigating the agentic data science multiverse.

Empirical conclusions depend not only on data but also on analytic decisions. Many-analyst studies have quantified this dependence: independent teams testing the same hypothesis on the same dataset regularly reach conflicting conclusions. But such studies require costly human coordination. We show that fully autonomous AI analysts built on large language models (LLMs) can, cheaply and at scale, produce the analytic dispersion observed in human many-analyst studies. In our framework, each AI anal
Empirical conclusions depend not only on data but also on analytic decisions. Many-analyst studies have quantified this dependence: independent teams testing the same hypothesis on the same dataset regularly reach conflicting conclusions. But such studies require costly human coordination. We show that fully autonomous AI analysts built on large language models (LLMs) can, cheaply and at scale, produce the analytic dispersion observed in human many-analyst studies. In our framework, each AI analyst independently executes a complete analysis pipeline on a fixed dataset and hypothesis; a separate AI auditor screens every run for methodological validity. Across three datasets, AI analyst-produced analyses exhibit substantial dispersion in effect sizes, [Formula: see text]-values, and conclusions. This dispersion can be traced to identifiable analytic choices in preprocessing, model specification, and inference that vary systematically across LLM and persona conditions. Critically, the outcomes are steerable: reassigning the analyst persona or LLM shifts the distribution of results even among methodologically sound runs. These results highlight a central challenge for AI-automated empirical science: when defensible analyses are cheap to generate, evidence becomes abundant and vulnerable to selective reporting. The same capability also helps address it: treating analyst results as distributions makes analytic uncertainty visible, and deploying AI analysts against a published specification can reveal how much disagreement stems from underspecified design choices. Taken together, our results motivate a transparency norm: AI-generated analyses should be accompanied by multiverse-style reporting and full disclosure of the prompts used, on par with code and data.




