Argo-Bench tests data agents on a 235-table, 7.5-billion-row warehouse

Argo-Bench tests data agents on a 235-table, 7.5-billion-row warehouse

Argo-Bench is an evaluation framework for data science and analytics agents, introduced in a paper listed on Hugging Face. It contains 210 tasks. The authors start from a complaint about existing text-to-SQL benchmarks: those benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Real enterprise warehouses are too sensitive to release, so the existing benchmarks are built on public datasets where a business event fits in a single table. Real enterprise work is different: it requires reasoning across dozens of tables, performing statistical analyses, and acting on the results.

To get around the data-privacy problem, the authors build a simulation. Drawing on public data, peer-reviewed industry literature and regulatory filings, they simulate a food delivery platform in New York City at true scale, with 81 million orders in 2024, grounded economics, fraud patterns and marketplace incentives. They then export this simulated world to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema.

The simulator's ground-truth state is withheld from the warehouse the agent sees. A task therefore requires the agent to reconstruct facts by navigating the warehouse before acting on them. Argo-Bench also goes beyond text-to-SQL: the agent files actions such as banning fraudulent accounts, allocating courier incentive budgets, or issuing back pay, and the grader scores each action by its consequences in the simulator. Every task has an executable reference solution that demonstrates solvability using only the warehouse.

The headline result: the strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points. The authors say they hope Argo-Bench drives progress toward agents that understand, navigate and act within real data environments.

Key facts

  • Argo-Bench has 210 data science and analytics tasks set in a simulated New York City food delivery platform with 81 million orders in 2024.
  • The simulated world is exported to an ERP warehouse of 235 tables and 7.5 billion rows, modeled on the Oracle E-Business Suite schema.
  • Agents must file actions, such as banning fraudulent accounts, allocating courier incentive budgets or issuing back pay, which are graded by their consequences in the simulator.
  • The simulator's ground-truth state is hidden from the agent, so it has to reconstruct facts by navigating the warehouse.
  • The strongest of 14 frontier and open-weight models scores 95 or higher on only 34.8% of tasks and averages 59.5 points.

Why it matters

The authors argue that established text-to-SQL benchmarks measure only query generation, that audits have found their answer keys frequently wrong, and that they sit on public datasets where one business event fits in a single table. Argo-Bench is aimed at the gap: a warehouse of 235 tables and 7.5 billion rows where facts must be pieced together across many tables, and where the agent's output is an action judged by what it does in a simulated business, not a query judged against an answer key. Because the warehouse is simulated from public sources, it offers enterprise-scale complexity without releasing a real company's data.

Who it affects

Teams building or buying data and analytics agents are the obvious audience, since the tasks mirror enterprise work: statistical analysis across many tables followed by operational actions. Researchers who evaluate agents on text-to-SQL benchmarks are also addressed, because the paper questions the quality of those benchmarks' answer keys.

How to use it

The abstract gives no release date, code or dataset link, or license, so there is nothing yet to tell a reader how to run the benchmark. What the abstract does describe is the setup an agent faces: it sees only the ERP warehouse, must navigate it to reconstruct facts, then files actions that the grader scores by their consequences in the simulator. Each task ships with an executable reference solution that shows it can be solved using only the warehouse.

How solid is it

This is a self-reported result from the benchmark's own authors, and what is available here is the paper's abstract. The numbers are specific: 210 tasks, 235 tables, 7.5 billion rows, 14 models evaluated. The executable reference solution per task is meant to show solvability. The abstract does not name the 14 models, including which one is strongest. It does not state the scoring scale, and the 59.5 average is not identified as a mean or a median, so it should not be read as a percentage. No scores are given for the other 13 models.

Risks and caveats

The data is simulated, not real: the food delivery platform is built from public data, industry literature and regulatory filings, so results may not carry over to a particular company's warehouse. Only the schema is modeled on Oracle E-Business Suite, per the authors. No details of how the grader computes scores are given beyond scoring by consequences in the simulator. The 34.8% figure is the share of tasks on which the best model scores 95 or higher, a different measure from its 59.5-point average, and the two should not be merged into one accuracy number.

“We hope Argo-Bench drives progress toward agents that understand, navigate, and act within real data environments.”

— Argo-Bench paper abstract