Synthesis Through Simulation: an LLM agent generates enterprise data via policy-enforcing APIs
Tool-calling agents have become central to enterprise AI, but the paper argues that training and evaluating them at scale is severely constrained by business and legal restrictions on enterprise systems, data and database schemas. The obvious workaround, tabular data synthesis, is limited by structural validity and by whether a schema is available at all. Procedure-based approaches have the opposite weakness: they typically lack distributional fidelity unless someone authors them for each domain.
The paper introduces Synthesis Through Simulation (STS), a schema-free paradigm. Instead of modelling a table and hoping the output is valid, an LLM agent generates data by executing operations against policy-enforcing APIs inside simulated enterprise environments. Because the data is produced through the same environment that defines what counts as valid, the authors say STS guarantees structural validity by construction. That also separates two problems that are usually tangled together: enforcing validity, and modelling the distribution of the data. Each can be addressed on its own.
The remaining problems, distributional fidelity and scalability, are handled by the Generalist Populator (GP), STS's domain-agnostic agent. The authors report that GP reaches 0.88 average marginal fidelity and 100% constraint satisfaction across all ten environments, without access to database schemas.
Two comparisons frame that result. Statistical synthesizers are inapplicable to seven of the ten environments because they need seed data. And schema-privileged agents, which are given the schema, fail 82% of trajectories on the airline environment's tightly coupled workflows, which the authors attribute to brittle task composition.
The authors open-source the full framework, all ten environments and the generated datasets at github.com/SAP/synthesis-through-simulation.
Key facts
- Synthesis Through Simulation (STS) is a schema-free paradigm: an LLM agent generates enterprise data by executing operations against policy-enforcing APIs in simulated environments.
- The authors say structural validity is guaranteed by construction, since data is produced through the same environment that defines what is valid.
- The Generalist Populator (GP), STS's domain-agnostic agent, is reported to reach 0.88 average marginal fidelity and 100% constraint satisfaction across all ten environments, without database schemas.
- Statistical synthesizers are inapplicable to seven of the ten environments because of seed data requirements; schema-privileged agents fail 82% of trajectories on the airline environment.
- The framework, all ten environments and the generated datasets are open-sourced in a repository under the SAP GitHub organisation.
Why it matters
Enterprise AI increasingly relies on tool-calling agents, yet the paper says training and evaluating them at scale is held back by business and legal limits on enterprise systems, data and schemas. Real data is often off limits, and the usual synthetic substitutes each fail in a different way: tabular synthesis needs a schema and struggles with structural validity, while procedure-based generation lacks distributional fidelity without per-domain authoring. STS offers a different route: let an agent act in a simulated system whose APIs enforce policy, so validity comes from the environment and the agent only has to worry about realistic distributions.
Who it affects
Teams building and testing tool-calling agents for enterprise settings are the obvious audience, especially where real systems, data and database schemas cannot be shared or used for training. Researchers working on synthetic data generation, and on agent benchmarks built from simulated environments, are also in scope, since the authors release ten environments and the datasets generated from them.
How to use it
The authors open-source the full framework, all ten environments and the generated datasets at https://github.com/SAP/synthesis-through-simulation. A practical starting point is to run GP in one of the released environments and inspect the data it produces. The paper's summary names no licence terms, costs or runtime figures, so check the repository for those.
How solid is it
The claims come from the paper's own abstract and are the authors' own measurements: 0.88 average marginal fidelity and 100% constraint satisfaction across ten environments. The validity guarantee is framed as holding by construction, a design property rather than an empirical result. The code and datasets are public, which lets others check the numbers. The abstract does not define how marginal fidelity is measured or what the scale of 0.88 is, and gives no baseline fidelity scores for other methods.
Risks and caveats
The abstract does not say which LLM powers the Generalist Populator. It does not name the ten environments except the airline one, nor say which seven are out of reach for statistical synthesizers. The 82% failure figure is stated only for schema-privileged agents on the airline environment, so it should not be read as a general failure rate. The abstract does not state that SAP is the employer of the authors or that SAP uses the system in production. Cost, runtime and data volume figures are not given.
“STS guarantees structural validity by construction”
— Synthesis Through Simulation paper abstract