Trace2Env lets an LLM agent act as the environment, built from historical traces

Trace2Env lets an LLM agent act as the environment, built from historical traces

The paper starts from a practical problem. Realistic replicas of environments are increasingly valuable for training and evaluating LLM agents, but the original systems may be inaccessible or impractical to reproduce.

The authors explore what they call agentic language world modeling. Instead of rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation. In other words, one agent plays the role of the system that another agent is trying to use.

They instantiate the idea with Trace2Env, a learning-free framework for settings where the original system is unavailable but historical interaction traces remain accessible. Trace2Env reconstructs those traces into a reusable environment worldbook. The worldbook contains three things: environment schemas, grounded evidence, and induced behavioral knowledge.

At runtime, the world model agent actively consults the worldbook together with persistent episodic state. From these it infers each action's observation and the lasting effects that the action has on state.

Across nine environments, the authors report that Trace2Env improves both next-observation fidelity and long-horizon interaction consistency over conventional prompt-based language world models (LWMs). They add a second check: in multi-turn interaction, task agent actions generated against Trace2Env remain valid more often when replayed in the real environment. They take this to indicate that its simulated dynamics better preserve the consequences of earlier actions across successive turns.

The authors conclude that these results establish agentic language world modeling as an alternative direction for building realistic environment replicas without reconstructing the original executable system.

Key facts

  • Agentic language world modeling: a world model agent acts as the environment for a task agent, rather than a rebuilt executable environment.
  • Trace2Env is a learning-free framework that works when the original system is unavailable but historical interaction traces remain accessible.
  • It turns traces into a reusable environment worldbook holding environment schemas, grounded evidence, and induced behavioral knowledge.
  • Across nine environments, the authors report better next-observation fidelity and long-horizon interaction consistency than conventional prompt-based LWMs.
  • Task agent actions generated against Trace2Env stay valid more often when replayed in the real environment.

Why it matters

Agent builders need realistic environments to train and test against, and the real systems are not always available or practical to reproduce. This paper proposes a different route: skip rebuilding the executable system and let a world model agent stand in for it, using only records of past interactions. If the approach holds up, old traces become a way to recreate an environment that can no longer be run.

Who it affects

Researchers and teams who train or evaluate LLM agents and cannot get at the original system. The setting the authors target is one where the system is unavailable but historical interaction traces remain accessible.

How to use it

The abstract describes the method rather than a product. Trace2Env takes historical interaction traces and reconstructs them into a worldbook of environment schemas, grounded evidence and induced behavioral knowledge. At runtime a world model agent consults that worldbook and persistent episodic state to produce each observation and state change for the task agent. It is learning-free, so no training step is part of the framework. No code, dataset or release information is mentioned.

How solid is it

The claims come from the authors' own abstract on a Hugging Face Papers listing. They report results across nine environments, plus a replay check in the real environment that they read as evidence the simulation preserves the consequences of earlier actions. The nine environments are not named or described. No numeric results (fidelity or consistency scores, margins of improvement) are given. The conventional prompt-based LWM baselines are not named, and the underlying LLM(s) used for the world model agent or task agent are not stated.

Risks and caveats

The conclusion that this is an alternative direction for building environment replicas is the authors' own framing. The number or size of historical traces required is not stated, and no limitations or failure cases are described. No authors or institutions are named in the source, so the work has no visible outside assessment here.

“rather than rebuilding an executable environment, a world model agent serves as the environment for a task agent and supports faithful and stateful simulation”

— Paper abstract, Hugging Face Papers