MIMESIS builds a 9B user simulator for training interactive agents

MIMESIS builds a 9B user simulator for training interactive agents

Training and evaluating interactive language agents needs rich user interactions, but collecting human feedback is expensive and hard to scale. Simulated users are the scalable alternative, and they have to do two things at once: resemble real user behavior and give agents useful learning experiences. The paper argues that most agent-training frameworks fall short here because they use off-the-shelf assistant LLMs as the simulated user. Those models' helpfulness can make them overly cooperative, explicit and behaviorally homogeneous compared with real users.

The authors introduce MIMESIS, a purpose-built user simulator. It is trained on human conversations with explicit reasoning supervision, and with 13 realistic behavioral patterns derived from real user interactions. The 9B model reaches a SOUL-Index of 65.7, which the authors say surpasses the strongest frontier model. Against Claude-Opus-5, the strongest baseline on RealUserSim and SimulatorArena, MIMESIS improves behavioral fidelity by 13.4 points and reduces Turing distance by 3.6 points, respectively.

The second step is to use the simulator as a training environment. The authors freeze MIMESIS and train an agent by having it interact with the frozen simulator through multi-turn reinforcement learning. Across eight environments, training with MIMESIS gives better agent performance than training with GPT-5.5 as the simulated user, and this holds under all nine unseen user simulators used for evaluation. The authors present this as stronger generalization to new user simulators.

The paper also proposes Coached On-Policy Self-Distillation (CSD). It takes the simulator's private reasoning traces and its subsequent utterances as feedback on how well the agent is addressing the user's needs. A coach converts that information into concise coaching notes describing how the agent can better anticipate user needs and adapt its behavior over the course of an interaction. CSD turns this feedback into dense, token-level supervision on top of sparse task rewards, and the authors report further gains across all nine evaluation user models.

Key facts

  • MIMESIS is a 9B user simulator trained on human conversations with explicit reasoning supervision and 13 behavioral patterns derived from real user interactions.
  • It scores 65.7 on the SOUL-Index, which the authors say surpasses the strongest frontier model.
  • Against Claude-Opus-5, the strongest baseline, it improves behavioral fidelity by 13.4 points on RealUserSim and reduces Turing distance by 3.6 points on SimulatorArena.
  • Across eight environments, agents trained against MIMESIS beat agents trained against GPT-5.5 under all nine unseen user simulators.
  • Coached On-Policy Self-Distillation (CSD) adds dense token-level supervision from simulator reasoning traces and coaching notes, with further gains across all nine evaluation user models.

Why it matters

Agents that talk to people need practice with people, and human feedback is expensive and hard to scale. The usual shortcut is to let a helpful assistant LLM play the user, but the authors argue such models are too cooperative, too explicit and too uniform to stand in for real users. MIMESIS tests a different idea: train a simulator on actual human conversations and make it reason privately about what it wants. The reported result is that agents trained against it generalize better to user simulators they have not seen.

Who it affects

Mainly researchers and teams that train interactive language agents with reinforcement learning and need a simulated user as the environment. It also touches anyone who builds user simulators for evaluating agents, since the paper reports results on two simulator benchmarks, RealUserSim and SimulatorArena.

How to use it

The abstract describes the recipe rather than a product. First train a simulator on human conversations with reasoning supervision and behavioral patterns. Then freeze it and train the agent against it with multi-turn reinforcement learning. Optionally add CSD, where a coach turns the simulator's private reasoning and later utterances into coaching notes that give token-level supervision. No release of code, weights or data is mentioned.

How solid is it

This rests on the paper's abstract alone, so every result is the authors' own claim. The numbers are concrete: a SOUL-Index of 65.7, a 13.4-point gain in behavioral fidelity on RealUserSim, a 3.6-point drop in Turing distance on SimulatorArena, and wins over GPT-5.5 as the simulated user in eight environments under nine unseen simulators. The abstract gives no scale, units or definition for SOUL-Index, behavioral fidelity or Turing distance, and no baseline SOUL-Index values. It does not state absolute scores for Claude-Opus-5.

Risks and caveats

The size of the agent-performance gains over GPT-5.5 and of the further CSD gains is not given, so how large the improvements are cannot be judged. The eight environments and nine user simulators are not named, and the agent model that is trained is not specified. No authors or institutions are named, and no base model for the 9B simulator is named. No limitations or caveats are stated in the abstract.

“whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real users”

— MIMESIS paper abstract, on off-the-shelf assistant LLMs used as simulated users