Learn2Play Bench tests whether LLM agents learn from experience in new text games

Learn2Play Bench is a benchmark for a specific question: can LLM agents learn from experience when they land in an unfamiliar environment? The authors argue that existing benchmarks mostly evaluate tasks whose rules are given in the instructions or are already familiar to pretrained models. That makes it hard to separate learning from interaction from simple reasoning with knowledge the model already has.
To close that gap, the benchmark consists of newly designed text-based games whose rules are novel or counterintuitive. Agents cannot lean on pretrained knowledge alone and have to acquire what they need by playing. The games give reproducible feedback and automatic scoring, which allows controlled evaluation of learning across repeated attempts. The authors also vary the game instances to test whether agents can apply what they learned to new situations.
Using the benchmark, the authors evaluate how backbone models, self-evolving methods and agent harnesses affect an agent's learning ability, and report three findings.
First, on experience retention: keeping complete records of actions and feedback can support more effective learning than summarizing those experiences into rules or strategies.
Second, a human-agent gap: top-performing human players achieve higher peak scores than the evaluated agents. Humans explore more varied strategies and repeat actions less.
Third, the harness matters: with the backbone fixed, changing the harness can improve performance while reducing estimated inference cost.
The authors say these findings give insight into how LLM agents learn from experience and point to directions for future work on improving that ability. The paper links a project website for the benchmark.
Key facts
- Learn2Play Bench is built from newly designed text-based games with novel or counterintuitive rules, so agents must learn through interaction rather than rely on pretrained knowledge.
- The games offer reproducible feedback and automatic scoring, and game instances are varied to test whether agents can apply what they learned to new situations.
- Finding 1: retaining complete records of actions and feedback can support more effective learning than summarizing them into rules or strategies.
- Finding 2: top-performing human players achieve higher peak scores than the evaluated agents; humans explore more varied strategies and repeat actions less.
- Finding 3: with the backbone fixed, changing the harness can improve performance while reducing estimated inference cost.
Why it matters
Agents that work in the real world will meet environments they were not trained on, so learning from experience is central to their usefulness. The authors' point is that current benchmarks blur the measurement: when rules are stated up front or already familiar to the model, a good score may reflect existing knowledge rather than learning. Games with novel or counterintuitive rules are meant to isolate the learning itself.
Who it affects
Researchers building and evaluating LLM agents, especially those working on self-evolving methods and agent harnesses, are the direct audience. The three findings touch design choices they make: how an agent stores its experience, which harness wraps a given backbone, and how far agents still trail strong human players.
How to use it
The abstract points to a project website at https://liushiliushi.github.io/learn2play-bench-website/ for the benchmark. Teams could use the setup to compare backbone models, self-evolving methods and harnesses on how well they learn across repeated attempts and on varied game instances. The abstract does not describe a release or installation process beyond that link.
How solid is it
This is a paper abstract, so the claims are the authors' own summary of their results. It gives no model names, scores, cost figures or numbers of games, and no numeric size for the human-agent gap, so the strength of each finding cannot be judged from it. Finding 1 is worded with 'can', not as a universal result, and the inference cost in Finding 3 is described as 'estimated' with no stated method.
Risks and caveats
Read the findings as directional. The abstract does not say which self-evolving methods or harnesses were tested, so it is unclear how far the harness result generalizes. The human-agent comparison concerns top-performing players and peak scores only. The finding that full logs beat summaries is stated as something that can happen, not a rule that holds everywhere.
“Retaining complete records of actions and feedback can support more effective learning than summarizing these experiences into rules or strategies.”
— Learn2Play Bench paper abstract