InterEvolve lets an LLM evolve reward programs for humanoid tasks

The paper studies what it calls test-time evolution for humanoid loco-manipulation: getting a controller to solve tasks it was never trained for by repurposing its existing skills, improving from its own attempts, and keeping what it learns, all without retraining. The authors' key insight is that a broad controller already holds much of the competence a new task needs. That competence becomes accessible, they say, through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience.
InterEvolve realizes this interface with two components. The first is an object-aware forward-backward (FB) behavioral foundation model. Its object residuals sit on top of a frozen body prior, and they turn a new reward about the body or objects into loco-manipulation behavior at test time. The second component is the task specification itself: tasks are written as reward programs, meaning staged rewards with completion conditions and tunable constants.
The evolution loop splits the work in two. A large language model agent revises the structure of the program in context, drawing on execution feedback and on a skill library of verified programs. A numerical optimizer tunes the program's constants. Every candidate is verified across parallel simulation scenarios. Through this loop the program explores new ways to induce, repurpose and compose the controller's existing motor competence for the task at hand, and it improves over iterations.
On results, the authors say experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. The method also produces behaviors for diverse tasks, complex scenes and long-horizon compositions in simulation. Evolved skills run autonomously on a physical Unitree G1 humanoid from egocentric onboard perception.
Key facts
- InterEvolve adapts a humanoid controller to tasks it was never trained for at test time, without retraining.
- It has two components: an object-aware forward-backward (FB) behavioral foundation model, and tasks written as reward programs (staged rewards with completion conditions and tunable constants).
- An LLM agent revises the program structure using execution feedback and a skill library of verified programs; a numerical optimizer tunes the constants; every candidate is verified across parallel simulation scenarios.
- The authors report that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, while evolved programs release it, sometimes through novel strategies.
- Evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception; long-horizon and complex-scene results are stated for simulation.
Why it matters
The paper's pitch is that a broad humanoid controller already contains much of the competence a new task needs, and that the bottleneck is how the task is specified. Instead of retraining, InterEvolve searches over reward programs, using an LLM to rewrite their structure and an optimizer to tune their numbers. The authors report that this unlocks behavior that human-designed rewards leave untapped, sometimes through novel strategies.
Who it affects
The work is aimed at researchers building humanoid loco-manipulation systems, particularly those working with behavioral foundation models and reward design. The physical demonstration used a Unitree G1.
How to use it
This is a research paper, not a product. The method pairs an object-aware FB behavioral foundation model with reward programs that an LLM agent and a numerical optimizer refine against parallel simulation scenarios. No code or dataset release is mentioned.
How solid is it
The claims come from the authors' own summary of the work. They say evolved skills run autonomously on a physical Unitree G1, and that diverse tasks, complex scenes and long-horizon compositions work in simulation. No quantitative results are given: no success rates, number of iterations or number of tasks, and the size of the gap between human-designed and evolved rewards is not quantified.
Risks and caveats
Real-robot evaluation is described only for evolved skills on a G1, and which real-world tasks it performed is not stated. Long-horizon and complex-scene results are stated for simulation only. The LLM used is not named. Every candidate is verified in simulation before use, so the approach depends on that simulated verification.
“a broad controller already holds much of the competence a new task needs”
— InterEvolve paper, on its key insight