Skill2Real framework lifts LIBERO-Pro Long success from 2.0% to 56.3%

A paper on Hugging Face Papers introduces Skill2Real, an agentic policy framework that learns executable robot-manipulation skills through a shared application programming interface (API) and carries them from simulation to a real robot. The opening premise is that transferring skills from simulation to reality requires task knowledge that stays usable across differences in perception, dynamics and embodiment.
The framework has two parts. A Proposer-Verifier-Governor (PVG) loop uses privileged simulation evidence to diagnose outcomes and validate updates, while keeping the learned skills grounded in public observations and API semantics. Skills are organised in two memories: the Cerebellum first acquires local manipulation skills, and the Brain then learns task-level composition with the Cerebellum frozen. According to the authors, both memories transfer to the real robot without task-policy fine-tuning or skill-memory updates.
The abstract gives four sets of results. First, GPT-5.6 Sol learns skills on LIBERO-90, and each frozen checkpoint is evaluated with GPT-6 Astra. Under that setup, LIBERO-Pro Long success rises from 2.0% to 56.3%, without any training on Pro Long. Second, independent Robosuite training reaches 85.1% mean success with Sol and 89.4% with Opus 5, across seven tasks. Third, frozen Sol-trained LIBERO-90 skills achieve 78.75% mean completion across four real-world manipulation tasks with Astra. Fourth, an ablation: removing the Verifier during LIBERO-90 training lowers final Pro Long success by 17.3 percentage points, and removing the Governor lowers it by 13.3 percentage points.
The authors conclude that these results support learning and transferring a hierarchy of executable skills through a common robot interface.
Key facts
- Skill2Real is an agentic policy framework that learns executable skills through a shared API; a Proposer-Verifier-Governor loop validates updates using privileged simulation evidence.
- Skills are split into a Cerebellum (local manipulation skills) and a Brain (task-level composition, learned with the Cerebellum frozen); both transfer to the real robot without task-policy fine-tuning or skill-memory updates.
- LIBERO-Pro Long success rises from 2.0% to 56.3% with GPT-5.6 Sol learning on LIBERO-90 and GPT-6 Astra evaluating each frozen checkpoint, with no training on Pro Long.
- Frozen Sol-trained LIBERO-90 skills reach 78.75% mean completion across four real-world manipulation tasks with Astra; Robosuite training reaches 85.1% (Sol) and 89.4% (Opus 5) mean success across seven tasks.
- Removing the Verifier cuts final Pro Long success by 17.3 percentage points; removing the Governor cuts it by 13.3 percentage points.
Why it matters
Moving a skill learned in simulation onto a physical robot is hard because perception, dynamics and embodiment all differ between the two. Skill2Real takes the approach of learning executable skills behind a shared API, so the knowledge is tied to public observations and API semantics rather than to simulator internals. The reported payoff is transfer to a real robot with no task-policy fine-tuning and no skill-memory updates. The jump on LIBERO-Pro Long, from 2.0% to 56.3% without training on that benchmark, is the headline figure.
Who it affects
Robotics researchers working on sim-to-real transfer and on skill libraries for manipulation are the direct audience. The abstract names the benchmarks and environments involved: LIBERO-90, LIBERO-Pro Long and Robosuite, plus four real-world manipulation tasks. It also names the models used as the skill learner and evaluator: GPT-5.6 Sol, GPT-6 Astra and Opus 5.
How to use it
This is a research result, not a product. The abstract describes the method: a Proposer-Verifier-Governor loop, a Cerebellum for local skills and a Brain for composition, all behind a common API. The source states that no code release, dataset release or peer-review status is mentioned, so anyone wanting to build on it would need the full paper.
How solid is it
The source is the paper abstract only, and every figure comes from the authors' own reporting. The numbers are specific, and the ablation gives some internal support: removing the Verifier or the Governor lowers final Pro Long success by 17.3 and 13.3 percentage points. The authors word their conclusion cautiously, saying the results support the idea of learning and transferring a hierarchy of executable skills. The abstract does not say how many trials were run in the real-world evaluation, and no baseline is given for the 78.75% figure.
Risks and caveats
The abstract does not say which real robot or which four real-world tasks were used, and it does not name the seven Robosuite tasks. No baseline is given for the real-world 78.75% mean completion. It is not stated whether the 2.0% starting point is a baseline method or the untrained state beyond the phrase 'raises ... from 2.0%'. The 56.3% result depends on a specific pairing, with Sol learning skills and Astra evaluating each frozen checkpoint, so it should not be read as a general figure for the framework. The Pro Long gain is also a success rate of 56.3%, which leaves a large share of tasks unsolved.
“Both memories transfer to the real robot without task-policy fine-tuning or skill-memory updates.”
— Skill2Real paper abstract