SQAM fine-tunes flow policies with a scalar adjoint, skipping per-step Jacobians

A paper on the Hugging Face papers page proposes Q-learning with Scalar Adjoint Matching, or SQAM, a way to fine-tune flow policies with off-policy reinforcement learning.
The starting point is that flow policies capture rich and diverse action distributions, and there is growing interest in fine-tuning them with off-policy RL so they improve beyond the demonstrations they were trained on. That is not trivial against a learned value function, the authors say, because the policy builds its action over many flow steps.
Adjoint matching is an existing, principled answer: it updates the flow model itself by propagating value information from the final action back to each flow step. Its cost is a vector-Jacobian product through the policy at every step, and that cost grows with both the number of flow steps and the size of the policy.
The authors make an observation: the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal. Motivated by this, they derive a closed-form scalar adjoint that scales the value gradient at the final action by the flow time. This removes the per-step vector-Jacobian products. They also find that controlling the critic's value at policy-generated actions is particularly important under the scalar adjoint. SQAM combines the scalar adjoint with a value penalty at those actions.
On results, the authors report that SQAM's gains concentrate on the four hardest OGBench domains, where its success rate exceeds that of the strongest baseline in each domain by 18 to 35 percentage points. To test whether the method extends to large pretrained policies, they also fine-tune a vision-language-action policy on a real bimanual robot. SQAM improves over supervised fine-tuning on all three tasks.
Key facts
- SQAM (Q-learning with Scalar Adjoint Matching) fine-tunes flow policies with off-policy RL using a closed-form scalar adjoint that scales the value gradient at the final action by the flow time.
- The scalar adjoint eliminates the per-step vector-Jacobian products that standard adjoint matching needs, a cost that grows with flow steps and policy size.
- The method rests on the authors' observation that the batch-averaged velocity Jacobian of pretrained flow policies concentrates on its diagonal.
- SQAM pairs the scalar adjoint with a value penalty at policy-generated actions, which the authors find particularly important.
- On the four hardest OGBench domains, SQAM's success rate exceeds the strongest baseline in each domain by 18 to 35 percentage points; on a real bimanual robot, a vision-language-action policy fine-tuned with SQAM improves over supervised fine-tuning on all three tasks.
Why it matters
Fine-tuning flow policies with off-policy RL is a way to push them beyond their demonstrations, but the policy builds its action over many flow steps, which makes optimising against a learned value function hard. Adjoint matching handles this in a principled way but pays a vector-Jacobian product at every step, with a cost that grows with the number of flow steps and the policy size. SQAM's pitch is that, because the velocity Jacobian of pretrained flow policies is concentrated on its diagonal on average, that per-step cost can be replaced by a closed-form scalar adjoint.
Who it affects
The work is aimed at researchers and engineers fine-tuning flow policies with off-policy RL, including those working on large pretrained policies such as vision-language-action models for robots. The authors tested one such policy on a real bimanual robot.
How to use it
The abstract describes a method, not a product. In practice SQAM means swapping the per-step vector-Jacobian products of adjoint matching for a scalar adjoint that scales the value gradient at the final action by the flow time, and adding a value penalty at policy-generated actions. The abstract mentions no code release, and the kind of penalty (its form or coefficient) is not specified.
How solid is it
This is a preprint, and the account here rests on its abstract alone. The headline result is concrete: on the four hardest OGBench domains, SQAM's success rate exceeds the strongest baseline in each domain by 18 to 35 percentage points, and the real-robot test shows gains over supervised fine-tuning on all three tasks. The abstract names no baselines and no individual OGBench domains, and gives no absolute success rates, only the margin over the strongest baseline per domain. It does not state the size of the improvement over supervised fine-tuning on the robot tasks, and the tasks are not named.
Risks and caveats
The authors say SQAM's gains concentrate on the four hardest OGBench domains, so the large margins should not be read as a uniform improvement across all domains. The real-robot test covers three tasks and compares against supervised fine-tuning. No speedup or compute-saving figure for removing the vector-Jacobian products is given, and the abstract mentions no limitations.
“SQAM improves over supervised fine-tuning on all three tasks.”
— Abstract of the SQAM paper