UMM-Reflection trains BAGEL to fix its own images with RL, +12.05 on GenEval

A new paper introduces UMM-Reflection, a way to teach a unified multimodal model to repair its own image generations. The starting point is that such models can both look at and render images, so in principle they can run a loop: diagnose what an image gets wrong, revise it, observe the result, and diagnose again.
The authors argue the two halves of that loop cannot be trained apart. Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop. They also say the usual shortcuts fall short. Supervised fine-tuning (SFT) on reflection trajectories gives a cold start but does not find the high-success repair paths. Naive reinforcement learning that optimizes only the renderer, or only one head, leaves most of the gain untapped.
UMM-Reflection instead applies RL to complete reflection trajectories inside one unified model. Sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies rather than different starting pictures. A single trajectory-level advantage then updates both the reflection tokens and the flow-based revisions. According to the authors, this avoids the combinatorial blow-up of assigning credit round by round. Unlike single-round editing or pipelines that rely on an external critic, credit flows across rounds and to both roles of the same model, and no verifier is needed at inference.
The reported results are on BAGEL. UMM-Reflection improves GenEval by 12.05 points over SFT. The gains transfer to three benchmarks that are not used in training: WISE (+10.97), OneIG-Bench (+3.48) and T2I-CompBench++ (+4.63).
Key facts
- UMM-Reflection applies reinforcement learning to complete reflection trajectories (diagnose, revise, observe, diagnose again) inside one unified multimodal model.
- Sibling trajectories share one initial image, so the group-relative advantage compares reflection strategies; one trajectory-level advantage updates both the reflection tokens and the flow-based revisions.
- On BAGEL, GenEval improves by 12.05 points over SFT.
- Gains transfer to WISE (+10.97), OneIG-Bench (+3.48) and T2I-CompBench++ (+4.63), none of which is used in training.
- No verifier is needed at inference, according to the authors.
Why it matters
Unified multimodal models can see and draw, which makes self-repair of generated images an obvious idea. The paper's contribution is a training recipe for it. The authors say SFT on reflection trajectories only gives a cold start, and RL that touches just the renderer or just one head leaves most of the gain untapped. Training the reflection text and the image generation together, over the whole loop, is their answer.
Who it affects
Researchers building unified multimodal models and image generators are the direct audience, since the method is framed around a single model that both critiques and renders. Teams that currently rely on an external critic pipeline or single-round editing are the ones the authors contrast their approach with.
How to use it
The source describes the method but does not mention a code or model release. In outline: sample sibling trajectories from one shared initial image, score them, compute a group-relative trajectory-level advantage, and use it to update both the reflection tokens and the flow-based revisions of the same model. At inference the authors say no verifier is needed.
How solid is it
The evidence is the reported benchmark deltas on BAGEL: +12.05 points on GenEval over SFT, and gains on WISE (+10.97), OneIG-Bench (+3.48) and T2I-CompBench++ (+4.63), which are not used in training. Those held-out gains are the stronger part of the case. The source gives only deltas, not absolute scores, and does not state the unit or reference point for the three transfer gains.
Risks and caveats
Results are reported only on BAGEL. The source does not give the number of reflection rounds, model size, training compute or dataset, and it does not quantify a comparison with external-critic pipelines or other published methods. The claims about credit assignment and inference without a verifier are the authors' own.
“Whether a revision helps is known only after it is rendered, so the reflection text and the image generation must be learned jointly, over the whole loop.”
— UMM-Reflection paper