Iris-3B pixel-space model rivals latent models, but fine-tunes show no gain

Latent diffusion models compress images through a VAE, which loses some detail. Pixel-space diffusion models skip that step, which suggests they should do better on downstream tasks where fine-grained detail matters. A new preprint tests that claim along both routes to a pixel-space backbone.
The first route is training from scratch. The authors pretrain Iris-3B, a 3B-parameter pixel-space text-to-image transformer, through a 256 to 512 to 1024 resolution curriculum. Before scaling, they ablated the prediction target and representation alignment at 256^2 to decide what to scale. Iris-3B uses the pixel-transformer (PiT) head of PixelDiT. The second route is conversion: the authors take a pretrained latent model, FLUX.2 Klein base 4B, and convert it to pixel space.
Both families were then fine-tuned for two tasks: monocular depth estimation, and image restoration or super-resolution. The headline finding is negative: the authors find no significant improvement from using a pixel-space generative prior.
For depth, with one matched direct-regression recipe, Iris-3B is level with the latent FLUX.2 Klein, while the converted pixel FLUX.2 Klein falls behind it. On 4x DIV2K restoration, neither pixel model beats a latent FLUX.2 Klein fine-tune, and the converted one trails it slightly. The authors say they document the recipes, the failure modes and the remaining confounds behind this negative result.
The paper still counts Iris-3B as a positive result for pixel-space pretraining itself. It shows that pretraining with the PiT head of PixelDiT scales to 3B parameters and reaches text-to-image quality competitive with latent models, matching Qwen-Image on OneIG under the official evaluators at 1024^2. The authors release the weights and training code, hoping to help pave the way for further work on pixel-space generation.
Key facts
- Iris-3B is a 3B-parameter pixel-space text-to-image transformer pretrained from scratch through a 256 to 512 to 1024 curriculum, using the PiT head of PixelDiT.
- The authors also convert the latent FLUX.2 Klein base 4B to pixel space and fine-tune both families for monocular depth estimation and image restoration/super-resolution.
- They find no significant improvement from a pixel-space generative prior: for depth, Iris-3B is level with latent FLUX.2 Klein and the converted pixel version falls behind; on 4x DIV2K restoration neither pixel model beats a latent fine-tune.
- On text-to-image, Iris-3B is competitive with latent models and matches Qwen-Image on OneIG under the official evaluators at 1024^2.
- Iris-3B weights and training code are released.
Why it matters
Pixel-space diffusion has an appealing argument: no lossy VAE, so fine detail should survive, and that should help tasks like depth and restoration. This paper tests the argument directly and does not confirm it. At the same time, Iris-3B shows that training a pixel-space text-to-image model from scratch can scale to 3B parameters and reach quality competitive with latent models, matching Qwen-Image on OneIG at 1024^2. Both outcomes are useful for anyone deciding whether to invest in pixel-space backbones.
Who it affects
Researchers building diffusion backbones, and teams choosing between latent and pixel-space priors for dense prediction tasks such as monocular depth estimation and for image restoration or super-resolution. Groups interested in pixel-space generation also get a released 3B model to build on.
How to use it
The authors release Iris-3B weights and training code. The abstract does not state a release location or license for the weights and code. The paper also documents its recipes and failure modes, including the matched direct-regression recipe used for depth fine-tuning, which is a starting point for anyone repeating the comparison.
How solid is it
This is a preprint abstract-level account, and the comparisons are stated only qualitatively: no numeric scores such as depth metrics, PSNR or OneIG values are given. The claim about Qwen-Image is specific: a match on OneIG under the official evaluators at 1024^2. The authors themselves describe the depth and restoration result as negative and mention remaining confounds, so it is a report of no significant improvement rather than proof that pixel space cannot help.
Risks and caveats
Only depth estimation and image restoration/super-resolution were tested as downstream tasks, so the result does not cover other uses. The abstract does not say that pixel-space models are worse in general; it reports no significant improvement and notes remaining confounds. The converted pixel FLUX.2 Klein did worse than the latent original on both tasks (behind on depth, slightly behind on 4x DIV2K), so conversion looks like the weaker route here. No training compute, dataset or timescale is given.
“We find no significant improvement from using a pixel-space generative prior.”
— Iris-3B paper abstract