Structured residual connections cut Diffusion Transformer training iterations by up to 1.73 times

Structured residual connections cut Diffusion Transformer training iterations by up to 1.73 times

Diffusion Transformers (DiTs) have become a scalable backbone for high-fidelity image synthesis. The paper starts from a contrast with U-Net based diffusion models, which rely on rigid, hand-crafted skip connections. DiTs, the authors say, predominantly use a uniform residual stream that integrates all preceding layers as a single monolithic state.

The authors propose to turn residual connections from passive summation into an active retrieval mechanism optimized for image denoising. They first analyse DiT's internal representation systematically and find a latent preference for early-layer feature reuse and symmetric layer guidance. Motivated by that, they introduce a structured connectivity design that combines local residual connections with long-range pathways.

Instead of static skip connections or dense all-layer routing, the method lets each transformer block selectively "attend" to critical earlier representations. Through direct, differentiable cross-depth paths, a block dynamically retrieves spatial and semantic cues from earlier layers.

The reported results: the adaptive connectivity converges faster, needing up to 1.73 times fewer training iterations, and brings significant gains in FID and visual quality with less than 0.1% additional parameters. On a strong REPA-XL/2 model, FID improves from 5.9 to 4.34 without guidance (lower is better), and the method reaches 1.39 FID with classifier-free guidance.

The authors conclude that their findings suggest adaptive cross-layer connectivity is a critical yet underexplored factor in diffusion transformers, and that structured information pathways are a simple and effective direction for improving scalable generative models.

Key facts

  • The paper replaces the uniform residual stream in Diffusion Transformers with a structured design that combines local residual connections and long-range pathways.
  • Each transformer block can selectively "attend" to critical earlier representations through direct, differentiable cross-depth paths.
  • Reported convergence gain: up to 1.73 times fewer training iterations, with less than 0.1% additional parameters.
  • On REPA-XL/2, FID improves from 5.9 to 4.34 without guidance; with classifier-free guidance the method reaches 1.39 FID.
  • The authors' analysis of DiT internals found a latent preference for early-layer feature reuse and symmetric layer guidance.

Why it matters

Diffusion Transformers are a scalable backbone for high-fidelity image synthesis, yet the paper argues their plain residual stream lumps all preceding layers into one monolithic state. U-Net diffusion models use hand-crafted skip connections instead. This work proposes a middle path: connections that are structured and learned, not static and not dense. If the reported gains hold, faster convergence and better FID come at under 0.1% extra parameters. The authors suggest that cross-layer connectivity is a critical yet underexplored factor in diffusion transformers.

Who it affects

Mainly researchers and engineers who train or design diffusion transformers for image generation, especially those building on REPA-XL/2, since that is the model the paper reports improving. Teams paying for long training runs have the most to gain from fewer training iterations.

How to use it

The paper describes an architectural change: keep local residual connections and add long-range pathways so each block can retrieve spatial and semantic cues from earlier layers. The source states that the extra cost is under 0.1% of parameters. No code or model release is mentioned.

How solid is it

The claims come from the paper's own experiments, with concrete figures: up to 1.73 times fewer training iterations, REPA-XL/2 FID from 5.9 to 4.34 without guidance, and 1.39 FID with classifier-free guidance. The paper also grounds the design in an analysis of DiT's internal representations. The dataset, image resolution and evaluation protocol behind the FID numbers are not stated. The abstract does not say which baseline the 1.73 factor is measured against or at what FID target. No comparison to other methods beyond REPA-XL/2 is given. The authors word their conclusion as a suggestion, not a proof.

Risks and caveats

The 1.73 figure is a maximum ("up to"), not a typical gain. No compute cost, wall-clock time or inference-speed effect is given, so the practical saving is unclear from the source. No authors or institutions are named, and the results have not been checked here beyond the paper's own report.

“propose to transform them from passive summation into an active retrieval mechanism optimized for image denoising”

— Paper abstract, on residual connections in diffusion transformers