The Lattice of Transition Laws: diffusion and AR decoding as paths on one lattice

Diffusion and autoregression (AR) have long been treated as different categories of generative models: diffusion for continuous fields, AR for discrete tokens. Recent hybrids try to combine the strengths of both, and each one fixes its decoding schedule by design. The paper "The Lattice of Transition Laws" asks a different question: can the performance of decoding schedules of one model be predicted before decoding, at a fixed number of steps?
To answer it, the authors describe diffusion, AR and the models in between as paths on one corruption lattice. They then define the cost of a schedule as the dependence its parallel steps discard. From that cost they argue that the fewest steps of a zero-cost schedule are set by the geometry of the data, in the same way for tokens and for continuous fields.
The sharpest statement concerns data that are Markov on a graph and dependent along its paths. For such data, the fewest steps equal the graph's treedepth, which is logarithmic in the length of a sequence and linear in the side length of a grid. With fewer steps than the treedepth, every schedule pays a positive cost. The authors say they predict the ranking of those costs before decoding, using a kernel of pairwise dependence estimated from pretrained weights.
Across text, image and video generation, they report verifying most of the predictions about how different schedules rank under different metrics and benchmarks. They present the work as a design principle for decoding in future AR models, diffusion models and anything in between. Code is published on GitHub at https://github.com/TSUITUENYUE/The-Lattice-of-Transition-Laws.
Key facts
- The paper describes diffusion, AR and models in between as paths on one corruption lattice.
- The cost of a decoding schedule is defined as the dependence its parallel steps discard.
- For data that are Markov on a graph and dependent along its paths, the fewest steps of a zero-cost schedule equal the graph's treedepth: logarithmic in sequence length, linear in a grid's side length.
- Below that step count every schedule pays a positive cost, and the ranking is predicted before decoding with a kernel of pairwise dependence estimated from pretrained weights.
- The authors report verifying most predictions across text, image and video generation; code is on GitHub.
Why it matters
Diffusion and AR have usually been seen as separate families, and recent hybrids each choose a decoding schedule by design. This paper offers one frame for both, in which a schedule is a path on a lattice and its cost is the dependence that parallel steps throw away. The claim that the minimum step count is set by the geometry of the data, the same way for tokens and for continuous fields, is what makes the framing more than a relabelling. If it holds, schedule choice could be reasoned about before decoding rather than found by trial.
Who it affects
Mainly researchers who design or compare decoding schedules for diffusion models, AR models and hybrids. The authors position the work as a design principle for future AR models, diffusion models and anything in between, across text, image and video generation.
How to use it
The authors have released code at https://github.com/TSUITUENYUE/The-Lattice-of-Transition-Laws. The method as described estimates a kernel of pairwise dependence from pretrained weights and uses it to rank schedules before decoding at a fixed number of steps. No claim is made that the approach yields faster or better generation in practice; the claim is about predicting schedule rankings and offering a design principle.
How solid is it
This account rests on the paper's abstract. The authors say they verify most of the predictions about schedule rankings across text, image and video generation, under different metrics and benchmarks. The share of predictions verified is not quantified beyond "most", and which ones failed is not said. No numerical results are given, and no specific models, datasets or benchmarks are named. No publication venue or peer-review status appears in the source text.
Risks and caveats
The sharpest result is stated for data that are Markov on a graph and dependent along its paths, so how far it carries to real data is for the experiments to show, and the authors report only that most predictions held. No comparison with the performance of existing hybrid diffusion and AR models is reported. Treat it as a theoretical framing with reported empirical checks, not as a demonstrated improvement in generation quality or speed.
“This work therefore provides a design principle for decoding for future AR models, diffusion models, and anything in between.”
— Abstract of "The Lattice of Transition Laws"