GRACE compresses video autoencoders, cutting Wan2.1-I2V-14B tokens 8x

The paper starts from a familiar idea: a highly compressed video autoencoder can speed up a video diffusion model, because the Diffusion Transformer (DiT) then works on far fewer tokens. The authors say the idea is hard to execute for three reasons. A higher compression ratio degrades reconstruction quality. Recovering that quality takes more channels, which is known to slow the convergence of the DiT. And the compressed latent differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost.
They also look at an apparent shortcut: compressing the very autoencoder the DiT was trained with seems to preserve compatibility. Yet optimizing that autoencoder for reconstruction alone still pushes the latent away from the distribution the DiT has learned.
To address this, the paper proposes GRACE, short for Generation-Aware Latent Compression for Efficient Video Generation. It is a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT.
In the first stage, the method keeps a frozen base latent from the pretrained encoder and learns a residual latent that carries the information lost under stronger compression. At the same time, it aligns the compressed latent with the pretrained latent in the feature space of the frozen DiT, so the autoencoder is optimized for generation rather than for reconstruction alone.
In the second stage, the DiT is adapted with lightweight fine-tuning and asymmetric denoising, in which the base latent is denoised ahead of the residual.
The headline result: GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.
Key facts
- GRACE (Generation-Aware Latent Compression for Efficient Video Generation) is a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT.
- Stage one keeps a frozen base latent from the pretrained encoder, learns a residual latent for information lost to stronger compression, and aligns the compressed latent with the pretrained one in the feature space of the frozen DiT.
- Stage two adapts the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual.
- On Wan2.1-I2V-14B, GRACE cuts the token count by 8x and latency by 11.1x at 480x832x81.
- The authors say VBench generation quality matches the pretrained pipeline before compression.
Why it matters
Fewer tokens is one of the most direct ways to make a video diffusion transformer cheaper to run, and an 8x token cut with an 11.1x latency cut is a large claimed gain. The paper's argument is that the obstacle has been compatibility: a more compressed latent usually forces the pretrained DiT to be retrained from scratch or adapted at considerable cost. GRACE is pitched as a way to compress while staying close to the latent the DiT already knows.
Who it affects
The result is reported for Wan2.1-I2V-14B, so people working with that image-to-video model, or building on pretrained video DiTs and their autoencoders, are the most directly concerned. Anyone weighing the cost of serving video generation will care about the latency claim.
How to use it
The source is a paper abstract describing a method, not a product. No code or weights release is mentioned. In practice the recipe has two steps: compress the pretrained autoencoder with a frozen base latent plus a learned residual, aligned to the pretrained latent in the frozen DiT's feature space; then fine-tune the DiT lightly with asymmetric denoising, base before residual.
How solid is it
The numbers are the authors' own claims. They report 8x fewer tokens, 11.1x lower latency at 480x832x81, and VBench quality matching the pretrained pipeline before compression. No VBench scores are given, only the statement that quality matches. No hardware, batch size or absolute latency figures are given for the 11.1x figure, and the source does not say whether it is a mean or a median. Results are reported only for Wan2.1-I2V-14B.
Risks and caveats
Matching on VBench is a benchmark statement, and the source gives no scores to inspect. The compression ratio of the autoencoder, its channel count and the fine-tuning cost are not stated, so the real price of adapting the DiT is unclear. Whether the approach carries over to other models is not addressed in the source, which mentions only Wan2.1-I2V-14B.
“GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.”
— GRACE paper abstract