Decoded image

64-token latent sequence

Global time

Global progress

We make high-compression latent spaces competitive by enabling channel scaling.

Problem

Higher compression needs more channels.

Higher latent compression makes diffusion cheaper, but preserving information requires more channels.

Higher compression contracts the latent grid and expands its channel dimension.
Wide latents are hard to generate under matched training and sampling compute.

Structured Latent Space

We break the latent into levels.

Early channels learn structure. Later channels restore detail.

Keep Level 0; zero Levels 1–2.

Keep through

Training

One global time. Different local noise.

A sampled global time maps to one noise amount for each latent level.

One shared Transformer

Sample one global time.

Result

More channels no longer mean collapse.

By generating later levels as refinements, APEX degrades far more slowly as channel capacity grows.

Existing approaches degrade sharply as channels grow; APEX degrades much more gracefully.

Compute

Then high compression pays off.

At limited training and inference budgets, APEX makes fewer tokens outperform less-compressed models.

Guided gFID versus relative inference FLOPs: APEX at 64 tokens outperforms an f32c32 256-token model at every budget.
With APEX, the highly compressed model leads across the evaluated compute range; the corresponding baseline does not.

BibTeX

@article{dinkevich2026apex,
  title={APEX: Asynchronous Prefix Denoising for Extreme Compression},
  author={Dinkevich, David and Chiprut, Nisan and HaCohen, Yoav and Lischinski, Dani},
  year={2026}
}

Examples

Prefixes and generations

Staged generation: successive prefixes decoded when each level reaches t=0.
Decode each prefix as its level completes. Early prefixes recover structure; later ones add detail.
Generated prefixes decoded as each level completes.
Abstract

Fewer latent tokens promise cheaper diffusion. Yet aggressive latent compression creates a reconstruction-generation dilemma: adding channels improves reconstruction but makes generation harder. We introduce APEX, short for Asynchronous Prefix denoising for EXtreme compression, which pairs a hierarchical latent space with asynchronous denoising over channel groups. Our autoencoder orders channel prefixes from structure to detail, while one shared denoiser generates them at staggered local times without adding tokens. On ImageNet-512, APEX outperforms existing methods at every evaluated channel count under 64× and 128× compression. At matched training and guided sampling compute, it also outperforms a 256-token model at every evaluated budget, with larger gains at lower compute.