Asynchronous Prefix denoising for EXtreme compression
1Lightricks 2Hebrew University of Jerusalem
Decoded image
64-token latent sequence
Global time
Global progress
We make high-compression latent spaces competitive by enabling channel scaling.
Structured Latent Space
Early channels learn structure. Later channels restore detail.
Keep Level 0; zero Levels 1–2.
Training
A sampled global time maps to one noise amount for each latent level.
One shared Transformer
Sample one global time.
Result
By generating later levels as refinements, APEX degrades far more slowly as channel capacity grows.
Compute
At limited training and inference budgets, APEX makes fewer tokens outperform less-compressed models.
@article{dinkevich2026apex,
title={APEX: Asynchronous Prefix Denoising for Extreme Compression},
author={Dinkevich, David and Chiprut, Nisan and HaCohen, Yoav and Lischinski, Dani},
year={2026}
}
Examples
Fewer latent tokens promise cheaper diffusion. Yet aggressive latent compression creates a reconstruction-generation dilemma: adding channels improves reconstruction but makes generation harder. We introduce APEX, short for Asynchronous Prefix denoising for EXtreme compression, which pairs a hierarchical latent space with asynchronous denoising over channel groups. Our autoencoder orders channel prefixes from structure to detail, while one shared denoiser generates them at staggered local times without adding tokens. On ImageNet-512, APEX outperforms existing methods at every evaluated channel count under 64× and 128× compression. At matched training and guided sampling compute, it also outperforms a 256-token model at every evaluated budget, with larger gains at lower compute.