Abstract
Diffusion Transformers (DiT) have demonstrated exceptional capabilities in visual generation, but their inherently sequential denoising process leads to substantial inference latency. Previous acceleration strategies, such as caching intermediate feature maps, often struggle with sharp irregularities across timesteps, leading to error accumulation and degraded generation fidelity.
In this work, we propose BRACE, a novel mathematical framework leveraging Barycentric Rational Forecasting to accurately extrapolate and tame these sharp irregularities. Extensive experiments across various generative tasks demonstrate that BRACE consistently outperforms existing methods in the latency-quality trade-off, enabling fast inference without structural modifications.
Qualitative Results
Qualitative comparison on Text-to-Video Generation.