SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation
arXiv cs.CV1mo4 min read
arXiv:2606.30849v1 Announce Type: new Abstract: Diffusion Transformers (DiTs) have significantly advanced audio-driven portrait animation, but their high computational cost leads to substantial inference latency. Although training-free diffusion caching accelerates inference significant, existing methods are primarily developed for text-conditioned generation and overlook the spatial and modality imbalances inherent in audio-driven portrait animation. In this paper, we propose SyncCache, a training-free caching acceleration method tailored for DiT-based portrait animation that explicitly explo
