ViDA: Video Diffusion Transformer Acceleration with Differential Approximation and Adaptive Dataflow
Li Ding, Jun Liu, Shan Huang, Guohao Dai · 2025
Recent advancements in Video Diffusion Transformer (VDiT) models have greatly promoted the development of video generation, as exemplified by Sora of OpenAI. However, there are still two challenges for VDiT: 1) There is still existing large inter-frame redundant computation. Previous works on reducing computation based on inter-frame similarity simply consider the Act-W operators. The remaining Act-Act operators still dominate the execution of VDiT (about 57%). 2) Operational intensity varies greatly, leading to under-utilization. There is a massive gap between the operational intensity of Act-W and Act-Act operators in VDiT with multiple frames. Previous works with the static hardware architecture and dataflow lead to under-utilization (<36.42%).