RedVTP: Training-Free Acceleration of Diffusion Vision-Language Models Inference via Masked Token-Guided Visual Token Pruning
CVPR 2026 Findings
RedVTP estimates visual-token importance from the attention of still-masked response tokens. It prunes less important visual tokens after the first diffusion inference step, reducing latency by up to 64.97% and increasing token-generation throughput by up to 186% on the evaluated models.
* Equal contribution.