Marianne Arriola
@mariannearr
Nice to see this architecture explored for video! We also found it effective for language: a large encoder processes context and a lightweight decoder quickly denoises token blocks x.com/mariannearr/st…
Xingjian Bai@SimulatedAnneal · Feb 19Do causal video diffusers really need dense causal attention at every layer, every denoising step?
We looked inside and found: no. Causality is separable from denoising.
Here are two surprising observations that hold across architectures, training objectives, and scales.
We looked inside and found: no. Causality is separable from denoising.
Here are two surprising observations that hold across architectures, training objectives, and scales.
Open quoted post →
1 25