SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
Junsong Chen et al.
arXiv:2509.24695
In short
SANA-Video is a small diffusion model for high-resolution, minute-long video built on linear attention, with a constant-memory cache for block-by-block generation. It trained for about 1% of MovieGen’s cost and is 16× faster than comparable small models.
Why it matters
Quality video generation that fits on a single consumer GPU.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 13 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Diffusion Model · read first ✓ understood
A generative model that learns to denoise data, achieving state-of-the-art image generation (Stable Diffusion, DALL-E 2).
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Feature ✓ understood
A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.
- Deep Learning ✓ understood
A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.
- Representation Learning ✓ understood
Learning useful features or representations of data automatically, rather than hand-crafting them.
- Embedding ✓ understood
A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.
- Attention Mechanism ✓ understood
A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.
- Self-Attention · read first ✓ understood
A mechanism where each token attends to all other tokens in the sequence to understand contextual relationships.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Inference ✓ understood
Running a trained model on new inputs to get predictions, with its weights frozen: the stage of a model's life that users actually interact with.
- Inference Latency · read first ✓ understood
The time delay between submitting input and receiving output from a deployed model, critical for real-time applications.
In the frontier
- Rank
- #67 of 100
- Citations
- 82
- as of Aug 9, 2026
- Published
- Sep 2025
Topics: Video generation and world models , Model architecture , Efficiency and serving
Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026