Landmark 8 stops to get here · leads to 5

Self-Attention

A mechanism where each token attends to all other tokens in the sequence to understand contextual relationships.

Your route here

8 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  4. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  5. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  6. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  7. Embedding ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  8. Attention Mechanism ✓ understood

    A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.

  9. Self-Attention · you are here ✓ understood

Picture it

EACH ROW ATTENDS TO THE COLUMNS · ROWS SUM TO 1thethe0.18catcat0.39satsat0.42onon0.26thethe0.18matmat0.38computed from toyword vectors:softmax(q·k / √d)
Each row shows how much one word attends to every word in the same sentence, including itself; the weights in each row sum to 1.

Where it sits

Explore nearby

In the research

All papers →

7 papers that build on Self-Attention ; showing 5, canon first.

Canon · 2017 Attention Is All You Need Every major LLM, and most modern vision and speech models, is a transformer. Canon · 2022 FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness It is a big part of why long context windows became affordable, and it ships inside nearly every LLM stack. Frontier · Dec 2025 · 671 citations DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models It shows an open model competing with the best closed ones on reasoning while getting cheaper to run. Frontier · Sep 2025 · 189 citations LongLive: Real-time Interactive Long Video Generation Minute-long, steerable video at 20 frames per second on one GPU makes interactive video generation real. Frontier · Oct 2025 · 116 citations Kimi Linear: An Expressive, Efficient Attention Architecture It is a credible claim that linear attention can replace full attention without a quality penalty.