Qwen3-VL Technical Report
Shuai Bai et al.
arXiv:2511.21631
In short
Qwen3-VL is an open family of vision-language models, from 2B dense up to a 235B mixture-of-experts, that reads text, images and video in one context of up to 256K tokens. The report covers upgrades to how positions and timestamps are encoded and how features from several vision-encoder layers are fed to the language model.
Why it matters
It is the most-cited paper of the year and a reference point for open multimodal models.
Read first
The 4 Field Guide ideas this paper leans on.
Starting from scratch? The full route 24 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Deep Learning ✓ understood
A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.
- Computer Vision ✓ understood
The field of AI that gets computers to extract meaning from images and video: what is in them, where it is, and how it moves.
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Feature ✓ understood
A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.
- Representation Learning ✓ understood
Learning useful features or representations of data automatically, rather than hand-crafting them.
- Embedding ✓ understood
A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.
- Attention Mechanism ✓ understood
A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.
- Transformer ✓ understood
A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.
- Vision Transformer · read first ✓ understood
Applying the transformer architecture to computer vision by treating image patches as tokens, achieving state-of-the-art results.
- Natural Language Processing ✓ understood
The field of AI that lets computers read, interpret, translate and generate human language, from spam filters and search to chatbots.
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- Tokenization ✓ understood
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
- Language Modeling ✓ understood
Learning probability distributions over sequences of words to predict what comes next.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Unsupervised Learning ✓ understood
Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.
- Self-Supervised Learning ✓ understood
Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).
- Pre-training ✓ understood
Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.
- Large Language Model ✓ understood
A neural network, almost always a transformer, trained on vast amounts of text to predict the next token, which lets it write, answer, summarize and follow instructions.
- Multimodal Model · read first ✓ understood
Models processing multiple data types (text, images, audio) jointly, like GPT-4V, Gemini, or CLIP.
- Context Window · read first ✓ understood
The maximum number of tokens an LLM can process at once, including both input prompt and generated output. Also called context length.
- Feedforward Network ✓ understood
A neural network where information flows in one direction from input to output without cycles.
- Mixture of Experts · read first ✓ understood
An architecture where multiple specialized sub-networks (experts) process inputs, with a gating network routing to relevant experts.
In the frontier
- Rank
- #1 of 100
- Citations
- 1.8K
- as of Aug 9, 2026
- Published
- Nov 2025
Topics: Vision-language models , Video understanding , Agent benchmarks and computer use
Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026