Vision & Multimodal Apr 2026 · #64 most cited · 87 citations

Qwen3.5-Omni Technical Report

Qwen Team

arXiv:2604.15804

In short

Qwen3.5-Omni scales the omni-modal Qwen line to hundreds of billions of parameters and a 256K context, trained on over 100 million hours of audio-visual data. It handles hours of audio and minutes of video, improves the stability of streaming speech, and can even write code from spoken and visual instructions.

Why it matters

The frontier of open omni-modal models, rivalling Gemini on audio.

Read first

The 4 Field Guide ideas this paper leans on.

Starting from scratch? The full route 26 ideas · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  4. Computer Vision ✓ understood

    The field of AI that gets computers to extract meaning from images and video: what is in them, where it is, and how it moves.

  5. Natural Language Processing ✓ understood

    The field of AI that lets computers read, interpret, translate and generate human language, from spam filters and search to chatbots.

  6. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  7. Tokenization ✓ understood

    Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.

  8. Language Modeling ✓ understood

    Learning probability distributions over sequences of words to predict what comes next.

  9. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  10. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  11. Unsupervised Learning ✓ understood

    Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.

  12. Self-Supervised Learning ✓ understood

    Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).

  13. Pre-training ✓ understood

    Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.

  14. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  15. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  16. Embedding ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  17. Attention Mechanism ✓ understood

    A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.

  18. Transformer ✓ understood

    A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.

  19. Large Language Model ✓ understood

    A neural network, almost always a transformer, trained on vast amounts of text to predict the next token, which lets it write, answer, summarize and follow instructions.

  20. Multimodal Model · read first ✓ understood

    Models processing multiple data types (text, images, audio) jointly, like GPT-4V, Gemini, or CLIP.

  21. Feedforward Network ✓ understood

    A neural network where information flows in one direction from input to output without cycles.

  22. Mixture of Experts · read first ✓ understood

    An architecture where multiple specialized sub-networks (experts) process inputs, with a gating network routing to relevant experts.

  23. Feature Engineering ✓ understood

    The process of selecting, transforming, and creating input features to improve model performance.

  24. Audio Processing ✓ understood

    Techniques for analyzing, transforming, and understanding audio signals for tasks like speech recognition and music generation.

  25. Speech Recognition · read first ✓ understood

    Converting spoken language into text using acoustic models and language models, now dominated by deep learning.

  26. Text-to-Speech · read first ✓ understood

    Synthesizing natural-sounding speech from text, using neural vocoders and attention-based models.

In the frontier

Rank
#64 of 100
Citations
87
as of Aug 9, 2026
Published
Apr 2026

Topics: Audio, speech, and omni-modal , Vision-language models , Model architecture

Selection: 1kpapers.com by Together AI, most-cited as of Aug 9, 2026

Nearby papers

Summary in our own words; read the paper for the details. ← All papers