About this project

Qwen3-VL is the multimodal large language model series from the Qwen team at Alibaba Cloud. The repository documents the model family, its capabilities, architecture updates, benchmarks, cookbooks, and quickstart code for inference with Transformers and ModelScope. Model family and releases The series spans multiple sizes and architectures, described as scaling from edge to cloud. Released checkpoints include Qwen3-VL-2B, 4B, 8B, 32B, 30B-A3B, and 235B-A22B, each offered in Instruct and Thinking editions. FP8 versions are also listed. Earlier generations referenced in the news section include Qwen2.5-VL, Qwen2-VL, and QvQ-72B-Preview. Documented capabilities - Visual agent behavior: operating PC and mobile GUIs by recognizing elements, understanding functions, invoking tools, and completing tasks. - Visual coding: generating Draw.io, HTML, CSS, and JS from images or videos. - Spatial perception: judging object positions, viewpoints, and occlusions, with 2D grounding and 3D grounding for spatial reasoning and embodied AI. - Long context and video: native 256K context expandable to 1M, handling books and hours-long video with second-level indexing. - Multimodal reasoning: STEM and math tasks with causal analysis and evidence-based answers. - Visual recognition: broad pretraining coverage for celebrities, anime, products, landmarks, and flora/fauna. - OCR: support for 32 languages, robustness in low light, blur, and tilt, plus long-document structure parsing. - Text understanding described as on par with pure LLMs through text-vision fusion. Architecture updates The README lists three architectural elements: Interleaved-MRoPE for full-frequency positional allocation across time, width, and height; DeepStack for fusing multi-level ViT features; and Text-Timestamp Alignment for timestamp-grounded event localization in video. Cookbooks A set of Jupyter notebooks covers omni recognition, document parsing, 2D grounding, OCR and key information extraction, video understanding, mobile agent, computer-use agent, 3D grounding, thinking with images, multimodal coding, long document understanding, and spatial understanding. Each is linked with a Colab badge. Quickstart Usage requires transformers >= 4.57.0. Examples show loading with AutoModelForImageTextToText and AutoProcessor, applying a chat template, and generating output. Additional sections cover multi-image inference, video inference, batch inference with left padding, and pixel budget control through the official processor for images and videos, including fps and num_frames settings. The qwen-vl-utils toolkit (version 0.0.14) is documented with image_patch_size and return_video_metadata options. Resources The repository links to Qwen Chat, Hugging Face and ModelScope collections, a blog post, an arXiv paper, a demo Space, WeChat, Discord, an API reference, and PAI-DSW.