You send it text, images, or video, and it answers, extracts, or writes code from what it sees. Video is the part that stands out: it accepts an hour-long recording as input and answers questions about what happened across the whole thing, not just a few sampled frames. It also reads about 260,000 tokens in one prompt, a few books' worth.
Every answer starts in thinking mode, and you set the effort: low for a quick caption, high for a chart that needs careful reading. Qwen reports it ahead of its own Qwen3.8-27B on long video understanding, real-world photo questions and chart analysis, and ahead of DeepSeek V4 Flash on most coding and agent tests.
It holds 125 billion parameters but uses only 6 billion for each word, which is why it responds like a much smaller model. The weights are free to use commercially unless you serve more than 100 million monthly users or sell model access itself.