Kurdish Speech Logo
Kurdish Speech
← Back to articles
Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use
Dialogue, Chatbots, QA & Agents

Alibaba Qwen Releases Qwen3.8-Omni-Flash: A 1M-Context Omni-Modal Model Built Around Agentic Audio-Video Understanding and Tool Use

Alibaba’s Qwen team has released Qwen3.8-Omni-Flash . They called it its first omni-modal model built around agentic capabilities. It accepts text, images, audio, and video, and it returns text. Audio-video understanding, reasoning, and tool use sit inside one model. The stated workflow is simple: understand the content, plan the task, execute with tools, deliver the result. Is it deployable? Yes, as a hosted API today. It is live on QwenCloud , Alibaba Cloud Model Studio , and Qwen Studio . No open weights were announced at launch, so self-hosting is not an option. What is Qwen3.8-Omni-Flash The Qwen3.8-Omni-Flash is built on the Qwen3.8-Flash-Next architecture. That base model shipped with open weights in August 2026. The context window is 1M tokens. QwenCloud lists 991K max input and 131K max output. Max reasoning length is 262K tokens. Output is text only. The Model Studio docs point developers to Qwen3.5-Omni when they need generated speech. Thinking is on by default, with reasoning_effort set to xhigh . Setting it to none disables thinking. The API follows both the DashScope and OpenAI protocols. It works with Chat Completions and the Responses API. Function calling, web search, structured outputs, context caching, and batch calls are supported. Agentic Perception for Long Video Most video models read a long file from start to finish. That holds even when the answer sits in 3 minutes of footage. Qwen research team describes a different path. The agent starts from the question. It decides what to watch and hear. It then gathers evidence over several coarse-to-fine roun

Source: MarkTechPost

Source: MarkTechPost