Qwen3.8-Omni-Flash Targets Agentic Audio and Video Workflows
Qwen has introduced Qwen3.8-Omni-Flash, a native omnimodal model for text, image, audio, and video inputs, with a 1M token context window and a focus on agentic production workflows such as video editing, meeting follow-up, localization, and real-time interaction.
Qwen has released Qwen3.8-Omni-Flash, described by QwenTeam as its next-generation native omnimodal model. According to the source, the model supports text, image, audio, and video input and provides a 1M token context window. Qwen frames the release less as a perception-only upgrade and more as a step toward agents that can plan tasks, call tools, and deliver finished work.
Source facts: Qwen says Qwen3.8-Omni-Flash improves the average score across 29 evaluations by more than 25% compared with Qwen3.5-Omni-Plus. It also claims API pricing per hour of audio input is down by more than 98%, while audio-video input pricing per hour is down by more than 93%. The source notes that these pricing estimates use the cost of two minutes of source material multiplied by 30; audio-video input uses 720p and 1 fps, with additional API assumptions listed in the original post.
The model is positioned for long audio and video understanding. Qwen reports gains of 36.5 points on WildClawBench-MM, 22.3 points on AgenticVBench, and a score of 69.6 on UniClawBench. For long media tasks, the post says Qwen3.8-Omni-Flash can decide what to watch or listen to across multiple rounds instead of processing every frame. On OmniVideoBench, Qwen reports that agentic understanding raises accuracy from 63.4 to 67.8 while reducing token use from 145,736 to 79,117, a reduction of about 45.7%.
For creators, the most relevant source examples are production pipelines. Qwen describes Music2MV workflows that analyze song structure, rhythm, mood, vocals, and instrumental changes, then support timestamped lyrics and shot planning through Qwen-MM-Plugins. It also describes short-drama translation workflows covering speaker-aware dialogue recognition, conversational translation, voice cloning and dubbing, mixing, and quality review. Another example covers long-film commentary, where an agent plans narration, selects key plot points, prepares voiceover and music, edits, renders, and reviews the final output.
Qwen also announced open-source tooling around the model. Qwen-MM-Plugins is expanded with long audio and long video perception, tool calling, and workflow execution. Qwen-Live Harness is introduced as a native runtime for continuous real-time omnimodal interaction. For real-time use, Qwen also introduces Qwen3.8-Omni-Flash-Realtime, designed to perceive and respond to live audio-video streams while calling tools and executing tasks.
AiPix analysis: this release matters because it points toward a creator workflow where video, audio, transcripts, references, and instructions become part of one agentic workspace. For visual creators, the practical question is not only whether a model can describe media, but whether it can preserve intent across editing, localization, summarization, and delivery. Qwen’s examples align with the same direction AiPix sees in AI Canvas workflows: creators increasingly want one place to inspect source media, issue natural-language changes, compare outputs, and move from idea to asset without stitching together many separate tools.
The source also reports a model-development experiment: Qwen3.8-Omni-Flash was tasked with improving Qwen2.5-Omni-3B’s Sichuan dialect speech recognition within 12 hours. Qwen says the agent selected WenetSpeech-Chuan, created 3,413 training samples over four experiment rounds, and reduced character error rate from 25.79% to 15.30%, a relative reduction of about 40.7%. This should be read as Qwen’s own reported experiment, not independent validation.
For teams building media products, the release is a signal that omnimodal agents are moving from passive understanding toward task completion. The strongest near-term opportunities are likely in repetitive creator operations: finding relevant clips, producing structured notes from video, translating short-form series, drafting commentary, and preparing reviewable cuts. The open questions remain deployment cost, latency, reliability, and how consistently these workflows perform outside Qwen’s showcased scenarios.
Sources
- Qwen 发布原生全模态模型 Qwen3.8-Omni-Flash,主打音视频智能体任务交付Qwen:Blog Retrieval(API)
Continue in AiPix
Turn this update into a practical image or canvas workflow.
Open AI Canvas