Alibaba Qwen has released Qwen3.8-Omni-Flash, a native omni-modal model that processes text, images, audio, and video in a single workflow and supports a 1-million-token context window2. The model is available now through the Qwen AI platform. Alibaba said API input pricing has been reduced to as low as RMB 0.8 per million tokens, and audio input cost has been cut by 98%4.
Alibaba said Qwen3.8-Omni-Flash improved its average score by more than 26% across 30 evaluations compared with Qwen3.5-Omni-Plus3. The company reported gains specifically in audio-video agents, coding, long-context tasks, and real-time multimodal interaction.
The model supports long-video analysis, meeting summaries, video research, and multimodal tool use. Alibaba Qwen also released two companion artifacts: Qwen-MM-Plugins, for long-running multimodal workflows, and Qwen-Live Harness, for real-time workflows.
ANALYSIS The 98% audio-cost reduction and RMB 0.8-per-million-token input pricing place Qwen3.8-Omni-Flash at the aggressive end of multimodal API pricing, particularly for audio- and video-heavy workloads. Pairing the model with dedicated plugin and harness tooling for agentic and real-time use cases positions the release as an integrated serving stack rather than a standalone checkpoint drop.