Qwen3.8-Omni-Flash undercuts Gemini 3.8 Flash on price, nears it on benchmarks

Qwen3.8-Omni-Flash undercuts Gemini 3.8 Flash on price, nears it on benchmarks

Qwen released Qwen3.8-Omni-Flash, described as its first multimodal model built for AI agents. The model processes audio and video together, draws conclusions from that combined input, and can use tools on its own to carry out tasks such as editing vlogs, translating short videos, and summarizing movies. Its context window spans one million tokens. On audio-video tasks, Qwen says the model comes close to matching Google's Gemini 3.8 Flash.

The pricing gap between the two is wide. Qwen3.8-Omni-Flash charges $0.15 per million input tokens and $0.47 per million output tokens through its API. Qwen estimates that audio input costs under $0.01 per hour to process, and that 720p video with audio at one frame per second costs about $0.20, not counting the cost of generating a response. Gemini 3.8 Flash, by contrast, charges $0.75 per million input tokens and $3.75 per million output tokens at its current introductory rate, and those prices are set to double on January 1, 2027.

Qwen3.8-Omni-Flash is available now through Qwen Studio, Qwen Cloud, and the API. Qwen has also released open-source Qwen-MM-Plugins that add video editing, speaker recognition, PDF-to-video notes, and reusable workflows to coding agents including Claude Code, Gemini CLI, and Qwen Code. A separate tool, Qwen-Live Harness, lets the model interact in real time using a camera and microphone.

Key facts

  • Qwen released Qwen3.8-Omni-Flash, its first multimodal model built for AI agents; it processes audio and video together and has a one million token context window.
  • The API charges $0.15 per million input tokens and $0.47 per million output tokens, undercutting Gemini 3.8 Flash's introductory $0.75 input and $3.75 output per million tokens.
  • Qwen estimates audio input costs under $0.01 per hour and 720p video with one frame-per-second audio costs about $0.20, not counting response generation.
  • Qwen says the model comes close to matching Gemini 3.8 Flash on audio-video tasks, without citing specific benchmark scores.
  • Gemini 3.8 Flash's introductory prices are set to double on January 1, 2027, while Qwen3.8-Omni-Flash ships now via Qwen Studio, Qwen Cloud, and the API, alongside the open-source Qwen-MM-Plugins and the Qwen-Live Harness real-time tool.

Why it matters

Qwen3.8-Omni-Flash is Qwen's first model built specifically for multimodal AI agents, combining audio and video understanding with tool use in a single system. The headline is the price: at about a fifth of Gemini 3.8 Flash's introductory rate for input tokens and about an eighth for output tokens, it pushes down the cost of running audio-video AI agents at scale, while Qwen says the model's performance stays close to Gemini's on the same tasks.

Who it affects

Developers building agents that watch or listen, editing vlogs, translating short videos, summarizing movies, stand to gain the most from the lower per-token cost. The open-source Qwen-MM-Plugins extend that reach to coding agents such as Claude Code, Gemini CLI, and Qwen Code, and Qwen-Live Harness targets anyone building real-time camera-and-microphone interactions. Teams currently paying Gemini 3.8 Flash's rates for similar workloads have a concrete price comparison to weigh.

How to use it

Qwen3.8-Omni-Flash is available now through Qwen Studio, Qwen Cloud, and the API, priced at $0.15 per million input tokens and $0.47 per million output tokens. Qwen-MM-Plugins, distributed as open source, add video editing, speaker recognition, PDF video notes, and reusable workflows on top of agents including Claude Code, Gemini CLI, and Qwen Code. Qwen-Live Harness enables real-time interaction through a camera and microphone.

How solid is it

The pricing figures come directly from each company's own published rates for its own product, which makes the cost comparison concrete. The performance claim is weaker: the source gives no specific benchmark names or scores, only Qwen's own qualitative statement that the model 'comes close to matching' Gemini 3.8 Flash on audio-video tasks. That is Qwen's characterization of its own model against a competitor's, not an independent evaluation.

Risks and caveats

No release date is given for Qwen3.8-Omni-Flash itself, so it is unclear how long the stated prices and capabilities have been in effect. Gemini 3.8 Flash's introductory prices are set to double on January 1, 2027, but the source does not say whether Qwen's own pricing carries a similar promotional limit. Because the performance comparison rests on Qwen's own wording rather than published benchmark numbers, the actual gap to Gemini 3.8 Flash could turn out smaller or larger than 'close' suggests once tested independently.