Qwen3.8-Omni-Flash Matches Gemini Flash Benchmarks at Lower Pricing
- Published
- Sep 19, 2026 — 14:30 UTC
On September 19, 2026, Qwen announced the Qwen3.8-Omni-Flash, the first multimodal model designed for AI agents. This model features a context window of one million tokens and offers competitive pricing through the Qwen API at $0.15 per million input tokens and $0.47 per million output tokens. In contrast, Google's Gemini 3.8 Flash charges $0.75 per million input tokens and $3.75 per million output tokens, with a price increase set to double these rates on January 1, 2027.
Qwen claims that the Qwen3.8-Omni-Flash closely matches the performance benchmarks of Gemini 3.8 Flash, making it a viable alternative for developers. The estimated costs for audio input are under $0.01 per hour, while 720p video with audio costs about $0.20 per hour. This pricing structure positions Qwen3.8-Omni-Flash as a cost-effective solution for developers looking to leverage multimodal capabilities without the impending price hikes associated with Gemini.
Qwen offers access to its models through Qwen Studio, Qwen Cloud, and the Qwen API, along with open-source Qwen-MM-Plugins that enhance functionality. Compatibility with agents like Claude Code, Gemini CLI, and Qwen Code allows for flexible integration into existing workflows. This follows previous announcements regarding Google's AI initiatives, emphasizing the competitive landscape in multimodal AI development.
By Callan Zhang · Sep 19, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: The Decoder
