Alibaba logo

Qwen3.8-Omni-Flash

AlibabaProprietaryPending Human Review

Qwen3.8-Omni-Flash is a native omnimodal agent model built on the Qwen3.8-Flash-Next architecture. It accepts text, images, audio, and video in a million-token context, extending coding, GUI operation, and knowledge work to audio-visual productivity. Standard APIs return text, while the subsequently available Realtime variant supports text and speech output, multichannel audio, video aggregation, and remote MCP tools. Hosted access and companion plugins support video editing, multimedia summarization, and real-time conversations; its weights are not publicly released.

2026-09-18
Hybrid-attention MoE with native multimodal processing
Proprietary

Specifications

Architecture
Hybrid-attention MoE with native multimodal processing
License
Proprietary
Context Window
1,000,000 tokens
Type
multimodal
Modalities
textimageaudiovideo

Benchmark Scores

Advanced Specifications

Model Family
Qwen
API Access
Available
Chat Interface
Not Available
Variants
Qwen3.8-Omni-Flash-Realtime (2026-09-21 API release)

Capabilities & Limitations

Capabilities
multimodal reasoningaudio-visual agentscodingGUI operationmultichannel audio understanding
Known Limitations
Standard and Realtime endpoints have different output capabilities.Reported agent results depend on companion tools and evaluation harnesses.
Notable Use Cases
multimedia analysisvideo-editing agentsreal-time multimodal assistants
Function Calling Support
Yes
Tool Use Support
Yes

Related Models