Omni-modal input
Natively accepts text, image, audio, and video input. Supports 2-channel and 4-channel spatial audio understanding for production media workflows.
Build with Qwen's first native omni-modal model for agentic work. Accepts text, image, audio, and video input with 1M-token context. Follow the quick guide below to get an API key and send your first request.
Qwen3.8-Omni-Flash is Alibaba Qwen's first native omni-modal model, purpose-built for agentic capabilities in real-world productivity scenarios. It natively accepts text, image, audio, and video input with a 1M-token context window, delivering significant performance gains in multimodal understanding while maintaining text performance comparable to text-only models of the same class.
Natively accepts text, image, audio, and video input. Supports 2-channel and 4-channel spatial audio understanding for production media workflows.
Built for coding, knowledge work, GUI interaction, and sustained agentic tasks. Includes function calling, structured outputs, and web search capabilities.
Process entire codebases, long videos, multi-document collections, and extended transcripts in a single request with 991K max input and 131K max output.
Designed for workloads that combine multiple input modalities with sustained agentic execution.
Register for TokenRa, open the dashboard, create an API key, confirm that Qwen3.8-Omni-Flash is enabled for your account, and then send an OpenAI-compatible request from your server. The model outputs text only — do not set audio modalities in your request.
Register, then open the TokenRa dashboard.
Open the API key or token section, create a key, and store it server-side.
Use the model ID qwen3.8-omni-flash, test with your preferred modality, and add production controls.
curl https://tokenra.io/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.8-omni-flash","messages":[{"role":"user","content":"Analyze this video transcript and summarize the key decisions made."}]}'Qwen3.8-Omni-Flash uses transparent usage-based pricing through TokenRa: $0.23 per 1M input tokens, $0.76 per 1M output tokens, and $0.03 per 1M cached read tokens. No credit card is required to get started. Access is subject to account eligibility, provider availability, and applicable rate limits; verify the live TokenRa console before production use.
Standard input pricing for text, image, audio, and video tokens.
Generated text output pricing — text-only output.
Check the TokenRa console for current RPM and quota limits.
Qwen3.8-Omni-Flash is developed and operated by Alibaba's Qwen team, built on the Qwen3.8-Flash-Next architecture. It is Qwen's first native omni-modal model designed for agentic capabilities. TokenRa provides API access as a routing gateway, and is not the model developer or owner.
Qwen3.8-Omni-Flash includes a full suite of production-ready features accessible through OpenAI-compatible and DashScope protocols.
Review the applicable provider and TokenRa terms before sending sensitive or regulated information. Prompts and completions are subject to the provider's data policy; verify current terms before production use.
Qwen3.8-Omni-Flash is Alibaba Qwen's first native omni-modal model built for agentic capabilities. It accepts text, image, audio, and video input with text-only output, supports a 1M-token context window, and is built on the Qwen3.8-Flash-Next architecture.
Register for TokenRa, open the dashboard, create an API key, confirm that Qwen3.8-Omni-Flash is enabled, and use the key in a server-side integration with OpenAI-compatible endpoints.
Qwen3.8-Omni-Flash costs $0.23 per 1M input tokens, $0.76 per 1M output tokens, and $0.03 per 1M cached read tokens. No credit card is required to get started. Confirm current availability, rate limits, and account restrictions in the TokenRa console.
It is developed and operated by Alibaba's Qwen team. Built on the Qwen3.8-Flash-Next architecture, it is Qwen's first native omni-modal model designed for agentic capabilities. TokenRa provides API access as a routing gateway, and is not the model developer or owner.
It natively accepts text, image, audio, and video input. It supports 2-channel and 4-channel spatial audio understanding. Output is text-only; do not set audio modalities in the request.
Qwen3.8-Omni-Flash supports a 1M-token context window. Standard route limits: max input 991K tokens, max output 131K tokens. Thinking mode has separate input limits.
Qwen3.8-Omni-Flash is developed and operated by Alibaba's Qwen team. Provider and related names belong to their respective rights holders. TokenRa provides API access as a routing gateway and is not the model developer or owner.