TokenRaAI Model Catalog
Qwen3.8-Omni-Flash API · Native omni-modal via TokenRa

Qwen3.8-Omni-Flash API: $0.23/M Input, $0.76/M Output

Build with Qwen's first native omni-modal model for agentic work. Accepts text, image, audio, and video input with 1M-token context. Follow the quick guide below to get an API key and send your first request.

$0.23 inputPer 1M input tokens.
$0.76 outputPer 1M output tokens.
1M contextNative text, image, audio & video.

Native omni-modal. Built for agentic delivery.

Qwen3.8-Omni-Flash is Alibaba Qwen's first native omni-modal model, purpose-built for agentic capabilities in real-world productivity scenarios. It natively accepts text, image, audio, and video input with a 1M-token context window, delivering significant performance gains in multimodal understanding while maintaining text performance comparable to text-only models of the same class.

01

Omni-modal input

Natively accepts text, image, audio, and video input. Supports 2-channel and 4-channel spatial audio understanding for production media workflows.

02

Agentic workflows

Built for coding, knowledge work, GUI interaction, and sustained agentic tasks. Includes function calling, structured outputs, and web search capabilities.

03

1M-token context

Process entire codebases, long videos, multi-document collections, and extended transcripts in a single request with 991K max input and 131K max output.

Where Qwen3.8-Omni-Flash fits

Designed for workloads that combine multiple input modalities with sustained agentic execution.

  • Video post-production, music video creation, and multimedia narration workflows.
  • Film and video production requiring integrated text, image, audio, and video processing.
  • Multimedia summarization and audio-video dialogue with spatial audio understanding.
  • Coding and software engineering with repository-level context and agentic tool use.
  • GUI interaction and knowledge work requiring structured outputs and function calling.

How to get a Qwen3.8-Omni-Flash API key

Register for TokenRa, open the dashboard, create an API key, confirm that Qwen3.8-Omni-Flash is enabled for your account, and then send an OpenAI-compatible request from your server. The model outputs text only — do not set audio modalities in your request.

1

Create a TokenRa account

Register, then open the TokenRa dashboard.

2

Create your API key

Open the API key or token section, create a key, and store it server-side.

3

Send a test request

Use the model ID qwen3.8-omni-flash, test with your preferred modality, and add production controls.

curl https://tokenra.io/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"qwen3.8-omni-flash","messages":[{"role":"user","content":"Analyze this video transcript and summarize the key decisions made."}]}'

Qwen3.8-Omni-Flash pricing: $0.23 input, $0.76 output

Qwen3.8-Omni-Flash uses transparent usage-based pricing through TokenRa: $0.23 per 1M input tokens, $0.76 per 1M output tokens, and $0.03 per 1M cached read tokens. No credit card is required to get started. Access is subject to account eligibility, provider availability, and applicable rate limits; verify the live TokenRa console before production use.

Input

$0.23 / 1M tokens

Standard input pricing for text, image, audio, and video tokens.

Output

$0.76 / 1M tokens

Generated text output pricing — text-only output.

Rate limits

2M TPM · 30K RPM

Check the TokenRa console for current RPM and quota limits.

Alibaba Qwen Team

Qwen3.8-Omni-Flash is developed and operated by Alibaba's Qwen team, built on the Qwen3.8-Flash-Next architecture. It is Qwen's first native omni-modal model designed for agentic capabilities. TokenRa provides API access as a routing gateway, and is not the model developer or owner.

Built-in tools and capabilities

Qwen3.8-Omni-Flash includes a full suite of production-ready features accessible through OpenAI-compatible and DashScope protocols.

  • Function calling — connect the model with external tools and systems for agentic workflows.
  • Structured outputs — ensure the model returns JSON in the expected format for downstream processing.
  • Web search — enable real-time retrieved data for answers that require up-to-date information.
  • Context caching — reduce repeated computation, improve latency, and lower cost for long-context requests.
  • Batch processing — asynchronously process requests in batches to reduce costs at scale.
  • Prefix completion — make the model continue strictly from your provided prefix text with Partial Mode.

Data handling

Review the applicable provider and TokenRa terms before sending sensitive or regulated information. Prompts and completions are subject to the provider's data policy; verify current terms before production use.

Production checklist

  • Keep API keys on the server and separate credentials by environment.
  • Confirm the 1M-token context and supported input modalities against the live integration before relying on them.
  • Run representative omni-modal workloads: coding, video analysis, audio-video dialogue, and agentic tool use.
  • Set explicit timeouts, bounded retries, tracing, and application-side budgets.
  • Review data handling, provider terms, privacy requirements, and output rights for your use case.

Qwen3.8-Omni-Flash API FAQ

What is Qwen3.8-Omni-Flash?

Qwen3.8-Omni-Flash is Alibaba Qwen's first native omni-modal model built for agentic capabilities. It accepts text, image, audio, and video input with text-only output, supports a 1M-token context window, and is built on the Qwen3.8-Flash-Next architecture.

How do I get a Qwen3.8-Omni-Flash API key?

Register for TokenRa, open the dashboard, create an API key, confirm that Qwen3.8-Omni-Flash is enabled, and use the key in a server-side integration with OpenAI-compatible endpoints.

How much does Qwen3.8-Omni-Flash cost?

Qwen3.8-Omni-Flash costs $0.23 per 1M input tokens, $0.76 per 1M output tokens, and $0.03 per 1M cached read tokens. No credit card is required to get started. Confirm current availability, rate limits, and account restrictions in the TokenRa console.

Who develops Qwen3.8-Omni-Flash?

It is developed and operated by Alibaba's Qwen team. Built on the Qwen3.8-Flash-Next architecture, it is Qwen's first native omni-modal model designed for agentic capabilities. TokenRa provides API access as a routing gateway, and is not the model developer or owner.

What input does Qwen3.8-Omni-Flash support?

It natively accepts text, image, audio, and video input. It supports 2-channel and 4-channel spatial audio understanding. Output is text-only; do not set audio modalities in the request.

What is the context window size?

Qwen3.8-Omni-Flash supports a 1M-token context window. Standard route limits: max input 991K tokens, max output 131K tokens. Thinking mode has separate input limits.

Qwen3.8-Omni-Flash is developed and operated by Alibaba's Qwen team. Provider and related names belong to their respective rights holders. TokenRa provides API access as a routing gateway and is not the model developer or owner.