Maximum throughput
Up to 200 tokens/s for latency-sensitive workloads: real-time chat, agentic loops, IDE copilots, and high-concurrency production.
Officially launched by Z.ai on September 18, 2026. GLM-5.3-FlashX delivers up to 200 tokens/s while retaining the same frontier intelligence as GLM-5.3-Flash. Backed by 100,000 domestic chips and optimized inference infrastructure. Usage-based pricing through TokenRa.
GLM-5.3-FlashX is the high-speed variant of GLM-5.3-Flash, officially launched by Z.ai on September 18, 2026. Previously previewed under the code name Ox Alpha, GLM-5.3-Flash has seen rapid adoption among developers worldwide. Facing surging demand, Z.ai invested further in Infra and inference optimization, backed by 100,000 domestic chips, to launch FlashX at up to 200 tokens/s.
Up to 200 tokens/s for latency-sensitive workloads: real-time chat, agentic loops, IDE copilots, and high-concurrency production.
FlashX retains the full reasoning, coding, and multimodal capabilities of GLM-5.3-Flash — no quality trade-off.
Same usage-based pricing as GLM-5.3-Flash: $0.14/M input, $0.52/M output. Speed without the premium.
FlashX is the faster variant. Both share the same architecture and intelligence. Choose based on your latency requirements.
| Dimension | GLM-5.3-Flash | GLM-5.3-FlashX |
|---|---|---|
| Speed | Standard | Up to 200 tok/s |
| Intelligence | Frontier | Frontier (same) |
| Architecture | 320B total / 18B active MoE | 320B total / 18B active MoE |
| Context | ~1.3M tokens | ~1.3M tokens |
| Modalities | Text · Image · Video → text | Text · Image · Video → text |
| Pricing (input) | $0.14 / 1M tokens | $0.14 / 1M tokens |
| Pricing (output) | $0.52 / 1M tokens | $0.52 / 1M tokens |
| Infrastructure | Standard inference | 100K domestic chips + optimized inference |
GLM-5.3-FlashX does not replace GLM-5.3-Flash. GLM-5.3-Flash remains available for workloads where standard speed is sufficient. Confirm current pricing and rate limits for both models in the live TokenRa console.
FlashX is optimized for workloads where throughput and latency are critical, without compromising on intelligence.
Register for TokenRa, open the dashboard, create an API key, confirm that GLM-5.3-FlashX is enabled for your account, and then send an OpenAI-compatible request from your server. Keep the key private and verify the live model identifier before production use.
Register, then open the TokenRa dashboard.
Open the API key section, create a key, and store it server-side.
Confirm the enabled model ID glm-5.3-flashx, test a representative workload, and add production controls.
curl https://tokenra.io/v1/chat/completions \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"glm-5.3-flashx","messages":[{"role":"user","content":"Write a Python function to sort a list"}]}'GLM-5.3-FlashX pricing is aligned with GLM-5.3-Flash: $0.14 per 1M input tokens, $0.52 per 1M output tokens, and $0.04 per 1M cached read tokens through TokenRa. No credit card is required to get started. Confirm current rates in the live TokenRa console before production use.
Standard input pricing for GLM-5.3-FlashX.
Generated output pricing for GLM-5.3-FlashX.
Check the TokenRa console for current RPM and quota limits.
GLM-5.3-FlashX is developed and operated by Z.ai (Zhipu AI). It was previously previewed under the code name Ox Alpha. TokenRa provides API access as a routing gateway, and is not the model developer or owner.
Prompts and completions are retained by the provider and are not used for training. Review the applicable provider and TokenRa terms before sending sensitive or regulated information.
GLM-5.3-FlashX is the high-speed variant of GLM-5.3-Flash, officially launched by Z.ai on September 18, 2026. It delivers up to 200 tokens/s while retaining the same frontier intelligence.
FlashX reaches up to 200 tokens/s, making it ideal for real-time and high-concurrency workloads. GLM-5.3-Flash runs at standard speed with the same intelligence.
No. Pricing is aligned with GLM-5.3-Flash: $0.14/M input and $0.52/M output tokens. Check the live console for current rates.
Register for TokenRa, create an API key, and use the model identifier glm-5.3-flashx in your requests.
Yes — text, image, and video input (to text output), subject to the enabled integration.
Z.ai (Zhipu AI). TokenRa is an API routing gateway and is not the model developer.
GLM-5.3-FlashX is developed by Z.ai (Zhipu AI). The provider and related names belong to their respective rights holders. TokenRa provides API access as a routing gateway and is not the model developer or owner.