TokenRaAI Model Catalog
GLM-5.3-FlashX API · Up to 200 tokens/s · Usage-based pricing via TokenRa

GLM-5.3-FlashX API — Up to 200 tok/s, Same Intelligence

Officially launched by Z.ai on September 18, 2026. GLM-5.3-FlashX delivers up to 200 tokens/s while retaining the same frontier intelligence as GLM-5.3-Flash. Backed by 100,000 domestic chips and optimized inference infrastructure. Usage-based pricing through TokenRa.

Up to 200 tok/sFastest in the GLM-5.3 family.
$0.14 inputPer 1M input tokens.
$0.52 outputPer 1M output tokens.

200 tok/s on the same frontier intelligence

GLM-5.3-FlashX is the high-speed variant of GLM-5.3-Flash, officially launched by Z.ai on September 18, 2026. Previously previewed under the code name Ox Alpha, GLM-5.3-Flash has seen rapid adoption among developers worldwide. Facing surging demand, Z.ai invested further in Infra and inference optimization, backed by 100,000 domestic chips, to launch FlashX at up to 200 tokens/s.

01

Maximum throughput

Up to 200 tokens/s for latency-sensitive workloads: real-time chat, agentic loops, IDE copilots, and high-concurrency production.

02

Same intelligence

FlashX retains the full reasoning, coding, and multimodal capabilities of GLM-5.3-Flash — no quality trade-off.

03

Aligned pricing

Same usage-based pricing as GLM-5.3-Flash: $0.14/M input, $0.52/M output. Speed without the premium.

GLM-5.3-FlashX vs GLM-5.3-Flash

FlashX is the faster variant. Both share the same architecture and intelligence. Choose based on your latency requirements.

DimensionGLM-5.3-FlashGLM-5.3-FlashX
SpeedStandardUp to 200 tok/s
IntelligenceFrontierFrontier (same)
Architecture320B total / 18B active MoE320B total / 18B active MoE
Context~1.3M tokens~1.3M tokens
ModalitiesText · Image · Video → textText · Image · Video → text
Pricing (input)$0.14 / 1M tokens$0.14 / 1M tokens
Pricing (output)$0.52 / 1M tokens$0.52 / 1M tokens
InfrastructureStandard inference100K domestic chips + optimized inference

Both models are available

GLM-5.3-FlashX does not replace GLM-5.3-Flash. GLM-5.3-Flash remains available for workloads where standard speed is sufficient. Confirm current pricing and rate limits for both models in the live TokenRa console.

Where GLM-5.3-FlashX fits

FlashX is optimized for workloads where throughput and latency are critical, without compromising on intelligence.

  • Real-time chat and conversational agents that require sub-second response times.
  • Agentic loops with multiple sequential tool calls where cumulative latency matters.
  • IDE copilots and inline code completions that demand instant feedback.
  • High-concurrency production serving with strict latency budgets.
  • Long-context reasoning where output token generation speed dominates end-to-end latency.

How to get a GLM-5.3-FlashX API key

Register for TokenRa, open the dashboard, create an API key, confirm that GLM-5.3-FlashX is enabled for your account, and then send an OpenAI-compatible request from your server. Keep the key private and verify the live model identifier before production use.

1

Create a TokenRa account

Register, then open the TokenRa dashboard.

2

Create your API key

Open the API key section, create a key, and store it server-side.

3

Send a test request

Confirm the enabled model ID glm-5.3-flashx, test a representative workload, and add production controls.

curl https://tokenra.io/v1/chat/completions \
  -H "Authorization: Bearer YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"glm-5.3-flashx","messages":[{"role":"user","content":"Write a Python function to sort a list"}]}'

GLM-5.3-FlashX pricing

GLM-5.3-FlashX pricing is aligned with GLM-5.3-Flash: $0.14 per 1M input tokens, $0.52 per 1M output tokens, and $0.04 per 1M cached read tokens through TokenRa. No credit card is required to get started. Confirm current rates in the live TokenRa console before production use.

Input

$0.14 / 1M tokens

Standard input pricing for GLM-5.3-FlashX.

Output

$0.52 / 1M tokens

Generated output pricing for GLM-5.3-FlashX.

Rate limits

Account-based

Check the TokenRa console for current RPM and quota limits.

Z.ai (Zhipu AI)

GLM-5.3-FlashX is developed and operated by Z.ai (Zhipu AI). It was previously previewed under the code name Ox Alpha. TokenRa provides API access as a routing gateway, and is not the model developer or owner.

Data handling

Prompts and completions are retained by the provider and are not used for training. Review the applicable provider and TokenRa terms before sending sensitive or regulated information.

Production checklist

  • Keep API keys on the server and separate credentials by environment.
  • Confirm the 200 tok/s throughput and supported input modalities against the live integration before relying on them.
  • Run representative coding, agentic, reasoning, and visual-context workloads.
  • Set explicit timeouts, bounded retries, tracing, and application-side budgets.
  • Review data handling, provider terms, privacy requirements, and output rights for your use case.

GLM-5.3-FlashX FAQ

What is GLM-5.3-FlashX?

GLM-5.3-FlashX is the high-speed variant of GLM-5.3-Flash, officially launched by Z.ai on September 18, 2026. It delivers up to 200 tokens/s while retaining the same frontier intelligence.

How fast is GLM-5.3-FlashX compared to GLM-5.3-Flash?

FlashX reaches up to 200 tokens/s, making it ideal for real-time and high-concurrency workloads. GLM-5.3-Flash runs at standard speed with the same intelligence.

Is GLM-5.3-FlashX more expensive?

No. Pricing is aligned with GLM-5.3-Flash: $0.14/M input and $0.52/M output tokens. Check the live console for current rates.

How do I get an API key?

Register for TokenRa, create an API key, and use the model identifier glm-5.3-flashx in your requests.

Does GLM-5.3-FlashX support the same input types?

Yes — text, image, and video input (to text output), subject to the enabled integration.

Who develops GLM-5.3-FlashX?

Z.ai (Zhipu AI). TokenRa is an API routing gateway and is not the model developer.

GLM-5.3-FlashX is developed by Z.ai (Zhipu AI). The provider and related names belong to their respective rights holders. TokenRa provides API access as a routing gateway and is not the model developer or owner.