GLM-5.3-FlashX

· Updated

Model ID: glm-5.3-flashx

Capabilities

Text

Pricing

Effective since 2026-09-23

Item Price
Input / 1M tokens 0.259 USD · Official 0.37 USD · Save 30% vs official
Output / 1M tokens 0.875 USD · Official 1.25 USD · Save 30% vs official
Cache read / 1M tokens 0.0525 USD · Official 0.075 USD · Save 30% vs official

Reasoning effort

The platform handles reasoning in two steps: it turns each client's setting into one level, then sends a level the upstream accepts. A level the upstream rejects becomes the next higher accepted level, or its highest one; rewrites show up in usage records.

Client setting → platform level

Client and API Sent as Platform level
Codex · Responses reasoning.effort same value
OpenCode / pi / Hermes · Chat reasoning_effort same value
Claude Code · Anthropic thinking.type=enabled, budget_tokens 1–512 minimal
Claude Code · Anthropic budget_tokens 513–1024 low
Claude Code · Anthropic budget_tokens 1025–8192 medium
Claude Code · Anthropic budget_tokens 8193–24576 high
Claude Code · Anthropic budget_tokens above 24576 xhigh
Claude Code · Anthropic thinking.type=enabled without budget_tokens auto
Claude Code · Anthropic thinking.type=adaptive with output_config.effort the output_config.effort value
Claude Code · Anthropic thinking.type=adaptive without an effort xhigh
Claude Code · Anthropic thinking.type=disabled none
Gemini · generateContent thinkingConfig.thinkingLevel same value
Gemini · generateContent thinkingConfig.thinkingBudget budget ranges above; -1 is auto, 0 is none

Platform level → level sent upstream

Platform level Sent upstream
none Reasoning cannot be turned off; returns 422 model.reasoning_off_unsupported
minimal low
low low
medium high
high high
xhigh max
max max
auto high

Usage

OpenAI-compatible base URL https://omnimodel.me/v1, set model to glm-5.3-flashx.