GLM-5.3-Flash

· 最后更新

模型 ID:glm-5.3-flash

能力

文本

价格

当前价格生效于 2026-09-23

计费项 价格
输入 / 百万 tokens 0.1095 USD · 官方 0.15 USD · 较官方省 27%
输出 / 百万 tokens 0.365 USD · 官方 0.5 USD · 较官方省 27%
缓存读取 / 百万 tokens 0.0219 USD · 官方 0.03 USD · 较官方省 27%

推理档位

平台分两步处理推理档位:先把各客户端的写法换成统一档位,再按上游实际接受的档位发送。上游不接受请求的档位时改用更高一档,高过最高档时取最高档;改写记在用量明细里。

客户端写法 → 平台档位

客户端与接口 请求里的写法 平台档位
Codex · Responses reasoning.effort 原值
OpenCode / pi / Hermes · Chat reasoning_effort 原值
Claude Code · Anthropic thinking.type=enabled,budget_tokens 1–512 minimal
Claude Code · Anthropic budget_tokens 513–1024 low
Claude Code · Anthropic budget_tokens 1025–8192 medium
Claude Code · Anthropic budget_tokens 8193–24576 high
Claude Code · Anthropic budget_tokens 大于 24576 xhigh
Claude Code · Anthropic thinking.type=enabled,不带 budget_tokens auto
Claude Code · Anthropic thinking.type=adaptive,带 output_config.effort output_config.effort 的值
Claude Code · Anthropic thinking.type=adaptive,不带 effort xhigh
Claude Code · Anthropic thinking.type=disabled none
Gemini · generateContent thinkingConfig.thinkingLevel 原值
Gemini · generateContent thinkingConfig.thinkingBudget 同上预算区间;-1 为 auto,0 为 none

平台档位 → 发给上游的档位

平台档位 发给上游
none 不支持关闭思考,返回 422 model.reasoning_off_unsupported
minimal low
low low
medium high
high high
xhigh xhigh / max(不同线路支持的档位不同)
max max
auto high

调用

OpenAI 兼容接口地址 https://omnimodel.me/v1,model 参数填 glm-5.3-flash。