Qwen3.7-Flash

· Updated

Model ID: qwen3.7-flash

Capabilities

Text

Pricing

Effective since 2026-09-23

Item Price
Input / 1M tokens 0.021 USD · Official 0.03 USD · Save 30% vs official
Output / 1M tokens 0.091 USD · Official 0.13 USD · Save 30% vs official
Cache read / 1M tokens 0.0042 USD · Official 0.006 USD · Save 30% vs official
Cache write / 1M tokens 0.02625 USD · Official 0.0375 USD · Save 30% vs official

Reasoning effort

The platform handles reasoning in two steps: it turns each client's setting into one level, then sends a level the upstream accepts. A level the upstream rejects becomes the next higher accepted level, or its highest one; rewrites show up in usage records.

Client setting → platform level

Client and API Sent as Platform level
Codex · Responses reasoning.effort same value
OpenCode / pi / Hermes · Chat reasoning_effort same value
Claude Code · Anthropic thinking.type=enabled, budget_tokens 1–512 minimal
Claude Code · Anthropic budget_tokens 513–1024 low
Claude Code · Anthropic budget_tokens 1025–8192 medium
Claude Code · Anthropic budget_tokens 8193–24576 high
Claude Code · Anthropic budget_tokens above 24576 xhigh
Claude Code · Anthropic thinking.type=enabled without budget_tokens auto
Claude Code · Anthropic thinking.type=adaptive with output_config.effort the output_config.effort value
Claude Code · Anthropic thinking.type=adaptive without an effort xhigh
Claude Code · Anthropic thinking.type=disabled none
Gemini · generateContent thinkingConfig.thinkingLevel same value
Gemini · generateContent thinkingConfig.thinkingBudget budget ranges above; -1 is auto, 0 is none

Platform level → level sent upstream

Platform level Sent upstream
none none
minimal minimal
low low
medium medium
high high
xhigh xhigh
max xhigh
auto medium

Usage

OpenAI-compatible base URL https://omnimodel.me/v1, set model to qwen3.7-flash.