205 callable models — 103 text across 12 vendors, 23 image, 71 video, generated from the pricing registry
The platform currently offers 205 callable models: 103 text chat models across 12 vendors, plus 23 image and 71 video generators, and 8 music / voice / avatar SKUs (last section on this page). Text pricing is in credits per 1M tokens; images bill per image and video per second.
This page is generated from the same pricing registry the gateway bills against, so every id here is callable and every price is the one you are charged.
⚠️ The 205 here and the 211 on the model plaza count different things — this is not drift: the plaza additionally shows not-yet-launched placeholder cards (badged "coming soon"), while this page lists only ids you can call today. A few API-only SKUs appear on neither. Live data: GET https://zhonkezhonkeapi.dflop.top/api/v1/models/public, or the model plaza.
Cached pricing needs no opt-in: when upstream usage reports cached_tokens, the gateway settles that portion of the input at the cached rate automatically — clients don't (and can't) enable it explicitly. Models without a cached figure bill all input at the Input price.
Only Anthropic Claude charges for cache writes. The first time Claude writes a prefix into the cache, those tokens bill at Input × 1.25 (cache_creation_input_tokens); later hits bill at the cached rate above (Input × 0.1). The cache lives about 5 minutes, refreshed on every hit — if no follow-up turn reuses the prefix within that window, the write premium buys nothing. Every other vendor (OpenAI / xAI / Alibaba / Zhipu / Moonshot / DeepSeek / MiniMax / Tencent / Volcengine) only has a hit discount and does not bill writes separately.
Limited-time promotion: the whole Anthropic Claude line is charged at 40% off (list price × 0.6) and the whole OpenAI GPT text line at 70% off (× 0.3), with cached rates discounted equally. The tables show list prices; actual billing always uses the discounted rate, for both API-key calls and in-platform conversations. Discounted models carry an "X% OFF" badge on the model plaza.
Long-context tier: for grok-4.5 / grok-4.6, a prompt of ≥200K tokens bills input, cache-hit and output all at 2× the table price. That is xAI's whole-turn doubling for very long requests, passed straight through (verified to the cent on 2026-08-23). The trigger is the total prompt size of the turn, independent of how much was generated. No other model has this tier.
Server-side tools bill per call: when a grok model runs web search / tools, the upstream charges 2.022 points per tool call on top of tokens. The count comes from the upstream's usage.num_server_side_tools_used and is written to the server_tool_calls column of every usage-log row, so you can re-derive your bill. Turns that use no tools incur no such charge.
gpt-image-2 also accepts POST /v1/images/edits for reference-image edits. The upstream ignores size — put the aspect ratio in the prompt; latency is 30–215s.
The two Midjourney entries are priced per image, but one request always produces 4 of them (a 2×2 grid) and n does not change that — a single request costs 161.76 (v8.1) / 129.41 (v7) credits. size only sets the aspect ratio, never exact pixels; the output tier goes at the end of the prompt (--sd / --hd). Integration details in Image / video APIs.
Since 2026-07-08 the three nano entries are the only way in: older ids (gemini-2.5-flash-image, gemini-3-pro-image-preview, gemini-3.1-flash-image(-preview), tvod-nano-*) keep working as aliases onto the canonical entries above.
Billed per second as an async task through POST /v1/videos/generations. Where tiers are listed (resolution, or operation for the subtitle SKU) you are billed at the tier you request. See Image / video APIs.
480p 33.16 / 720p 59.45 / 1080p 147.61 / 2k 291.17 / 4k 355.87 per second; token-billed turns (reference video present; resolution required) by delivery tier: 480p 3301.5532/1M default, 2009.6411/1M with video input, 720p 2752.1667/1M default, 1675.232/1M with video input, 1080p 3037.1605/1M default, 1848.7064/1M with video input
doubao-seedance-2.0-fast
Seedance 2.0 Fast
480p 23.86 / 720p 47.72 / 1080p 117.28 / 2k 141.54 / 4k 169.85 per second; token-billed turns (reference video present; resolution required) by delivery tier: 480p 2375.5078/1M default, 1445.9613/1M with video input, 720p 2209.2223/1M default, 1344.7441/1M with video input, 1080p 2413.0865/1M default, 1468.8353/1M with video input
doubao-seedance-2.0-fast-lite
Seedance 2.0 Fast Lite
720p 24.67 / 1080p 52.57 per second; token-billed turns (reference video present; resolution required) by delivery tier: 720p 2456.0335/1M default, 1494.977/1M with video input, 1080p 2433.8889/1M default, 1481.4976/1M with video input
doubao-seedance-2.0-lite
Seedance 2.0 Lite
720p 29.93 / 1080p 64.7 per second; token-billed turns (reference video present; resolution required) by delivery tier: 720p 2979.4505/1M default, 1813.5786/1M with video input, 1080p 2995.5556/1M default, 1823.3817/1M with video input
doubao-seedance-2.0-mini
Seedance 2.0 Mini
480p 14.96 / 720p 29.93 per second; token-billed turns (reference video present; resolution required) by delivery tier: 480p 1489.7253/1M default, 906.7894/1M with video input, 720p 1385.4445/1M default, 843.3141/1M with video input
doubao-seedance-2.0-mini-lite
Seedance 2.0 Mini Lite
720p 16.18 / 1080p 34.78 per second; token-billed turns (reference video present; resolution required) by delivery tier: 720p 1610.5138/1M default, 980.3128/1M with video input, 1080p 1610.1112/1M default, 980.0677/1M with video input
doubao-seedance-2.5
Seedance 2.5
480p 42.06 / 720p 90.59 / 1080p 226 per second; token-billed variants 4165.32/1M (default) and 2499.19/1M (with video input); the 1080p tier bills at its own rate — 4650.21/1M (default) and 2790.12/1M (with video input)
doubao-seedance-2.5-lite
Seedance 2.5 Lite
720p 44.48 / 1080p 95.84 per second; token-billed variants 4424.14/1M (default) and 2758.01/1M (with video input)
Both Wan 3.0 cards share one capability set: 2-30s, 480P/720P/1080P, 30fps, aspect 16:9 / 4:3 / 1:1 / 3:4 / 9:16 / adaptive; text-to-video, image-to-video (first frame, first+last frame), and image / video / audio references. Prime is the accelerated tier (same quality, markedly faster) at 1.5x the standard rate. The default delivery is 1080P — pass resolution explicitly for a cheaper tier.
Different models support different client SDKs, and the compatibility matrix is authoritative on what is actually reachable. The catalog's supported_protocols field reflects only the native / primary path; many models are additionally reachable on other protocols through gateway translation (the response then carries an X-Protocol-Translation header). Native paths, counted from the registry that generated this page:
OpenAI Chat (/v1/chat/completions) — all 103 text models
For a model × protocol combination with no reachable channel, the gateway returns 503 no_channel_available (the model exists but has no channel on that protocol — not a 404); switch to a supported protocol.
135 older ids still resolve onto the canonical entries above, so integrations pinned to them keep working. An id that is neither in the tables above nor an alias returns model_not_found — the catalog endpoint is the source of truth.
The response looks like { "models": [...] }. Note that this endpoint returns every registered entry — including placeholder SKUs that aren't live yet and non-chat SKUs — so filter on the three fields below when consuming it from a script rather than looping over the whole list:
Field
Meaning
callable
false means a placeholder SKU (listed but not live); calling it returns model_not_found
endpoint_type
null means an ordinary chat model. images_generations, videos_generations, contents_generations_tasks and similar mean the model uses its own dedicated endpoint and cannot be sent to /v1/chat/completions — see Image / video / music APIs
supported_protocols
the client protocols available for this model (see above)
To take only the models you can send straight to /v1/chat/completions: