On this page
- Measure the workload first#
- What I would trial first#
- Normalize the subscription costs#
- Start with the API baseline#
- Compare with OpenRouter too#
- Convert the allowance before comparing it#
- Native coding subscriptions belong in the trial too#
- Temporary offers, kept separate#
- The next analysis should measure accepted work#
- Reproduce and update the comparison#
- Changelog#
- 2026-09-25 — Every tracked subscription, first-party only#
- 2026-09-25 — First-party re-verification#
- 2026-09-07 — Published as a data artifact#
- 2026-09-07 — Display the analysis date#
- 2026-09-07 — Index by subscription and model#
- 2026-09-07 — OpenRouter route comparison#
- 2026-09-07 — Normalize subscription costs#
- 2026-09-07 — Measured usage, supported models and capacity#
- 2026-09-07 — Initial research draft#
curl -s https://jedarden.com/data/inference-plans.md Once coding workers can keep themselves busy, inference becomes a purchasing decision. A subscription buys an allowance that expires. An API bills for the work requested. Both can be economical, and both can be wasted.
In The unit economics of running cattle, I argued for measuring cost per completed outcome. This artifact applies that discipline to buying inference: which offers deserve a trial, what their advertised units mean, and where the apparent savings disappear.
Analysis date: September 25, 2026. Prices, quotas and routes were rechecked against the providers’ own pages that day; the usage measurement below is from September 7.
This is a maintained research artifact. Prices are in USD before tax; individual source checks and usage captures retain their own dates in the dataset. The calculations model billing; they do not establish model quality, sustained throughput, or the cost of a shipped application. The changelog records substantive revisions. Publication dates, editorial updates, and source-verification dates serve different purposes.
Measure the workload first#
The baseline now comes from an actual local usage scan captured September 7, 2026 at 14:55 UTC, using Tokscale 4.15.1. It covers the locally discovered Claude Code, Codex and ZCode histories in one scan, across 871,701 recorded messages. This is an accumulated history snapshot, not one month of usage or a deduplicated total across every machine. The aggregate measurement contains the exact counters and derivation; no prompts or session records are included.
Fresh input2,293,232,314
- Share of all tokens
- 2.794%
- Per 1M output
- 5.985M
Cache reads78,848,276,612
- Share of all tokens
- 96.082%
- Per 1M output
- 205.776M
Cache writes538,768,695
- Share of all tokens
- 0.657%
- Per 1M output
- 1.406M
Output, including billed reasoning383,174,954
- Share of all tokens
- 0.467%
- Per 1M output
- 1.000M
82.06B measured tokens; 213.2:1 total input/output; 96.53% of input served from cache. All input includes fresh input, cache reads and cache writes.
Tokscale’s model report keeps separately reported reasoning outside its output field. I add that bucket once to obtain billable output; reasoning already included by a client stays inside output. Cache reads and writes are separate from fresh input. This normalization follows the versioned scanner implementation.
These are measured usage proportions, transferred into hypothetical purchases below. Another model, tokenizer, provider cache or session pattern can change them. In particular, the high cache-hit rate is not guaranteed on another service. Most processed tokens here are repeated context; maximizing that total alone would reward repetition rather than useful applications.
What I would trial first#
My shortlist from the published economics is OpenCode Go for routine coding, Z.ai for sustained GLM-5.3 work, and Synthetic when its model selection or additional concurrency justifies a pack. Keep a capped API balance available for overflow and difficult tasks. These are candidates to measure in the actual harness, not a ranking of coding quality. The September 25 recheck keeps the shortlist but narrows Go’s role: OpenRouter’s cheapest GLM-5.3 route now undercuts Go’s GLM-5.3 allowance, so Go earns its place on MiMo, MiniMax and Flash-class models, and Z.ai remains the GLM-5.3 option.
Fifteen subscriptions were added on September 25, including the native coding products, China-region editions and several credit-based plans. Three of them publish enough to calculate and are worth measuring. StepFun’s Step Plan states the largest allowance per dollar here: its Credits convert at “approximately $1 ≈ 7M Credits” of list-price usage, with no five-hour or weekly window, so the $29 tier covers about 47B processed Step 3.5 Flash tokens at this mix. Its models have not been benchmarked here. Command Code sells $70 of list-price usage for $10 on GOAT, with per-model caps. GitHub Copilot joins annual Kilo Pass as a quantifiable way to buy Claude Sonnet 5 below Anthropic’s list price, and it also covers Opus 5.5, at 33–50% below list depending on tier. Its credits above the fee are a flex allotment that “may change over time”. ChatGPT, Claude, Cursor, Google AI, Devin and SuperGrok publish limits but no convertible allowance. Their rows quote those limits; turning them into token capacity would take a measurement of my own usage, which this revision does not claim.
The key eligibility question is whether the allowance covers the workload. A developer running a supported coding tool, an unattended coding worker, and a finished application serving customers can fall under different rules. An API-shaped endpoint does not establish permission for all three.
Each expanded plan reports total processed tokens, with output in parentheses, at the measured mix above. Each example spends the entire allowance on its named model; examples within one plan are alternatives. Weekly quotas are prorated to 30 days and assume full utilization within shorter limits. B means billion and M means million tokens. This is quota capacity, not a measured throughput promise or a model’s per-request context window.
OpenCode Go$1033 models / variants · 3 calculations
Supported models / catalog
- MiMo V2.6 Flash
- MiniMax M3
- Kimi K3
- DeepSeek V4 Flash
- GLM-5.3
- GPT-5.6 Luna
- GPT 6 Luna
- MiMo V2.6 Pro
- MiMo V2.5
- MiMo V2.5 Pro
- GLM-5.3-Flash
- DeepSeek V4 Pro
- DeepSeek V4.1 Flash
- Grok 4.6
- Grok 4.7
- GLM-5.2
- GLM-5.1
- Kimi K2.7 Code
- Kimi K2.6
- MiniMax M2.7
- MiniMax M2.5
- Qwen 3.8 Max
- Qwen 3.8 Flash
- Qwen 3.7 Max
- Qwen 3.7 Plus
- Qwen 3.6 Plus
- LongCat 2.0
- Muse Spark 1.3 Contributor
- Muse Spark 1.2 Contributor
- DeepSeek V4 Flash Vision Exp
- Hy4 preview
- Hy3
- Space Bunny
Estimated processed tokens / 30 days
- MiMo V2.6 Flash6.80B total (31.73M output)
- MiniMax M3815.16M total (3.81M output)
- Kimi K332.48M total (0.15M output)
Allowance and practical limit
Monthly API allowance per model: $60 for MiMo V2.6 Flash, MiniMax M3, GLM-5.3-Flash, GLM-5.2/5.1, Kimi K2.7 Code/K2.6, Qwen3.7 Plus and LongCat 2.0; $30 for DeepSeek V4 Flash; $15 for Kimi K3, GLM-5.3, GPT 6 Luna and DeepSeek V4 Pro. One shared pool; each model is also capped at 20% per five hours and 50% per week. Coding-agent traffic only.
Z.ai Lite / Pro / Max$18 / $80 / $1682 models / variants · 3 calculations
Estimated processed tokens / 30 days
- Lite, GLM-5.3 off-peak432.12M total (2.02M output)
- Pro, GLM-5.3 off-peak2.59B total (12.11M output)
- Max, GLM-5.3 off-peak6.05B total (28.25M output)
Allowance and practical limit
10k / 60k / 140k credits weekly; model and cache weights apply. Off-peak usage costs half the credits; peak is weekdays 14:00-18:00 UTC+8. Supported coding tools; dynamic concurrency.
Synthetic$30 per pack11 models / variants · 2 calculations
Supported models / catalog
- syn:large:text
- syn:small:text
- syn:large:vision
- syn:small:vision
- openai/gpt-oss-120b
- zai-org/GLM-5.3-Flash
- deepseek-ai/DeepSeek-V4.1-Flash
- moonshotai/Kimi-K3
- Qwen/Qwen3.8-27B
- zai-org/GLM-4.7-Flash
- nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4
Estimated processed tokens / 30 days
- GLM-5.3-Flash2.24B total (10.45M output)
- Kimi-K3169.75M total (0.79M output)
Allowance and practical limit
$24 of its own API credits weekly; 500 weighted requests per five hours; one concurrent request per model per pack. More packs increase each limit.
camelStream$5 per stream4 models / variants · 1 calculation
Estimated processed tokens / 30 days
- Unmetered; throughput unknown
Allowance and practical limit
Unmetered tokens; one generation per stream. Provider selects the model. Published throughput targets are not a service-level agreement; standard terms permit training on submitted content.
MiniMax Token Plan$22 / $55 / $1322 models / variants · 1 calculation
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
Five-hour and weekly windows; text, image and speech share quota, deducted at pay-as-you-go list prices. The provider publishes conflicting token figures and none for current prices, so no API-equivalent is calculated.
Kimi membership / CodePlus $19 / Pro $39 / Max $99 / Ultra $1993 models / variants · 1 calculation
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
Agent credits scale 1x / 2x / 5x / 10x by tier and share the membership monthly quota; the base amount is not published. New members have only a rolling five-hour window. Extra usage bills near Open Platform API prices. China-region CNY prices are served only to China-region sessions; the Chinese docs say pricing is unchanged under the new tier names.
Alibaba Cloud Coding Plan$5010 models / variants · 1 calculation
Supported models / catalog
- qwen3.7-plus
- qwen3.6-plus
- kimi-k2.5
- glm-5
- MiniMax-M2.5
- qwen3.5-plus
- qwen3-max-2026-01-23
- qwen3-coder-next
- qwen3-coder-plus
- glm-4.7
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
6k requests per five hours, 45k per week, 90k per month; a task can use 5-30+ model calls. Explicitly excludes automated scripts, application backends and noninteractive use. Limited slots, restocked daily.
deepseekv4pro.comCoding $19.90 / Coding Max $49.90 / Agent $49.90 / Coding Pro Max $1995 models / variants · 1 calculation
Estimated processed tokens / 30 days
- Not quantifiable from published terms
Allowance and practical limit
Independent reseller, not affiliated with DeepSeek. Recurring promotional prices and inventory-gated slots. Quotas are published in unweighted "plan tokens"; weighting and upstream plan rules apply.
MiMo Token Plan$6 / $16 / $50 / $1002 models / variants · 4 calculations
Estimated processed tokens / 30 days
- Lite, V2.6 Flash daytime650.12M total (3.04M output)
- Standard, V2.6 Flash daytime1.74B total (8.14M output)
- Pro, V2.6 Flash daytime6.03B total (28.13M output)
- Max, V2.6 Flash daytime13.00B total (60.71M output)
Allowance and practical limit
4.1B / 11B / 38B / 82B credits, with model-specific deductions. These are not raw tokens. Programming-tool restrictions apply.
Chutes$10 / $2014 models / variants · 2 calculations
Supported models / catalog
- google/gemma-4-31B-turbo-TEE
- Qwen/Qwen3.6-27B-TEE
- Qwen/Qwen3.8-27B-TEE
- Qwen/Qwen3.5-397B-A17B-TEE
- zai-org/GLM-5.1-TEE
- deepseek-ai/DeepSeek-V3.2-TEE
- zai-org/GLM-5.2-TEE
- moonshotai/Kimi-K2.6-TEE
- deepseek-ai/DeepSeek-V4-Flash-0731-TEE
- unsloth/Mistral-Nemo-Instruct-2407-TEE
- moonshotai/Kimi-K3-TEE
- Qwen/Qwen3-235B-A22B-Thinking-2507-TEE
- Qwen/Qwen3-32B-TEE
- Nemotron-3-Nano-Omni-30B-TEE
Estimated processed tokens / 30 days
- Plus, Flash conditional ceiling785.87M total (3.67M output)
- Pro, Flash conditional ceiling1.57B total (7.34M output)
Allowance and practical limit
Daily quotas; overage is 6% (Plus) or 10% (Pro) below PAYG rates. The February policy caps subscription value at 5x its PAYG rates. Per-tier quotas are not published.
NanoGPT$12296 models / variants · 2 calculations
Supported models / catalog
- GLM 5.3 Flash
- MiniMax M3
- MiMo V2.6 Flash
- MiMo V2.6 Pro
- GLM 5.3
- DeepSeek V4 Pro 0813
- DeepSeek V4.1 Flash
Showing 7 compared routes from 296 recorded models / variants. Browse the complete catalog.
Estimated processed tokens / 30 days
- 1× input weight258.35M total (1.21M output)
- 2× input weight129.17M total (0.60M output)
Allowance and practical limit
60M input-token units weekly, including cached input, with model multipliers; resets Monday 00:00 UTC. Shared or commercial team workloads must use pay-as-you-go.
Kilo Pass$19 / $49 / $199394 models / variants · 3 calculations
Supported models / catalog
- DeepSeek: DeepSeek V4.1 Flash
- Z.ai: GLM 5.3 Flash
- MoonshotAI: Kimi K3
- MiniMax: MiniMax M3
- Xiaomi: MiMo-V2.6-Flash
- Xiaomi: MiMo-V2.6-Pro
- Z.ai: GLM 5.3
- DeepSeek: DeepSeek V4 Pro 0813
- Anthropic: Claude Sonnet 5
Showing 9 compared routes from 394 recorded models / variants. Browse the complete catalog.
Estimated processed tokens / 30 days
- Starter, Sonnet 5 annual91.59M total (0.43M output)
- Pro, Sonnet 5 annual236.21M total (1.10M output)
- Expert, Sonnet 5 annual959.32M total (4.48M output)
Allowance and practical limit
Credits plus welcome/loyalty bonuses; annual billing gives 50% extra monthly credits. Full use yields a 33.3% effective discount. Bonus credits expire monthly; Gateway supported.
ChatGPT (Codex)Go $8 / Plus $20 / Pro from $1009 models / variants · 1 calculation
Supported models / catalog
- GPT-6 Astra
- GPT-6 Sol
- GPT-6 Luna
- GPT-5.6 Sol
- GPT-5.6 Terra
- GPT-5.6 Luna
- GPT-5.5
- GPT-5.4
- GPT-5.4 mini
Estimated processed tokens / 30 days
- All tiers
Allowance and practical limit
Model-dependent five-hour message allowances, with weekly limits possible; Pro gives 5x or 20x Plus. Extra usage can be bought as ChatGPT credits. Per-token credit rates are published, but the included allowance is not.
Claude (Claude Code)Pro $20 / Max 5x $100 / Max 20x $2003 models / variants · 1 calculation
Estimated processed tokens / 30 days
- All tiers
Allowance and practical limit
Five-hour session limits plus a weekly limit on Max; Max is 5x or 20x Pro, whose allowance is not published. Limits assume ordinary, individual use; developers building products must use API keys.
CursorPro $20 / Pro+ $60 / Ultra $20013 models / variants · 1 calculation
Supported models / catalog
- Grok 4.7
- Grok 4.6
- Grok 4.5
- Composer 2.5
- Claude Fable 5.1
- Claude Opus 5.5
- Claude Sonnet 5
- Gemini 3.1 Pro
- Gemini 3.8 Flash
- GPT-5.6 Sol
- GPT-5.6 Terra
- GPT-5.6 Luna
- Muse Spark 1.3
Estimated processed tokens / 30 days
- All tiers
Allowance and practical limit
Two monthly usage pools, one for Cursor's own models and one for third-party models at their API prices. Per-model rates are published; the included dollar amount per plan is not.
GitHub CopilotPro $10 / Pro+ $39 / Max $1003 models / variants · 3 calculations
Estimated processed tokens / 30 days
- Pro, Claude Sonnet 548.21M total (0.23M output)
- Pro+, Claude Sonnet 5224.97M total (1.05M output)
- Max, Claude Sonnet 5642.76M total (3.00M output)
Allowance and practical limit
1,500 / 7,000 / 20,000 monthly AI credits (1 credit = $0.01), spent at per-token model prices. Base credits equal the fee; the rest is a flex allotment that "may change over time." Chat, CLI and cloud agent draw on it.
Google AI Pro / UltraPro $19.99 / Ultra $99.99 (5x) or $199.99 (20x)3 models / variants · 3 calculations
Supported models / catalog
- Gemini CLI (Gemini model family)
- Antigravity agent models
- Jules
Estimated processed tokens / 30 days
- Pro: 1,500 CLI requests/day
- Ultra 5x: 2,000 CLI requests/day
- Ultra 20x: 2,000 CLI requests/day
Allowance and practical limit
Gemini CLI: 1,500 (Pro) or 2,000 (Ultra) model requests per user per day. Jules: 100 or 300 tasks per rolling 24 hours. Antigravity: tiered rate limits plus an AI credit pool; no token amounts published.
Devin Pro / MaxPro $20 / Max $2002 models / variants · 2 calculations
Estimated processed tokens / 30 days
- Pro: SWE-2 free until Oct 10
- Max: weekly quota, size unpublished
Allowance and practical limit
Pro has a daily and weekly usage quota; Max a larger weekly quota with no daily cap. Quota sizes are not published. Extra usage is prepaid on-demand credit at API pricing. SWE-2 is free in Devin Desktop and CLI through October 10, 2026.
SuperGrokSuperGrok $30 / Plus $100; Lite and Heavy prices not shown on the pricing page2 models / variants · 1 calculation
Estimated processed tokens / 30 days
- Limits not quantified
Allowance and practical limit
Tiers: Lite, SuperGrok, Plus and Heavy. Limits are described only as "higher rate limits" and "significantly higher usage across Chat, Imagine, Voice & Build". No quantities are published.
Ollama Cloud Pro / MaxPro $20 ($200/yr) / Max $1006 models / variants · 3 calculations
Supported models / catalog
- GLM-5.3
- GLM-5.3-Flash
- Kimi K3
- MiniMax M3
- DeepSeek V4.1 Flash
- DeepSeek V4 Pro
Estimated processed tokens / 30 days
- Pro, GLM-5.3188.28M total (0.88M output)
- Pro, DeepSeek V4.1 Flash off-peak5.52B total (25.80M output)
- Max, GLM-5.3941.41M total (4.40M output)
Allowance and practical limit
Pro includes $60 and Max $300 of usage credits per month, deducted at the published per-token model prices; no rollover. Concurrency: Pro 3, Max 10 requests. Off-peak rates for DeepSeek apply outside weekday 12:00-18:00 UTC.
Command CodeGo $1 / GOAT $10 / Pro $20 / Max 10x $100 / Max 20x $2008 models / variants · 4 calculations
Supported models / catalog
- DeepSeek V4 Flash
- GLM-5.2
- Kimi K2.7 Code
- MiMo V2.6 Flash
- Kimi K3
- GLM-5.3
- Claude Sonnet 5
- Claude Opus 5.5
Estimated processed tokens / 30 days
- GOAT, GLM-5.2219.66M total (1.03M output)
- GOAT, DeepSeek V4 Flash5.52B total (25.80M output)
- Pro, Claude Sonnet 564.28M total (0.30M output)
- Max 10x, Claude Opus 5.5232.47M total (1.09M output)
Allowance and practical limit
Monthly credit pools of $10 / $70 / $80 / $150 / $300 at Command Code's per-model token prices, with per-model caps on GOAT and Pro and separate standard/premium limits on Max. Every plan also caps 5-hour and weekly spend. Payment processing fee extra.
StepFun Step PlanFlash Mini $6.99 / Plus $9.99 / Pro $29 / Max $994 models / variants · 3 calculations
Estimated processed tokens / 30 days
- Flash Mini, Step 3.5 Flash2.37B total (11.09M output)
- Flash Plus, Step 3.5 Flash9.50B total (44.34M output)
- Flash Max, Step 5 Preview60.05B total (280.39M output)
Allowance and practical limit
400M / 1,600M / 8,000M / 40,000M Credits per month, deducted at list price (about 7M Credits per $1). Monthly pool only, no five-hour or weekly window; booster packs available.
StepFun Step Plan (China)Flash Mini ¥49 / Plus ¥99 / Pro ¥199 / Max ¥6992 models / variants · 2 calculations
Estimated processed tokens / 30 days
- Flash Mini, Step 3.5 Flash2.37B total (11.09M output)
- Flash Max, Step 3.5 Flash237.42B total (1.11B output)
Allowance and practical limit
Same Credit pools as the international plan; 1M Credits = ¥1 of list-price usage. USD equivalents use 6.7489 CNY/USD (CFETS, 2026-09-24).
GLM Coding Plan (China)Lite ¥118 / Pro ¥538 / Max ¥1,0782 models / variants · 2 calculations
Estimated processed tokens / 30 days
- Lite, GLM-5.3 off-peak432.12M total (2.02M output)
- Max, GLM-5.3 off-peak6.05B total (28.25M output)
Allowance and practical limit
bigmodel.cn edition of the Z.ai plan: same 10k / 60k / 140k weekly credits and coefficients, priced in CNY (USD at 6.7489, CFETS 2026-09-24). Prices are rendered client-side; figures are the credits-based set in the official site code.
MiniMax Token Plan (China)Plus ¥49 / Max ¥119 / Ultra ¥4692 models / variants · 1 calculation
Estimated processed tokens / 30 days
- All tiers
Allowance and practical limit
Five-hour and weekly windows; text, image and speech share one quota deducted at pay-as-you-go prices. No token allowance published except an approximate Ultra figure.
Alibaba Bailian Coding Plan (China)Pro ¥200 (first month ¥39.90)1 model / variant · 1 calculation
Estimated processed tokens / 30 days
- Pro
Allowance and practical limit
CNY edition of the Alibaba Coding Plan: same request limits and model list. Explicitly excludes automated scripts, backends and noninteractive use; inputs are used for model improvement.
Alibaba Bailian Token Plan (China)Lite ¥39 / Essential ¥79 / Standard ¥139 / Pro ¥499 (limited-time; list ¥60 / ¥120 / ¥180 / ¥600)3 models / variants · 1 calculation
Estimated processed tokens / 30 days
- All tiers
Allowance and practical limit
Credit-based successor to the Coding Plan: monthly Credit caps, no weekly limit since 2026-09-22. Credits-per-token not published. Beijing region only; interactive tools only, no automation; one per verified identity.
Several qualifications matter more than another decimal place in the price:
- Z.ai supports specific coding tools and uses dynamic concurrency limits. Its published examples include automated development tasks, but that does not establish unrestricted custom-backend access. Usage policy
- camelStream’s standard terms permit retention and training on prompts and outputs, without an opt-out. Model choice, queue time and throughput are not guaranteed. It merits a small experiment with suitable non-sensitive work. Terms
- Kimi Code’s personal benefits differ from production API access. The membership page alone does not establish a fixed Code token allowance or identical context limits across products. Code benefits
- NanoGPT’s new subscription terms explicitly exclude building commercial products. They apply to new users now and existing users from September 14. PAYG permits commercial use. Terms, quota accounting
- The DeepSeek reseller needs further clarification. Its pricing cards and DeepSeek integration page give different allowances. Its separate Agent Plan documentation describes an upstream Volcengine product. I have not treated either as an official DeepSeek subscription or assigned it a reliable savings multiplier.
- Chutes’ old unlimited-value anecdotes are obsolete. Its published revision caps monthly subscription benefit at five times its own PAYG value and permits shorter rolling limits.
The capacity CSV includes additional models, peak/off-peak variants and all four token buckets. The catalog snapshot preserves the large public model lists. Model availability, subscription inclusion and a selectable API alias are distinct; unresolved tier or alias details are marked in the comparison.
Normalize the subscription costs#
The same measured mix makes every quantifiable allowance comparable in token units: 1M billable output tokens requires approximately 213.167M input tokens, for 214.167M total processed tokens. The input is divided among the fresh, cache-read and cache-write buckets above before applying each provider’s prices or credit weights. Dividing a subscription’s fee by its advertised credit count would skip that conversion.
output_M = allowance_for_period / deductions_per_1M_output_and_its_input
processed_M = output_M × (1 + measured_input_output_ratio)
cost_per_1M_processed = fee_for_period / processed_M
cost_per_1M_output_with_input = fee_for_period / output_M
These are fully utilized, steady-state 30-day estimates from the dated September 25 price snapshot. Output costs include the associated input and cache processing; they are not the provider’s output-only API rate. Processed-token costs include cache reads, which dominate this workload. Both measures describe the same workload and produce the same cost ordering. Differences in model quality still need an accepted-work benchmark.
OpenCode Go: GoMiMo V2.6 Flash$0.32 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 6.80B
- Output / 30 days
- 31.73M
- $ / 1M processed
- $0.00147
- $ / 1M output + its input
- $0.32
- Saving vs direct API
- 83.3% less
OpenCode Go: GoMiniMax M3$2.63 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 815.16M
- Output / 30 days
- 3.81M
- $ / 1M processed
- $0.01227
- $ / 1M output + its input
- $2.63
- Saving vs direct API
- 83.3% less
OpenCode Go: GoKimi K3$65.94 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 32.48M
- Output / 30 days
- 0.15M
- $ / 1M processed
- $0.30788
- $ / 1M output + its input
- $65.94
- Saving vs direct API
- 33.3% less
Synthetic: One packGLM-5.3-Flash$2.87 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 2.24B
- Output / 30 days
- 10.45M
- $ / 1M processed
- $0.01340
- $ / 1M output + its input
- $2.87
- Saving vs direct API
- 63.1% less
Synthetic: One packKimi-K3$37.85 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 169.75M
- Output / 30 days
- 0.79M
- $ / 1M processed
- $0.17673
- $ / 1M output + its input
- $37.85
- Saving vs direct API
- 61.7% less
MiMo Token Plan: LiteMiMo V2.6 Flash daytime$1.98 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 650.12M
- Output / 30 days
- 3.04M
- $ / 1M processed
- $0.00923
- $ / 1M output + its input
- $1.98
- Saving vs direct API
- 4.5% more expensive
MiMo Token Plan: StandardMiMo V2.6 Flash daytime$1.96 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 1.74B
- Output / 30 days
- 8.14M
- $ / 1M processed
- $0.00917
- $ / 1M output + its input
- $1.96
- Saving vs direct API
- 3.9% more expensive
MiMo Token Plan: ProMiMo V2.6 Flash daytime$1.78 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 6.03B
- Output / 30 days
- 28.13M
- $ / 1M processed
- $0.00830
- $ / 1M output + its input
- $1.78
- Saving vs direct API
- 6.0% less
MiMo Token Plan: MaxMiMo V2.6 Flash daytime$1.65 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 13.00B
- Output / 30 days
- 60.71M
- $ / 1M processed
- $0.00769
- $ / 1M output + its input
- $1.65
- Saving vs direct API
- 12.9% less
Chutes: PlusDeepSeek V4 Flash · conditional ceiling$2.73 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 785.87M
- Output / 30 days
- 3.67M
- $ / 1M processed
- $0.01272
- $ / 1M output + its input
- $2.73
- Saving vs direct API
- 17.2% more expensive
Chutes: ProDeepSeek V4 Flash · conditional ceiling$2.73 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 1.57B
- Output / 30 days
- 7.34M
- $ / 1M processed
- $0.01272
- $ / 1M output + its input
- $2.73
- Saving vs direct API
- 17.2% more expensive
NanoGPT: ProIncluded model at 1x input weight$9.95 / 1M output
- Fee / month
- $12.00
- Processed / 30 days
- 258.35M
- Output / 30 days
- 1.21M
- $ / 1M processed
- $0.04645
- $ / 1M output + its input
- $9.95
- Saving vs direct API
- Not established
NanoGPT: ProIncluded model at 2x input weight$19.90 / 1M output
- Fee / month
- $12.00
- Processed / 30 days
- 129.17M
- Output / 30 days
- 0.60M
- $ / 1M processed
- $0.09290
- $ / 1M output + its input
- $19.90
- Saving vs direct API
- Not established
Kilo Pass: Starter annual billingAnthropic: Claude Sonnet 5$44.43 / 1M output
- Fee / month
- $19.00
- Processed / 30 days
- 91.59M
- Output / 30 days
- 0.43M
- $ / 1M processed
- $0.20744
- $ / 1M output + its input
- $44.43
- Saving vs direct API
- Not established
Kilo Pass: Pro annual billingAnthropic: Claude Sonnet 5$44.43 / 1M output
- Fee / month
- $49.00
- Processed / 30 days
- 236.21M
- Output / 30 days
- 1.10M
- $ / 1M processed
- $0.20744
- $ / 1M output + its input
- $44.43
- Saving vs direct API
- Not established
Kilo Pass: Expert annual billingAnthropic: Claude Sonnet 5$44.43 / 1M output
- Fee / month
- $199.00
- Processed / 30 days
- 959.32M
- Output / 30 days
- 4.48M
- $ / 1M processed
- $0.20744
- $ / 1M output + its input
- $44.43
- Saving vs direct API
- Not established
GitHub Copilot: ProClaude Sonnet 5$44.43 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 48.21M
- Output / 30 days
- 0.23M
- $ / 1M processed
- $0.20744
- $ / 1M output + its input
- $44.43
- Saving vs direct API
- 33.3% less
GitHub Copilot: Pro+Claude Sonnet 5$37.13 / 1M output
- Fee / month
- $39.00
- Processed / 30 days
- 224.97M
- Output / 30 days
- 1.05M
- $ / 1M processed
- $0.17336
- $ / 1M output + its input
- $37.13
- Saving vs direct API
- 44.3% less
GitHub Copilot: MaxClaude Sonnet 5$33.32 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 642.76M
- Output / 30 days
- 3.00M
- $ / 1M processed
- $0.15558
- $ / 1M output + its input
- $33.32
- Saving vs direct API
- 50.0% less
Ollama Cloud Pro / Max: ProGLM-5.3$22.75 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 188.28M
- Output / 30 days
- 0.88M
- $ / 1M processed
- $0.10622
- $ / 1M output + its input
- $22.75
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: MaxGLM-5.3$22.75 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 941.41M
- Output / 30 days
- 4.40M
- $ / 1M processed
- $0.10622
- $ / 1M output + its input
- $22.75
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: ProDeepSeek V4.1 Flash off-peak$0.78 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 5.52B
- Output / 30 days
- 25.80M
- $ / 1M processed
- $0.00362
- $ / 1M output + its input
- $0.78
- Saving vs direct API
- 66.7% less
Command Code: GOATGLM-5.2$9.75 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 219.66M
- Output / 30 days
- 1.03M
- $ / 1M processed
- $0.04552
- $ / 1M output + its input
- $9.75
- Saving vs direct API
- Not established
Command Code: GOATDeepSeek V4 Flash off-peak$0.39 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 5.52B
- Output / 30 days
- 25.80M
- $ / 1M processed
- $0.00181
- $ / 1M output + its input
- $0.39
- Saving vs direct API
- 83.3% less
Command Code: ProClaude Sonnet 5$66.64 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 64.28M
- Output / 30 days
- 0.30M
- $ / 1M processed
- $0.31116
- $ / 1M output + its input
- $66.64
- Saving vs direct API
- 0.0% less
Command Code: Max 10xClaude Opus 5.5$92.12 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 232.47M
- Output / 30 days
- 1.09M
- $ / 1M processed
- $0.43015
- $ / 1M output + its input
- $92.12
- Saving vs direct API
- 0.0% more expensive
StepFun Step Plan: Flash MiniStep 3.5 Flash$0.63 / 1M output
- Fee / month
- $6.99
- Processed / 30 days
- 2.37B
- Output / 30 days
- 11.09M
- $ / 1M processed
- $0.00294
- $ / 1M output + its input
- $0.63
- Saving vs direct API
- 87.8% less
StepFun Step Plan (China): Flash MiniStep 3.5 Flash$0.65 / 1M output
- Fee / month
- $7.26
- Processed / 30 days
- 2.37B
- Output / 30 days
- 11.09M
- $ / 1M processed
- $0.00306
- $ / 1M output + its input
- $0.65
- Saving vs direct API
- Not established
StepFun Step Plan: Flash PlusStep 3.5 Flash$0.23 / 1M output
- Fee / month
- $9.99
- Processed / 30 days
- 9.50B
- Output / 30 days
- 44.34M
- $ / 1M processed
- $0.00105
- $ / 1M output + its input
- $0.23
- Saving vs direct API
- 95.6% less
StepFun Step Plan: Flash MaxStep 5 Preview$0.35 / 1M output
- Fee / month
- $99.00
- Processed / 30 days
- 60.05B
- Output / 30 days
- 280.39M
- $ / 1M processed
- $0.00165
- $ / 1M output + its input
- $0.35
- Saving vs direct API
- 98.3% less
StepFun Step Plan (China): Flash MaxStep 3.5 Flash$0.09 / 1M output
- Fee / month
- $103.57
- Processed / 30 days
- 237.42B
- Output / 30 days
- 1.11B
- $ / 1M processed
- $0.00044
- $ / 1M output + its input
- $0.09
- Saving vs direct API
- Not established
GLM Coding Plan (China): LiteGLM-5.3 (off-peak)$8.66 / 1M output
- Fee / month
- $17.48
- Processed / 30 days
- 432.12M
- Output / 30 days
- 2.02M
- $ / 1M processed
- $0.04045
- $ / 1M output + its input
- $8.66
- Saving vs direct API
- 88.3% less
GLM Coding Plan (China): MaxGLM-5.3 (off-peak)$5.65 / 1M output
- Fee / month
- $159.73
- Processed / 30 days
- 6.05B
- Output / 30 days
- 28.25M
- $ / 1M processed
- $0.02640
- $ / 1M output + its input
- $5.65
- Saving vs direct API
- 92.3% less
Z.ai Lite / Pro / Max: LiteGLM-5.3 (off-peak)$8.92 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 432.12M
- Output / 30 days
- 2.02M
- $ / 1M processed
- $0.04166
- $ / 1M output + its input
- $8.92
- Saving vs direct API
- 86.9% less
Z.ai Lite / Pro / Max: ProGLM-5.3 (off-peak)$6.61 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 2.59B
- Output / 30 days
- 12.11M
- $ / 1M processed
- $0.03086
- $ / 1M output + its input
- $6.61
- Saving vs direct API
- 90.3% less
Z.ai Lite / Pro / Max: MaxGLM-5.3 (off-peak)$5.95 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 6.05B
- Output / 30 days
- 28.25M
- $ / 1M processed
- $0.02777
- $ / 1M output + its input
- $5.95
- Saving vs direct API
- 91.3% less
The comparison covers the quantifiable examples from the plan overview. Other model allocations and time windows use the same calculation:
Show the remaining model and time-of-day calculations
OpenCode Go: GoDeepSeek V4 Flash off-peak$0.78 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 2.76B
- Output / 30 days
- 12.90M
- $ / 1M processed
- $0.00362
- $ / 1M output + its input
- $0.78
- Saving vs direct API
- 66.7% less
OpenCode Go: GoGLM-5.3$45.50 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 47.07M
- Output / 30 days
- 0.22M
- $ / 1M processed
- $0.21245
- $ / 1M output + its input
- $45.50
- Saving vs direct API
- 33.3% less
OpenCode Go: GoGPT 6 Luna$2.22 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 964.14M
- Output / 30 days
- 4.50M
- $ / 1M processed
- $0.01037
- $ / 1M output + its input
- $2.22
- Saving vs direct API
- Not established
OpenCode Go: GoGPT-5.6 Luna$4.58 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 468.02M
- Output / 30 days
- 2.19M
- $ / 1M processed
- $0.02137
- $ / 1M output + its input
- $4.58
- Saving vs direct API
- Not established
Synthetic: One packgpt-oss-120b$1.45 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 4.45B
- Output / 30 days
- 20.76M
- $ / 1M processed
- $0.00675
- $ / 1M output + its input
- $1.45
- Saving vs direct API
- Not established
Synthetic: One packDeepSeek-V4.1-Flash$3.44 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 1.87B
- Output / 30 days
- 8.71M
- $ / 1M processed
- $0.01608
- $ / 1M output + its input
- $3.44
- Saving vs direct API
- 48.1% more expensive
Synthetic: One packQwen3.8-27B$7.01 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 916.11M
- Output / 30 days
- 4.28M
- $ / 1M processed
- $0.03275
- $ / 1M output + its input
- $7.01
- Saving vs direct API
- Not established
Synthetic: One packGLM-4.7-Flash$1.56 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 4.11B
- Output / 30 days
- 19.21M
- $ / 1M processed
- $0.00729
- $ / 1M output + its input
- $1.56
- Saving vs direct API
- Not established
Synthetic: One packNVIDIA-Nemotron-3-Super-120B-A12B-NVFP4$4.54 / 1M output
- Fee / month
- $30.00
- Processed / 30 days
- 1.42B
- Output / 30 days
- 6.61M
- $ / 1M processed
- $0.02120
- $ / 1M output + its input
- $4.54
- Saving vs direct API
- Not established
MiMo Token Plan: LiteMiMo V2.6 Flash off-peak$1.58 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 812.66M
- Output / 30 days
- 3.79M
- $ / 1M processed
- $0.00738
- $ / 1M output + its input
- $1.58
- Saving vs direct API
- 16.4% less
MiMo Token Plan: LiteMiMo V2.6 Pro daytime$4.88 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 263.55M
- Output / 30 days
- 1.23M
- $ / 1M processed
- $0.02277
- $ / 1M output + its input
- $4.88
- Saving vs direct API
- 1.0% more expensive
MiMo Token Plan: LiteMiMo V2.6 Pro off-peak$3.90 / 1M output
- Fee / month
- $6.00
- Processed / 30 days
- 329.44M
- Output / 30 days
- 1.54M
- $ / 1M processed
- $0.01821
- $ / 1M output + its input
- $3.90
- Saving vs direct API
- 19.2% less
MiMo Token Plan: StandardMiMo V2.6 Flash off-peak$1.57 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 2.18B
- Output / 30 days
- 10.18M
- $ / 1M processed
- $0.00734
- $ / 1M output + its input
- $1.57
- Saving vs direct API
- 16.9% less
MiMo Token Plan: StandardMiMo V2.6 Pro daytime$4.85 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 707.10M
- Output / 30 days
- 3.30M
- $ / 1M processed
- $0.02263
- $ / 1M output + its input
- $4.85
- Saving vs direct API
- 0.4% more expensive
MiMo Token Plan: StandardMiMo V2.6 Pro off-peak$3.88 / 1M output
- Fee / month
- $16.00
- Processed / 30 days
- 883.87M
- Output / 30 days
- 4.13M
- $ / 1M processed
- $0.01810
- $ / 1M output + its input
- $3.88
- Saving vs direct API
- 19.7% less
MiMo Token Plan: ProMiMo V2.6 Flash off-peak$1.42 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 7.53B
- Output / 30 days
- 35.17M
- $ / 1M processed
- $0.00664
- $ / 1M output + its input
- $1.42
- Saving vs direct API
- 24.8% less
MiMo Token Plan: ProMiMo V2.6 Pro daytime$4.38 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 2.44B
- Output / 30 days
- 11.41M
- $ / 1M processed
- $0.02047
- $ / 1M output + its input
- $4.38
- Saving vs direct API
- 9.2% less
MiMo Token Plan: ProMiMo V2.6 Pro off-peak$3.51 / 1M output
- Fee / month
- $50.00
- Processed / 30 days
- 3.05B
- Output / 30 days
- 14.26M
- $ / 1M processed
- $0.01638
- $ / 1M output + its input
- $3.51
- Saving vs direct API
- 27.3% less
MiMo Token Plan: MaxMiMo V2.6 Flash off-peak$1.32 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 16.25B
- Output / 30 days
- 75.89M
- $ / 1M processed
- $0.00615
- $ / 1M output + its input
- $1.32
- Saving vs direct API
- 30.3% less
MiMo Token Plan: MaxMiMo V2.6 Pro daytime$4.06 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 5.27B
- Output / 30 days
- 24.61M
- $ / 1M processed
- $0.01897
- $ / 1M output + its input
- $4.06
- Saving vs direct API
- 15.8% less
MiMo Token Plan: MaxMiMo V2.6 Pro off-peak$3.25 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 6.59B
- Output / 30 days
- 30.77M
- $ / 1M processed
- $0.01518
- $ / 1M output + its input
- $3.25
- Saving vs direct API
- 32.6% less
Kilo Pass: Starter annual billingMiniMax: MiniMax M3$10.51 / 1M output
- Fee / month
- $19.00
- Processed / 30 days
- 387.20M
- Output / 30 days
- 1.81M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
Kilo Pass: Pro annual billingMiniMax: MiniMax M3$10.51 / 1M output
- Fee / month
- $49.00
- Processed / 30 days
- 998.57M
- Output / 30 days
- 4.66M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
Kilo Pass: Expert annual billingMiniMax: MiniMax M3$10.51 / 1M output
- Fee / month
- $199.00
- Processed / 30 days
- 4.06B
- Output / 30 days
- 18.94M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
GitHub Copilot: ProClaude Opus 5.5$61.42 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 34.87M
- Output / 30 days
- 0.16M
- $ / 1M processed
- $0.28677
- $ / 1M output + its input
- $61.42
- Saving vs direct API
- 33.3% less
GitHub Copilot: ProGPT-6 Luna$2.22 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 964.14M
- Output / 30 days
- 4.50M
- $ / 1M processed
- $0.01037
- $ / 1M output + its input
- $2.22
- Saving vs direct API
- 33.3% less
GitHub Copilot: Pro+Claude Opus 5.5$51.33 / 1M output
- Fee / month
- $39.00
- Processed / 30 days
- 162.73M
- Output / 30 days
- 0.76M
- $ / 1M processed
- $0.23966
- $ / 1M output + its input
- $51.33
- Saving vs direct API
- 44.3% less
GitHub Copilot: Pro+GPT-6 Luna$1.86 / 1M output
- Fee / month
- $39.00
- Processed / 30 days
- 4.50B
- Output / 30 days
- 21.01M
- $ / 1M processed
- $0.00867
- $ / 1M output + its input
- $1.86
- Saving vs direct API
- 44.3% less
GitHub Copilot: MaxClaude Opus 5.5$46.06 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 464.95M
- Output / 30 days
- 2.17M
- $ / 1M processed
- $0.21508
- $ / 1M output + its input
- $46.06
- Saving vs direct API
- 50.0% less
GitHub Copilot: MaxGPT-6 Luna$1.67 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 12.86B
- Output / 30 days
- 60.02M
- $ / 1M processed
- $0.00778
- $ / 1M output + its input
- $1.67
- Saving vs direct API
- 50.0% less
Ollama Cloud Pro / Max: ProGLM-5.3-Flash$2.59 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 1.65B
- Output / 30 days
- 7.71M
- $ / 1M processed
- $0.01211
- $ / 1M output + its input
- $2.59
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: MaxGLM-5.3-Flash$2.59 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 8.26B
- Output / 30 days
- 38.55M
- $ / 1M processed
- $0.01211
- $ / 1M output + its input
- $2.59
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: ProKimi K3$32.97 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 129.92M
- Output / 30 days
- 0.61M
- $ / 1M processed
- $0.15394
- $ / 1M output + its input
- $32.97
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: MaxKimi K3$32.97 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 649.61M
- Output / 30 days
- 3.03M
- $ / 1M processed
- $0.15394
- $ / 1M output + its input
- $32.97
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: ProMiniMax M3$10.51 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 407.58M
- Output / 30 days
- 1.90M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
Ollama Cloud Pro / Max: MaxMiniMax M3$10.51 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 2.04B
- Output / 30 days
- 9.52M
- $ / 1M processed
- $0.04907
- $ / 1M output + its input
- $10.51
- Saving vs direct API
- 33.3% less
Ollama Cloud Pro / Max: MaxDeepSeek V4.1 Flash off-peak$0.78 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 27.62B
- Output / 30 days
- 128.98M
- $ / 1M processed
- $0.00362
- $ / 1M output + its input
- $0.78
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: ProDeepSeek V4 Pro off-peak$3.80 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 1.13B
- Output / 30 days
- 5.27M
- $ / 1M processed
- $0.01772
- $ / 1M output + its input
- $3.80
- Saving vs direct API
- 66.7% less
Ollama Cloud Pro / Max: MaxDeepSeek V4 Pro off-peak$3.80 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 5.64B
- Output / 30 days
- 26.35M
- $ / 1M processed
- $0.01772
- $ / 1M output + its input
- $3.80
- Saving vs direct API
- 66.7% less
Command Code: GoDeepSeek V4 Flash off-peak$0.23 / 1M output
- Fee / month
- $1.00
- Processed / 30 days
- 920.77M
- Output / 30 days
- 4.30M
- $ / 1M processed
- $0.00109
- $ / 1M output + its input
- $0.23
- Saving vs direct API
- 90.0% less
Command Code: GOATKimi K2.7 Code$8.35 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 256.39M
- Output / 30 days
- 1.20M
- $ / 1M processed
- $0.03900
- $ / 1M output + its input
- $8.35
- Saving vs direct API
- Not established
Command Code: GOATMiMo V2.6 Flash$0.95 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 2.27B
- Output / 30 days
- 10.58M
- $ / 1M processed
- $0.00441
- $ / 1M output + its input
- $0.95
- Saving vs direct API
- 50.0% less
Command Code: GOATKimi K3$49.45 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 43.31M
- Output / 30 days
- 0.20M
- $ / 1M processed
- $0.23091
- $ / 1M output + its input
- $49.45
- Saving vs direct API
- 50.0% less
Command Code: GOATGLM-5.3$34.12 / 1M output
- Fee / month
- $10.00
- Processed / 30 days
- 62.76M
- Output / 30 days
- 0.29M
- $ / 1M processed
- $0.15934
- $ / 1M output + its input
- $34.12
- Saving vs direct API
- 50.0% less
Command Code: ProDeepSeek V4 Flash off-peak$0.66 / 1M output
- Fee / month
- $20.00
- Processed / 30 days
- 6.45B
- Output / 30 days
- 30.10M
- $ / 1M processed
- $0.00310
- $ / 1M output + its input
- $0.66
- Saving vs direct API
- 71.4% less
Command Code: Max 10xDeepSeek V4 Flash off-peak$1.55 / 1M output
- Fee / month
- $100.00
- Processed / 30 days
- 13.81B
- Output / 30 days
- 64.49M
- $ / 1M processed
- $0.00724
- $ / 1M output + its input
- $1.55
- Saving vs direct API
- 33.3% less
Command Code: Max 20xDeepSeek V4 Flash off-peak$1.55 / 1M output
- Fee / month
- $200.00
- Processed / 30 days
- 27.62B
- Output / 30 days
- 128.98M
- $ / 1M processed
- $0.00724
- $ / 1M output + its input
- $1.55
- Saving vs direct API
- 33.3% less
Command Code: Max 20xClaude Opus 5.5$92.12 / 1M output
- Fee / month
- $200.00
- Processed / 30 days
- 464.95M
- Output / 30 days
- 2.17M
- $ / 1M processed
- $0.43015
- $ / 1M output + its input
- $92.12
- Saving vs direct API
- 0.0% more expensive
StepFun Step Plan: Flash MiniStep 5 Preview$2.49 / 1M output
- Fee / month
- $6.99
- Processed / 30 days
- 600.51M
- Output / 30 days
- 2.80M
- $ / 1M processed
- $0.01164
- $ / 1M output + its input
- $2.49
- Saving vs direct API
- 87.8% less
StepFun Step Plan: Flash PlusStep 5 Preview$0.89 / 1M output
- Fee / month
- $9.99
- Processed / 30 days
- 2.40B
- Output / 30 days
- 11.22M
- $ / 1M processed
- $0.00416
- $ / 1M output + its input
- $0.89
- Saving vs direct API
- 95.6% less
StepFun Step Plan (China): Flash PlusStep 3.5 Flash$0.33 / 1M output
- Fee / month
- $14.67
- Processed / 30 days
- 9.50B
- Output / 30 days
- 44.34M
- $ / 1M processed
- $0.00154
- $ / 1M output + its input
- $0.33
- Saving vs direct API
- Not established
StepFun Step Plan: Flash ProStep 3.5 Flash$0.13 / 1M output
- Fee / month
- $29.00
- Processed / 30 days
- 47.48B
- Output / 30 days
- 221.72M
- $ / 1M processed
- $0.00061
- $ / 1M output + its input
- $0.13
- Saving vs direct API
- 97.5% less
StepFun Step Plan: Flash ProStep 5 Preview$0.52 / 1M output
- Fee / month
- $29.00
- Processed / 30 days
- 12.01B
- Output / 30 days
- 56.08M
- $ / 1M processed
- $0.00241
- $ / 1M output + its input
- $0.52
- Saving vs direct API
- 97.5% less
StepFun Step Plan (China): Flash ProStep 3.5 Flash$0.13 / 1M output
- Fee / month
- $29.49
- Processed / 30 days
- 47.48B
- Output / 30 days
- 221.72M
- $ / 1M processed
- $0.00062
- $ / 1M output + its input
- $0.13
- Saving vs direct API
- Not established
StepFun Step Plan: Flash MaxStep 3.5 Flash$0.09 / 1M output
- Fee / month
- $99.00
- Processed / 30 days
- 237.42B
- Output / 30 days
- 1.11B
- $ / 1M processed
- $0.00042
- $ / 1M output + its input
- $0.09
- Saving vs direct API
- 98.3% less
GLM Coding Plan (China): LiteGLM-5.3 (peak)$17.33 / 1M output
- Fee / month
- $17.48
- Processed / 30 days
- 216.06M
- Output / 30 days
- 1.01M
- $ / 1M processed
- $0.08090
- $ / 1M output + its input
- $17.33
- Saving vs direct API
- 76.6% less
GLM Coding Plan (China): ProGLM-5.3 (off-peak)$6.59 / 1M output
- Fee / month
- $79.72
- Processed / 30 days
- 2.59B
- Output / 30 days
- 12.11M
- $ / 1M processed
- $0.03075
- $ / 1M output + its input
- $6.59
- Saving vs direct API
- 91.1% less
GLM Coding Plan (China): ProGLM-5.3 (peak)$13.17 / 1M output
- Fee / month
- $79.72
- Processed / 30 days
- 1.30B
- Output / 30 days
- 6.05M
- $ / 1M processed
- $0.06150
- $ / 1M output + its input
- $13.17
- Saving vs direct API
- 82.2% less
GLM Coding Plan (China): MaxGLM-5.3 (peak)$11.31 / 1M output
- Fee / month
- $159.73
- Processed / 30 days
- 3.02B
- Output / 30 days
- 14.12M
- $ / 1M processed
- $0.05281
- $ / 1M output + its input
- $11.31
- Saving vs direct API
- 84.7% less
Z.ai Lite / Pro / Max: LiteGLM-5.3 (peak)$17.84 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 216.06M
- Output / 30 days
- 1.01M
- $ / 1M processed
- $0.08331
- $ / 1M output + its input
- $17.84
- Saving vs direct API
- 73.9% less
Z.ai Lite / Pro / Max: LiteGLM-5.3-Flash list (off-peak)$2.94 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 1.31B
- Output / 30 days
- 6.11M
- $ / 1M processed
- $0.01375
- $ / 1M output + its input
- $2.94
- Saving vs direct API
- 62.2% less
Z.ai Lite / Pro / Max: LiteGLM-5.3-Flash list (peak)$5.89 / 1M output
- Fee / month
- $18.00
- Processed / 30 days
- 654.52M
- Output / 30 days
- 3.06M
- $ / 1M processed
- $0.02750
- $ / 1M output + its input
- $5.89
- Saving vs direct API
- 24.3% less
Z.ai Lite / Pro / Max: ProGLM-5.3 (peak)$13.22 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 1.30B
- Output / 30 days
- 6.05M
- $ / 1M processed
- $0.06171
- $ / 1M output + its input
- $13.22
- Saving vs direct API
- 80.6% less
Z.ai Lite / Pro / Max: ProGLM-5.3-Flash list (off-peak)$2.18 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 7.85B
- Output / 30 days
- 36.67M
- $ / 1M processed
- $0.01019
- $ / 1M output + its input
- $2.18
- Saving vs direct API
- 72.0% less
Z.ai Lite / Pro / Max: ProGLM-5.3-Flash list (peak)$4.36 / 1M output
- Fee / month
- $80.00
- Processed / 30 days
- 3.93B
- Output / 30 days
- 18.34M
- $ / 1M processed
- $0.02037
- $ / 1M output + its input
- $4.36
- Saving vs direct API
- 43.9% less
Z.ai Lite / Pro / Max: MaxGLM-5.3 (peak)$11.89 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 3.02B
- Output / 30 days
- 14.12M
- $ / 1M processed
- $0.05554
- $ / 1M output + its input
- $11.89
- Saving vs direct API
- 82.6% less
Z.ai Lite / Pro / Max: MaxGLM-5.3-Flash list (off-peak)$1.96 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 18.33B
- Output / 30 days
- 85.57M
- $ / 1M processed
- $0.00917
- $ / 1M output + its input
- $1.96
- Saving vs direct API
- 74.8% less
Z.ai Lite / Pro / Max: MaxGLM-5.3-Flash list (peak)$3.93 / 1M output
- Fee / month
- $168.00
- Processed / 30 days
- 9.16B
- Output / 30 days
- 42.79M
- $ / 1M processed
- $0.01833
- $ / 1M output + its input
- $3.93
- Saving vs direct API
- 49.5% less
“Saving vs direct API” compares the same named model against the direct API row below, using ordinary list prices for GLM Flash and off-peak prices for DeepSeek. Serving configurations and cache behavior can differ. “Not established” means this dataset lacks a defensible same-model baseline; it does not mean zero savings. Chutes estimates remain conditional on its published ceiling and model eligibility. Kilo examples assume annual billing with the full monthly bonus consumed before expiry. Alternative model allocations within a subscription cannot be added together.
The capacity CSV contains all four token buckets, both normalized costs, direct API savings and break-even utilization. It also separates a discount against the serving provider’s own PAYG rates from a discount against the original model API. Missing or conflicting quotas remain blank; camelStream’s unmetered service has no defensible fixed token capacity or normalized cost here. MiniMax, Kimi and the DeepSeek reseller need quota clarification; Alibaba needs tokens per billable request.
At 50% utilization, each subscription’s effective cost per token doubles. Break-even utilization is the fraction of modeled allowance that must replace direct API spending to recover the fee. A figure above 100% means even full utilization is more expensive than that API baseline. Real concurrency, request and reset limits can prevent reaching the modeled allowance.
Start with the API baseline#
The priced workload is 1M billable output tokens plus the measured proportions of fresh input, cache reads and cache writes shown above. Tokenizers differ, so equal token counts do not mean identical amounts of source code. These are estimates at current prices, not the historical cash cost of the measured sessions.
For prices stated per million tokens:
cost = fresh_input_M × input_price
+ cache_read_M × cache_read_price
+ cache_write_M × cache_creation_price
+ output_M × output_price
+ additional_charges
The comparison excludes paid tools, search, additional retries, infrastructure and long-context premiums. Cache creation uses the ordinary input price unless the serving provider specifies a separate full creation rate. Free cache storage or a waived write surcharge does not make the initial input processing free. Kilo’s Sonnet 5 example, for instance, uses its explicit cache-creation rate in the capacity calculator.
MiMo V2.6 Flash$1.89 / 1M output
- Input / M
- $0.14
- Cached / M
- $0.0028
- Output / M
- $0.28
- Cost per 1M output + its input
- $1.89
- $100 processed capacity
- 11.33B total (52.88M output)
DeepSeek V4.1 Flash off-peak$2.33 / 1M output
- Input / M
- $0.15
- Cached / M
- $0.003
- Output / M
- $0.60
- Cost per 1M output + its input
- $2.33
- $100 processed capacity
- 9.21B total (42.99M output)
GPT-6 Luna$3.33 / 1M output
- Input / M
- $0.1
- Cached / M
- $0.01
- Output / M
- $0.50
- Cost per 1M output + its input
- $3.33
- $100 processed capacity
- 6.43B total (30.01M output)
MiMo V2.6 Pro$4.83 / 1M output
- Input / M
- $0.435
- Cached / M
- $0.0036
- Output / M
- $0.87
- Cost per 1M output + its input
- $4.83
- $100 processed capacity
- 4.44B total (20.72M output)
Step 3.5 Flash$5.15 / 1M output
- Input / M
- $0.1
- Cached / M
- $0.02
- Output / M
- $0.30
- Cost per 1M output + its input
- $5.15
- $100 processed capacity
- 4.15B total (19.40M output)
GLM-5.3-Flash list$7.78 / 1M output
- Input / M
- $0.15
- Cached / M
- $0.03
- Output / M
- $0.50
- Cost per 1M output + its input
- $7.78
- $100 processed capacity
- 2.75B total (12.85M output)
DeepSeek V4 Pro off-peak$11.39 / 1M output
- Input / M
- $0.66
- Cached / M
- $0.022
- Output / M
- $1.98
- Cost per 1M output + its input
- $11.39
- $100 processed capacity
- 1.88B total (8.78M output)
MiniMax M3$15.76 / 1M output
- Input / M
- $0.3
- Cached / M
- $0.06
- Output / M
- $1.20
- Cost per 1M output + its input
- $15.76
- $100 processed capacity
- 1.36B total (6.34M output)
Step 5 Preview$20.38 / 1M output
- Input / M
- $1
- Cached / M
- $0.05
- Output / M
- $2.70
- Cost per 1M output + its input
- $20.38
- $100 processed capacity
- 1.05B total (4.91M output)
Claude Sonnet 5$66.64 / 1M output
- Input / M
- $2
- Cached / M
- $0.2
- Output / M
- $10.00
- Cost per 1M output + its input
- $66.64
- $100 processed capacity
- 321.38M total (1.50M output)
GLM-5.3$68.25 / 1M output
- Input / M
- $1.4
- Cached / M
- $0.26
- Output / M
- $4.40
- Cost per 1M output + its input
- $68.25
- $100 processed capacity
- 313.80M total (1.47M output)
GLM-5.3 (bigmodel.cn)$73.89 / 1M output
- Input / M
- $1.18538
- Cached / M
- $0.29634
- Output / M
- $4.15
- Cost per 1M output + its input
- $73.89
- $100 processed capacity
- 289.85M total (1.35M output)
Claude Opus 5.5$92.12 / 1M output
- Input / M
- $4
- Cached / M
- $0.2
- Output / M
- $20.00
- Cost per 1M output + its input
- $92.12
- $100 processed capacity
- 232.47M total (1.09M output)
Kimi K3$98.91 / 1M output
- Input / M
- $3
- Cached / M
- $0.3
- Output / M
- $15.00
- Cost per 1M output + its input
- $98.91
- $100 processed capacity
- 216.54M total (1.01M output)
Without cache hits, the same workload costs $30.12 on MiMo V2.6 Flash, $32.58 on off-peak DeepSeek V4.1 Flash, and $654.50 on Kimi K3. A plan that counts cached input at full token weight can lose much of its advantage against an API with inexpensive cache reads. Repeated context also inflates token totals without representing newly generated work.
DeepSeek’s listed prices apply outside weekday peak windows of 01:00–04:00 and 06:00–10:00 UTC, and all day on Chinese public holidays; peak rates double. DeepSeek has retired its V4 Flash model names and now serves them with V4.1 Flash at the V4.1 price. MiniMax M3’s listed rate applies through 512K input per request. These timing and context conditions belong in the comparison, not in a footnote to the invoice. DeepSeek pricing, MiniMax pricing
Compare with OpenRouter too#
OpenRouter narrows some subscription discounts because its alternative providers can undercut the model author’s API. The result depends heavily on cached-input pricing. The following snapshot was captured on September 25, 2026, at 13:49 UTC from OpenRouter’s public model-endpoint API. Each quote keeps one provider’s input, cache and output rates together.
The comparison is grouped by subscription, then model, then tier or time window. All twenty-seven researched coding subscriptions appear, including those with unknown costs. Models from the saved catalogs remain visible in the index; large unpriced catalogs expand underneath their subscription. Catalog presence, subscription eligibility and model pinning are separate questions, so each group retains those conditions. Native workflow subscriptions are discussed separately below.
The unit-cost fields report cash cost per 1M output tokens plus 213.167M associated input tokens, using the measured cache mix. Subscription summaries show the monthly fee, and capacity details show total processed tokens with output in parentheses. Each calculated scenario spends the full allowance on its named model; scenarios within one subscription are alternative allocations. “Not established” means the snapshot lacks a defensible entitlement or calculation; “Not quoted” means no matching OpenRouter route was priced.
OpenRouter figures include its standard card-purchase fee: 5.5%, with a $0.80 minimum. The reference purchase is $100 of inference credits plus $5.50 in fees, so quoted inference cost is multiplied by 1.055 once. A $10 credit purchase instead incurs an 8% effective fee. Taxes and special billing arrangements are excluded. OpenRouter fee policy
Jump to subscription: OpenCode Go · Z.ai Lite / Pro / Max · Synthetic · camelStream · MiniMax Token Plan · Kimi membership / Code · Alibaba Cloud Coding Plan · deepseekv4pro.com · MiMo Token Plan · Chutes · NanoGPT · Kilo Pass · ChatGPT (Codex) · Claude (Claude Code) · Cursor · GitHub Copilot · Google AI Pro / Ultra · Devin Pro / Max · SuperGrok · Ollama Cloud Pro / Max · Command Code · StepFun Step Plan · StepFun Step Plan (China) · GLM Coding Plan (China) · MiniMax Token Plan (China) · Alibaba Bailian Coding Plan (China) · Alibaba Bailian Token Plan (China)
OpenCode Go33 recorded models / aliases$10 / month
One shared allowance. Regional and experimental conditions apply; uncalculated models retain unknown capacity.
MiMo V2.6 Flash$0.32 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 6.80B (31.73M output)
- Subscription $ / 1M output + input
- $0.32
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 84.2% less
MiniMax M3$2.63 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 815.16M (3.81M output)
- Subscription $ / 1M output + input
- $2.63
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- 80.3% less
Kimi K3$65.94 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 32.48M (0.15M output)
- Subscription $ / 1M output + input
- $65.94
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- 9.7% less
DeepSeek V4 Flash$0.78 subscription
- Tier / conditions
- Go; off-peak
- Processed / 30 days (output)
- 2.76B (12.90M output)
- Subscription $ / 1M output + input
- $0.78
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 36.8% less
GLM-5.3$45.50 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 47.07M (0.22M output)
- Subscription $ / 1M output + input
- $45.50
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 66.8% more expensive
GPT-5.6 Luna$4.58 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 468.02M (2.19M output)
- Subscription $ / 1M output + input
- $4.58
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT 6 Luna$2.22 subscription
- Tier / conditions
- Go
- Processed / 30 days (output)
- 964.14M (4.50M output)
- Subscription $ / 1M output + input
- $2.22
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiMo V2.6 Pro$5.09 OpenRouter
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- Not established
MiMo V2.5Pricing not established
- Tier / conditions
- Go; deprecated 2026-10-21
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiMo V2.5 ProPricing not established
- Tier / conditions
- Go; deprecated 2026-10-21
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GLM-5.3-Flash$2.67 OpenRouter
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- Not established
DeepSeek V4 Pro$6.25 OpenRouter
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $6.25Baidu, FP8
- Subscription saving
- Not established
DeepSeek V4.1 Flash$1.23 OpenRouter
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- Not established
Grok 4.6Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Grok 4.7Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GLM-5.2Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GLM-5.1Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Kimi K2.7 CodePricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Kimi K2.6Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax M2.7Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax M2.5Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.8 MaxPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.8 FlashPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.7 MaxPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.7 PlusPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen 3.6 PlusPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
LongCat 2.0Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Muse Spark 1.3 ContributorPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Muse Spark 1.2 ContributorPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
DeepSeek V4 Flash Vision ExpPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Hy4 previewPricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Hy3Pricing not established
- Tier / conditions
- Go
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Space BunnyPricing not established
- Tier / conditions
- Go; free for a limited time
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Z.ai Lite / Pro / Max2 recorded models / aliases$18 / $80 / $168 / month
Older aliases route to these models. Each row allocates the entire tier allowance to one model and time window.
GLM-5.36 scenarios
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 432.12M (2.02M output)
- Subscription $ / 1M output + input
- $8.92
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 67.3% less
- Tier / conditions
- Lite; peak
- Processed / 30 days (output)
- 216.06M (1.01M output)
- Subscription $ / 1M output + input
- $17.84
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 34.6% less
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 2.59B (12.11M output)
- Subscription $ / 1M output + input
- $6.61
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 75.8% less
- Tier / conditions
- Pro; peak
- Processed / 30 days (output)
- 1.30B (6.05M output)
- Subscription $ / 1M output + input
- $13.22
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 51.5% less
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 6.05B (28.25M output)
- Subscription $ / 1M output + input
- $5.95
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 78.2% less
- Tier / conditions
- Max; peak
- Processed / 30 days (output)
- 3.02B (14.12M output)
- Subscription $ / 1M output + input
- $11.89
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 56.4% less
GLM-5.3-Flash6 scenarios
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 1.31B (6.11M output)
- Subscription $ / 1M output + input
- $2.94
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 10.3% more expensive
- Tier / conditions
- Lite; peak
- Processed / 30 days (output)
- 654.52M (3.06M output)
- Subscription $ / 1M output + input
- $5.89
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 120.6% more expensive
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 7.85B (36.67M output)
- Subscription $ / 1M output + input
- $2.18
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 18.3% less
- Tier / conditions
- Pro; peak
- Processed / 30 days (output)
- 3.93B (18.34M output)
- Subscription $ / 1M output + input
- $4.36
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 63.4% more expensive
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 18.33B (85.57M output)
- Subscription $ / 1M output + input
- $1.96
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 26.5% less
- Tier / conditions
- Max; peak
- Processed / 30 days (output)
- 9.16B (42.79M output)
- Subscription $ / 1M output + input
- $3.93
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 47.1% more expensive
Synthetic11 recorded models / aliases$30 per pack / month
The saved served catalog includes four automatic aliases. Catalog availability alone does not establish subscription eligibility.
syn:large:textPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
syn:small:textPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
syn:large:visionPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
syn:small:visionPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
openai/gpt-oss-120b$1.45 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 4.45B (20.76M output)
- Subscription $ / 1M output + input
- $1.45
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
zai-org/GLM-5.3-Flash$2.87 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 2.24B (10.45M output)
- Subscription $ / 1M output + input
- $2.87
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 7.5% more expensive
deepseek-ai/DeepSeek-V4.1-Flash$3.44 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 1.87B (8.71M output)
- Subscription $ / 1M output + input
- $3.44
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 180.7% more expensive
moonshotai/Kimi-K3$37.85 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 169.75M (0.79M output)
- Subscription $ / 1M output + input
- $37.85
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- 48.2% less
Qwen/Qwen3.8-27B$7.01 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 916.11M (4.28M output)
- Subscription $ / 1M output + input
- $7.01
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
zai-org/GLM-4.7-Flash$1.56 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 4.11B (19.21M output)
- Subscription $ / 1M output + input
- $1.56
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4$4.54 subscription
- Tier / conditions
- One pack
- Processed / 30 days (output)
- 1.42B (6.61M output)
- Subscription $ / 1M output + input
- $4.54
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
camelStream4 recorded models / aliases$5 per stream / month
Unmetered service has no fixed token entitlement. Named routes are possible automatic choices, not selectable allocations.
DeepSeek V4.1 Flash$1.23 OpenRouter
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- Not established
GLM-5.3-Flash$2.67 OpenRouter
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- Not established
GPT-5.6 LunaPricing not established
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Muse Spark 1.3Pricing not established
- Tier / conditions
- One stream; automatic routing, no model pinning
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax Token Plan2 recorded models / aliases$22 / $55 / $132 / month
Text models only. Image and speech share quota but are outside this token comparison; published numerical allowance is insufficient.
MiniMax M3$13.30 OpenRouter
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- Not established
MiniMax M2.7Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Kimi membership / Code3 recorded models / aliasesPlus $19 / Pro $39 / Max $99 / Ultra $199 / month
No fixed comparable entitlement established. Membership K3 and Code aliases are distinguished; alias costs are not assumed to equal K3.
Kimi K3$73.04 OpenRouter
- Tier / conditions
- Plus and above; 1M context on Pro and above
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- Not established
kimi-for-codingPricing not established
- Tier / conditions
- Code alias; currently K2.8 Preview
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
kimi-for-coding-highspeedPricing not established
- Tier / conditions
- Code HighSpeed alias; currently K2.7 Code; Pro and above
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Alibaba Cloud Coding Plan10 recorded models / aliases$50 / month
Request quotas cannot be converted without tokens per billable request. Automated scripts and application backends are excluded.
qwen3.7-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3.6-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
kimi-k2.5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
glm-5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax-M2.5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3.5-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3-max-2026-01-23Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3-coder-nextPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
qwen3-coder-plusPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
glm-4.7Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
deepseekv4pro.com5 recorded models / aliasesCoding $19.90 / Coding Max $49.90 / Agent $49.90 / Coding Pro Max $199 / month
Independent reseller; published quotas conflict. Models added by the Agent plan are marked separately.
DeepSeek V4.1 Flash$1.23 OpenRouter
- Tier / conditions
- Coding tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- Not established
DeepSeek V4 Pro$6.25 OpenRouter
- Tier / conditions
- Agent plan (as DeepSeek V4)
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $6.25Baidu, FP8
- Subscription saving
- Not established
GLM-5.3$27.28 OpenRouter
- Tier / conditions
- Agent plan only
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- Not established
MiniMax M3$13.30 OpenRouter
- Tier / conditions
- Agent plan only
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- Not established
Kimi K3$73.04 OpenRouter
- Tier / conditions
- Agent plan only
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- Not established
MiMo Token Plan2 recorded models / aliases$6 / $16 / $50 / $100 / month
Text models only. Credit pools are shared across models; daytime and off-peak rows are alternative allocations.
MiMo V2.6 Flash8 scenarios
- Tier / conditions
- Lite; daytime
- Processed / 30 days (output)
- 650.12M (3.04M output)
- Subscription $ / 1M output + input
- $1.98
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 0.9% less
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 812.66M (3.79M output)
- Subscription $ / 1M output + input
- $1.58
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 20.7% less
- Tier / conditions
- Standard; daytime
- Processed / 30 days (output)
- 1.74B (8.14M output)
- Subscription $ / 1M output + input
- $1.96
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 1.5% less
- Tier / conditions
- Standard; off-peak
- Processed / 30 days (output)
- 2.18B (10.18M output)
- Subscription $ / 1M output + input
- $1.57
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 21.2% less
- Tier / conditions
- Pro; daytime
- Processed / 30 days (output)
- 6.03B (28.13M output)
- Subscription $ / 1M output + input
- $1.78
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 10.9% less
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 7.53B (35.17M output)
- Subscription $ / 1M output + input
- $1.42
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 28.7% less
- Tier / conditions
- Max; daytime
- Processed / 30 days (output)
- 13.00B (60.71M output)
- Subscription $ / 1M output + input
- $1.65
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 17.4% less
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 16.25B (75.89M output)
- Subscription $ / 1M output + input
- $1.32
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 33.9% less
MiMo V2.6 Pro8 scenarios
- Tier / conditions
- Lite; daytime
- Processed / 30 days (output)
- 263.55M (1.23M output)
- Subscription $ / 1M output + input
- $4.88
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 4.2% less
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 329.44M (1.54M output)
- Subscription $ / 1M output + input
- $3.90
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 23.4% less
- Tier / conditions
- Standard; daytime
- Processed / 30 days (output)
- 707.10M (3.30M output)
- Subscription $ / 1M output + input
- $4.85
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 4.8% less
- Tier / conditions
- Standard; off-peak
- Processed / 30 days (output)
- 883.87M (4.13M output)
- Subscription $ / 1M output + input
- $3.88
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 23.9% less
- Tier / conditions
- Pro; daytime
- Processed / 30 days (output)
- 2.44B (11.41M output)
- Subscription $ / 1M output + input
- $4.38
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 13.9% less
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 3.05B (14.26M output)
- Subscription $ / 1M output + input
- $3.51
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 31.1% less
- Tier / conditions
- Max; daytime
- Processed / 30 days (output)
- 5.27B (24.61M output)
- Subscription $ / 1M output + input
- $4.06
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 20.2% less
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 6.59B (30.77M output)
- Subscription $ / 1M output + input
- $3.25
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- 36.2% less
Chutes14 recorded models / aliases$10 / $20 / month
Served catalog; exact subscription tier eligibility needs account confirmation. Flash capacity remains a conditional ceiling.
google/gemma-4-31B-turbo-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3.6-27B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3.8-27B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3.5-397B-A17B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
zai-org/GLM-5.1-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
deepseek-ai/DeepSeek-V3.2-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
zai-org/GLM-5.2-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
moonshotai/Kimi-K2.6-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
deepseek-ai/DeepSeek-V4-Flash-0731-TEE2 scenarios
- Tier / conditions
- Pinned V4 Flash 0731 checkpoint; compared with the V4.1 Flash routes DeepSeek now serves at the same list price; conditional ceiling
- Processed / 30 days (output)
- 785.87M (3.67M output)
- Subscription $ / 1M output + input
- $2.73
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 122.1% more expensive
- Tier / conditions
- Pinned V4 Flash 0731 checkpoint; compared with the V4.1 Flash routes DeepSeek now serves at the same list price; conditional ceiling
- Processed / 30 days (output)
- 1.57B (7.34M output)
- Subscription $ / 1M output + input
- $2.73
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 122.1% more expensive
unsloth/Mistral-Nemo-Instruct-2407-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
moonshotai/Kimi-K3-TEE$73.04 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- Not established
Qwen/Qwen3-235B-A22B-Thinking-2507-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Qwen/Qwen3-32B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Nemotron-3-Nano-Omni-30B-TEEPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
NanoGPT296 recorded models / aliases$12 / month
Subscription-included model/variant catalog. Per-model multipliers were not captured, so the 1x/2x quota examples cannot be assigned to individual IDs.
Unassigned quota examples
- Included model at 1x input weight$9.95 / 1M output + input
- Included model at 2x input weight$19.90 / 1M output + input
Per-model weights remain unknown.
GLM 5.3 Flash$2.67 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- Not established
MiniMax M3$13.30 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- Not established
MiMo V2.6 Flash$1.99 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- Not established
MiMo V2.6 Pro$5.09 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- Not established
GLM 5.3$27.28 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- Not established
DeepSeek V4 Pro 0813$6.25 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $6.25Baidu, FP8
- Subscription saving
- Not established
DeepSeek V4.1 Flash$1.23 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- Not established
Show 289 additional NanoGPT catalog models / variants
These IDs have no matched allowance or OpenRouter quote in this snapshot. Catalog coverage: Subscription-included endpoint.
- Ternary Bonsai 2 27B
- Clover 1 150B Preview
- Gemma 4 31B Split-Untied
- Muse Spark 1.3 Contributor
- Mercury 2.5 Preview
- Granite 4.2 8B
- DiffusionGemma
- Qwen 3.8 27B Cybersecurity
- Gemma 4 26B A4B Cybersecurity
- GLM 5.3 Flash Cybersecurity
- Nemotron 3.5 Content Safety
- TheDrummer/Artemis v1.1
- GLM 5.3 Flash Uncensored
- Qwen 3.8 27B Uncensored
- Qwen 3.8 27B Queen
- Qwen 3.8 27B Hemingway
- Gemma 4 12B Semancer
- Gemma 4 12B StationKeeper
- Qwen 3.8 27B Fable
- Qwen 3.8 27B Uncensored Thinking
- Qwen 3.8 27B Obliterated
- Qwen 3.8 27B Obliterated Thinking
- Gemma 4 31B MeroMero v2
- Gemma 4 31B MeroMero v2 Thinking
- Gemma 4 26B A4B MeroMero
- Gemma 4 26B A4B MeroMero Thinking
- Gemma 4 26B A4B Musica
- Gemma 4 26B A4B Shadow Siren
- Gemma 4 26B A4B Chimera X
- Gemma 4 26B A4B Luminous Mirror
- Gemma 4 26B A4B Dark Soul
- Gemma 4 26B A4B Moonlight Dusk
- Gemma 4 26B A4B Opus Distill
- Gemma 4 31B Fabled
- Gemma 4 31B DarkIdol
- Gemma 4 31B Garnet
- Gemma 4 31B Novelist
- Gemma 4 31B Isometry
- Gemma 4 31B Gembrain
- Gemma 4 31B Gemsicle
- Ornith 1.5 35B
- Ornith 1.5 35B Thinking
- Qwen3.8 27B
- Qwen3.8 27B Thinking
- Gemma 4 12B Instruct
- Laguna S 2.1
- Laguna S 2.1 Thinking
- Nvidia Nemotron 3.5 Lightning
- Nvidia Nemotron 3.5 Lightning Thinking
- Nvidia Nemotron 3 Ultra 550B
- Nvidia Nemotron 3 Ultra 550B Thinking
- LFM2.5 2.6B
- Muse Glimmer 30B
- Muse Spark 1.2 Contributor
- LongCat 2.0
- LongCat 2.0 Thinking
- Step 3.7 Flash Thinking
- Doubao Seed Character
- NanoGPT Help
- Auto model
- Auto model (Basic)
- Auto model (Standard)
- Auto model (Premium)
- Claw High
- Claw Medium
- Claw Low
- Hermes High
- Hermes Medium
- Hermes Low
- GPT OSS 120B
- GPT OSS 20B
- Amoral Gemma3 27B v2
- Mistral Devstral Small 2505
- Veiled Calla 12B
- Qwen: QvQ Max
- Step 3.5 Flash 2603
- Step 3.5 Flash
- Nex N2.5 Mini
- Nex N2.5 Pro
- Qwen 3 Coder 480B
- Llama 4 Maverick
- Llama 4 Scout
- DeepSeek R1 0528
- Kimi K2 Thinking
- Kimi K2.5
- Kimi K2.5 Thinking
- Kimi K2.6
- Kimi K2.7 Code
- Kimi K2.6 Thinking
- Ministral 3 14B
- Mistral Small 4 119B
- Mistral Small 4 119B Thinking
- Devstral 2 123B
- Hermes 4 Large (Thinking)
- OpenReasoning Nemotron 32B
- DeepSeek R1
- DeepSeek V3/Deepseek Chat
- MiniMax M2.5
- Qwen 3 235b A22B
- Qwen3.5 9B
- Qwen 3 32b
- Qwen 3 14b
- Qwen3 30B A3B
- Qwen3 Coder 30B A3B Instruct
- Qwen 3 235b A22B 2507
- Qwen 3 235b A22B 2507 Thinking
- Qwen3 Next 80B A3B (Instruct)
- Qwen3 Next 80B A3B (Thinking)
- MiniMax M2
- MiniMax M3 Thinking
- MiniMax M2.7
- MiniMax Latest
- MiMo V2.5
- MiMo V2.5 Thinking
- MiMo V2.5 Pro
- MiMo V2.5 Pro Thinking
- MiniMax M2.1
- GLM 4.6
- GLM 4.6 Thinking
- GLM 5
- GLM 5 Thinking
- GLM 5.1
- GLM Latest
- GLM 5.1 Thinking
- GLM 5.2
- GLM 5.2 Thinking
- GLM 5.3 Thinking
- GLM 4.7 Flash
- GLM 4.7 Flash Thinking
- GLM 4.7
- GLM 4.7 Thinking
- GLM 4.6V
- Qwen3 30B A3B Instruct 2507
- Llama 3.3 70b Instruct
- Nvidia Nemotron 70b
- Sao10K Stheno 8b
- Grayline Qwen3 8B
- Hermes 4 Large
- Hermes 3 70B
- Qwen3.8 Flash
- Qwen3.5 122B A10B
- Qwen3.5 122B A10B Thinking
- Qwen3.5 27B
- Qwen3.5 27B Thinking
- Qwen3.5 35B A3B
- Qwen3.5 35B A3B Thinking
- Qwen3.6 35B A3B
- Qwen3.6 35B A3B Thinking
- Qwen3.6 27B
- Qwen3.6 27B Thinking
- DeepSeek V3.2 Exp
- DeepSeek V3.2 Exp Thinking
- DeepSeek V3.2
- DeepSeek V3.2 Thinking
- DeepSeek V4 Flash
- DeepSeek V4 Flash Vision Exp
- DeepSeek V4 Flash 0731
- DeepSeek V4 Flash Latest
- DeepSeek V4 Flash 0731 (Thinking)
- DeepSeek V4 Flash (Thinking)
- DeepSeek V4 Pro 0813 Thinking
- DeepSeek V4 Pro
- DeepSeek Latest
- DeepSeek V4 Pro (Thinking)
- Qwen3.5 397B A17B
- Qwen3.5 397B A17B Thinking
- DeepSeek V4.1 Flash Thinking
- Qwen 2.5 Coder 32b
- Phi 4 Multimodal
- Phi 4 Mini
- The Drummer Cydonia 24B v2
- The Drummer Cydonia 24B v4
- The Drummer Cydonia 24B v4.1
- The Drummer Cydonia 24B v4.3
- The Drummer Magidonia 24B v4.3
- MS3.2 24B Magnum Diamond
- Omega Directive 24B Unslop v2.0
- EVA Llama 3.33 70B
- Steelskull Nevoria 70b
- Steelskull Nevoria R1 70b
- Steelskull Electra R1 70b
- Qwen2 72B Dracarys
- Lumimaid v0.2
- DeepSeek V3/Chat Cheaper
- Llama 3.3 70B Instruct abliterated
- MythoMax 13B
- Qwen2.5 72B
- EVA-Qwen2.5-32B-v0.2
- TheDrummer Skyfall 36B V2
- Qwen 3 8B
- K2-Think
- DeepSeek V3.1
- DeepSeek V3.1 Thinking
- DeepSeek V3.1 Terminus
- DeepSeek V3.1 Terminus (Thinking)
- DeepSeek Chat 0324
- GLM 4.5 (Thinking)
- GLM 4.5
- GLM 4.5 Air
- GLM 4.5 Air (Thinking)
- MN-LooseCannon-12B-v1
- EVA-Qwen2.5-72B-v0.2
- EVA-LLaMA-3.33-70B-v0.1
- Llama 3.1 8b Instruct
- ReMM SLERP 13B
- Mistral Saba
- Neural Daredevil 8B abliterated
- Llama 3 70B abliterated
- Magnum V2 72B
- Mistral Nemo
- DeepSeek Reasoner
- Llama 3.05 Storybreaker Ministral 70b
- Nemotron Tenyxchat Storybreaker 70b
- Mag Mell R1
- Qwerky 72B
- Anubis 70B v1
- Anubis 70B v1.1
- Llama 3.2 3b Instruct
- Llama 3.1 8B (decentralized)
- Llama 3.1 70B Hanami
- Rocinante 12b
- Llama 3.3 70B Euryale
- Llama 3.1 70B Euryale
- Llama 3.3 70B Cu Mai
- UnslopNemo 12b v4
- NemoMix 12B Unleashed
- Mistral Nemo Starcannon 12b v1
- Llama 3.1 70B Celeste v0.1
- DeepSeek R1 Qwen Abliterated
- DeepSeek R1 Llama 70B Abliterated
- Qwen 2.5 32B Abliterated
- Deepseek R1 Cheaper
- Llama 3.3 70B Wayfarer
- Gemma 3 27B IT
- Gemma 3 12B IT
- Gemma 3 4B IT
- Qwen25 VL 72b
- Holo3-35B-A3B
- Holo3-35B-A3B Thinking
- Cogito v1 Preview Qwen 32B
- Llama-xLAM-2 70B fc-r
- Mistral Small 3.1 24B (2503)
- Mistral Small 3.2 24B (2506)
- Nvidia Nemotron Super 49B
- Shisa V2 Llama 3.3 70B
- Shisa V2.1 Llama 3.3 70B
- GLM 4 9B 0414
- GLM 4 32B 0414
- Qwen3.5 27B Blossom V6.4 Derestricted
- Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted
- Qwen3.5 27B Claude 4.6 Opus Reasoning Distilled Derestricted Lite
- Gemma 4 31B Agares v1
- Gemma 4 31B Animus V14.1
- Gemma 4 31B AssGuard
- Gemma 4 31B Dark Gemistry
- Gemma 4 31B Gembrain Uncensored Heretic
- Gemma 4 31B Gembrain X Core
- Gemma 4 31B Isometry RP
- Gemma 4 31B Novelist (ArliAI)
- Gemma 4 31B SDFT Heretic RP
- Gemma 4 31B StyleTune
- Qwen3.5 27B BlueStar v3 Derestricted
- Qwen3.5 27B Queen Derestricted
- Gemma 4 31B Claude 4.6 Opus Reasoning Distilled
- Gemma 4 31B Cognitive Unshackled
- Gemma 4 31B DarkIdol (ArliAI)
- Gemma 4 31B Fabled (ArliAI)
- Gemma 4 31B Garnet V2
- Gemma 4 31B K1 v5
- Gemma 4 31B MeroMero
- Gemma 4 31B Queen
- GLM 4.6 Derestricted v5
- Venice Uncensored
- Gemma 4 26B A4B
- Gemma 4 26B A4B Thinking
- Tencent Hy3
- Qwen3 Coder Next
- Ling 3.0 Flash VL
- Ling 3.0 Flash
- Ling 3.0 Flash Thinking
- Gemma 4 31B
- Gemma 4 31B Thinking
- Nvidia Nemotron 3 Nano 30B
- Nvidia Nemotron 3 Super 120B
- Nvidia Nemotron 3 Super 120B Thinking
- Manta Mini 1.0
- Mistral Code Agent Latest
- Synth 2.5 Flash Preview
- Synth 2.5 Pro Preview
Kilo Pass394 recorded models / aliases$19 / $49 / $199 / month
Gateway catalog includes free and paid variants. Annual credit bonus is modeled only where provider rates were captured; catalog presence alone does not establish eligibility.
DeepSeek: DeepSeek V4.1 Flash$1.23 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- Not established
Z.ai: GLM 5.3 Flash$2.67 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- Not established
MoonshotAI: Kimi K3$73.04 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- Not established
MiniMax: MiniMax M33 scenarios
- Tier / conditions
- Starter annual billing
- Processed / 30 days (output)
- 387.20M (1.81M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- 21.0% less
- Tier / conditions
- Pro annual billing
- Processed / 30 days (output)
- 998.57M (4.66M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- 21.0% less
- Tier / conditions
- Expert annual billing
- Processed / 30 days (output)
- 4.06B (18.94M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- 21.0% less
Xiaomi: MiMo-V2.6-Flash$1.99 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- Not established
Xiaomi: MiMo-V2.6-Pro$5.09 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $5.09Xiaomi, FP8
- Subscription saving
- Not established
Z.ai: GLM 5.3$27.28 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- Not established
DeepSeek: DeepSeek V4 Pro 0813$6.25 OpenRouter
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $6.25Baidu, FP8
- Subscription saving
- Not established
Anthropic: Claude Sonnet 53 scenarios
- Tier / conditions
- Starter annual billing
- Processed / 30 days (output)
- 91.59M (0.43M output)
- Subscription $ / 1M output + input
- $44.43
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 36.8% less
- Tier / conditions
- Pro annual billing
- Processed / 30 days (output)
- 236.21M (1.10M output)
- Subscription $ / 1M output + input
- $44.43
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 36.8% less
- Tier / conditions
- Expert annual billing
- Processed / 30 days (output)
- 959.32M (4.48M output)
- Subscription $ / 1M output + input
- $44.43
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 36.8% less
Show 385 additional Kilo Pass catalog models / variants
These IDs have no matched allowance or OpenRouter quote in this snapshot. Catalog coverage: Served catalog; consult plan terms for eligibility.
- Auto Efficient
- Auto Free
- Space Bunny Alpha (new)
- Poolside: Laguna S 2.1 (free)
- NVIDIA: Nemotron 3 Ultra (free)
- Dots Studio: Dots3-Note Preview (free)
- Anthropic: Claude Opus 5.5 (new)
- OpenAI: GPT-6 Sol (new)
- Fireworks: Ember-1
- Z.ai: GLM 5.3 Prime
- Qwen: Qwen3.8 Max Prime
- AionLabs: Aion 3.5 Mini
- AionLabs: Aion 3.5
- Upstage: Solar Mini 4
- Cohere: Command A+
- OpenAI: GPT-6 Luna Pro
- OpenAI: GPT-6 Luna
- OpenAI: GPT-6 Sol Pro
- Xiaomi: MiMo-V2.6-Pro-UltraSpeed
- SpaceXAI: Grok 4.7
- Qwen: Qwen3.8 Omni Flash
- PrismML: Ternary Bonsai 2 27B
- Z.ai: GLM 5.3 FlashX
- Pareto
- DeepSeek: DeepSeek Pro Latest
- DeepSeek: DeepSeek Flash Latest
- Inference.net: Schematron V2 Turbo
- Inference.net: Schematron V2 Small
- OpenAI: GPT Astra Latest ($$$$)
- OpenAI: GPT Sol Latest
- OpenAI: GPT Terra Latest
- OpenAI: GPT Luna Latest
- Sakana: Fugu Ultra v2
- Sakana: Fugu Max
- inclusionAI: Ling 3.0 Flash VL
- Inception: Mercury 2.5
- OpenAI: GPT-6 Astra ($$$$)
- OpenAI: GPT-6 Astra Pro ($$$$)
- inclusionAI: Ling 3.0 Flash Sante (free)
- Qwen: Qwen3.8 Max (0902)
- Meta: Muse Spark 1.3 Contributor
- Meta: Muse Spark 1.3
- Google: Gemini 3.8 Flash (50% off)
- Anthropic: Claude Fable 5.1 ($$$$)
- IBM: Granite 4.2 8B
- Tencent: Hy4 preview
- inclusionAI: Ling 3.0 Flash Fin
- inclusionAI: Ling 3.0 Flash Fin (free)
- Z.ai: GLM Flash Latest
- Qwen: Qwen3.8 Flash
- Meta: Muse Spark 1.2 Contributor
- DeepSeek: DeepSeek V4 Flash Vision Exp
- Tencent: Hy-MT2-1.8B
- Tencent: Hy-MT2-30B-A3B
- Z.ai: GLM Latest
- Tencent: Hy-MT2-7B
- Qwen: Qwen3.8 27B
- Qwen: Qwen3.8 27B (free)
- Google: Gemini 3.7 Flash
- ByteDance Seed: Seed 2.1 Turbo
- Qwen: Qwen3.8 2.4T A95B
- ByteDance Seed: Seed-2.0-Code
- SpaceXAI: Grok 4.6
- LiquidAI: LFM2.5-2.6B (free)
- NVIDIA: Nemotron 3.5 Lightning
- NVIDIA: Nemotron 3.5 Lightning (free)
- Sakana: Sakana Namazu
- Upstage: Solar Pro 4
- Meta: Muse Glimmer 30B
- Meta: Muse Spark 1.2
- DeepSeek: DeepSeek V4 Flash Latest
- DeepSeek: DeepSeek V4 Flash 0731
- Thinking Machines: Inkling Small
- Thinking Machines: Inkling Small (free)
- Qwen: Qwen3.7 Flash
- Anthropic: Claude Opus 5
- inclusionAI: Ling 3.0 Flash
- Poolside: Laguna S 2.1
- Google: Gemini 3.6 Flash
- Google: Gemini 3.5 Flash Lite
- Meituan: LongCat 2.0
- Thinking Machines: Inkling
- OpenRouter Auto Router (Beta)
- Meta: Muse Spark 1.1
- Kwaipilot: KAT-Coder-Pro V2.5
- OpenAI: GPT-5.6 Luna Pro
- OpenAI: GPT-5.6 Luna
- OpenAI: GPT-5.6 Terra Pro
- OpenAI: GPT-5.6 Terra
- OpenAI: GPT-5.6 Sol Pro
- OpenAI: GPT-5.6 Sol
- SpaceXAI: Grok 4.5
- xAI: Grok Latest
- AionLabs: Aion-3.0-Mini
- AionLabs: Aion-3.0
- Tencent: Hy3
- Poolside: Laguna XS 2.1
- Poolside: Laguna XS 2.1 (free)
- Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
- Sakana: Fugu Ultra
- Google: Nano Banana 2 (Gemini 3.1 Flash Image)
- Google: Nano Banana Pro (Gemini 3 Pro Image)
- Cohere: North Mini Code (free)
- Z.ai: GLM 5.2
- Z.ai: GLM 5.2 (free)
- OpenRouter: Fusion
- MoonshotAI: Kimi K2.7 Code
- Anthropic: Claude Fable Latest ($$$$)
- Anthropic: Claude Fable 5 ($$$$)
- NVIDIA: Nemotron 3.5 Content Safety
- NVIDIA: Nemotron 3.5 Content Safety (free)
- NVIDIA: Nemotron 3 Ultra
- Qwen: Qwen3.7 Plus
- StepFun: Step 3.7 Flash
- Anthropic: Claude Opus 4.8
- Qwen: Qwen3.7 Max
- SpaceXAI: Grok Build 0.1
- Google: Gemini 3.5 Flash
- Perceptron: Perceptron Mk1
- Google: Gemini 3.1 Flash Lite
- OpenAI: GPT Chat Latest
- SpaceXAI: Grok 4.3
- Mistral: Mistral Medium 3.5
- NVIDIA: Nemotron 3 Nano Omni (free)
- Anthropic: Claude Haiku Latest
- OpenAI: GPT Mini Latest
- Google: Gemini Pro Latest
- MoonshotAI: Kimi Latest
- Google: Gemini Flash Latest
- Anthropic: Claude Sonnet Latest
- Qwen: Qwen3.5 Plus 2026-04-20
- Qwen: Qwen3.6 Flash
- Qwen: Qwen3.6 35B A3B
- Qwen: Qwen3.6 Max Preview (retires Oct 9)
- Qwen: Qwen3.6 27B
- OpenAI: GPT-5.5 Pro ($$$$)
- OpenAI: GPT-5.5
- DeepSeek: DeepSeek V4 Pro 0423
- DeepSeek: DeepSeek V4 Flash 0423
- Tencent: Hy3 preview
- Xiaomi: MiMo-V2.5-Pro
- Xiaomi: MiMo-V2.5
- OpenAI: GPT-5.4 Image 2
- Anthropic: Claude Opus Latest
- OpenRouter Pareto Code Router
- MoonshotAI: Kimi K2.6
- Anthropic: Claude Opus 4.7
- Z.ai: GLM 5.1
- Google: Gemma 4 26B A4B
- Google: Gemma 4 31B
- Qwen: Qwen3.6 Plus
- Z.ai: GLM 5V Turbo
- Arcee AI: Trinity Large Thinking
- SpaceXAI: Grok 4.20 Multi-Agent
- SpaceXAI: Grok 4.20
- Google: Lyria 3 Pro Preview
- Google: Lyria 3 Clip Preview
- Reka Edge
- MiniMax: MiniMax M2.7
- OpenAI: GPT-5.4 Nano
- OpenAI: GPT-5.4 Mini
- Mistral: Mistral Small 4
- Z.ai: GLM 5 Turbo
- NVIDIA: Nemotron 3 Super
- NVIDIA: Nemotron 3 Super (free)
- ByteDance Seed: Seed-2.0-Lite
- Qwen: Qwen3.5-9B
- OpenAI: GPT-5.4 Pro ($$$$)
- OpenAI: GPT-5.4
- Inception: Mercury 2
- Google: Gemini 3.1 Flash Lite Preview
- ByteDance Seed: Seed-2.0-Mini
- Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)
- Qwen: Qwen3.5-35B-A3B
- Qwen: Qwen3.5-27B
- Qwen: Qwen3.5-122B-A10B
- Qwen: Qwen3.5-Flash
- Google: Gemini 3.1 Pro Preview Custom Tools
- OpenAI: GPT-5.3-Codex
- AionLabs: Aion-2.0
- Google: Gemini 3.1 Pro Preview
- Anthropic: Claude Sonnet 4.6
- Qwen: Qwen3.5 Plus 2026-02-15
- Qwen: Qwen3.5 397B A17B
- MiniMax: MiniMax M2.5
- Z.ai: GLM 5
- Qwen: Qwen3 Max Thinking (retires Oct 9)
- Anthropic: Claude Opus 4.6
- Qwen: Qwen3 Coder Next
- OpenRouter Free Models Router
- StepFun: Step 3.5 Flash
- MoonshotAI: Kimi K2.5
- Upstage: Solar Pro 3
- MiniMax: MiniMax M2-her
- Writer: Palmyra X5
- OpenAI: GPT Audio
- OpenAI: GPT Audio Mini
- Z.ai: GLM 4.7 Flash
- OpenAI: GPT-5.2-Codex
- ByteDance Seed: Seed 1.6 Flash
- ByteDance Seed: Seed 1.6
- MiniMax: MiniMax M2.1 (retires Oct 8)
- Z.ai: GLM 4.7
- Google: Gemini 3 Flash Preview
- NVIDIA: Nemotron 3 Nano 30B A3B
- OpenAI: GPT-5.2 Chat
- OpenAI: GPT-5.2 Pro ($$$$)
- OpenAI: GPT-5.2
- Mistral: Devstral 2 2512
- Relace: Relace Search
- Z.ai: GLM 4.6V
- OpenRouter Body Builder (beta)
- OpenAI: GPT-5.1-Codex-Max
- Amazon: Nova 2 Lite
- Mistral: Ministral 3 14B 2512
- Mistral: Ministral 3 8B 2512
- Mistral: Ministral 3 3B 2512
- Mistral: Mistral Large 3 2512
- DeepSeek: DeepSeek V3.2 (retires Sep 28)
- Anthropic: Claude Opus 4.5
- Google: Nano Banana Pro (Gemini 3 Pro Image Preview)
- OpenAI: GPT-5.1
- OpenAI: GPT-5.1-Codex
- OpenAI: GPT-5.1-Codex-Mini
- MoonshotAI: Kimi K2 Thinking
- Amazon: Nova Premier 1.0
- Perplexity: Sonar Pro Search
- Mistral: Voxtral Small 24B 2507
- OpenAI: gpt-oss-safeguard-20b
- MiniMax: MiniMax M2
- Qwen: Qwen3 VL 32B Instruct (retires Oct 9)
- IBM: Granite 4.0 Micro
- OpenAI: GPT-5 Image Mini
- Anthropic: Claude Haiku 4.5
- Qwen: Qwen3 VL 8B Thinking (retires Oct 9)
- Qwen: Qwen3 VL 8B Instruct (retires Oct 9)
- OpenAI: GPT-5 Image ($$$$)
- Google: Nano Banana (Gemini 2.5 Flash Image)
- Qwen: Qwen3 VL 30B A3B Thinking (retires Oct 9)
- Qwen: Qwen3 VL 30B A3B Instruct
- OpenAI: GPT-5 Pro ($$$$)
- Z.ai: GLM 4.6
- Anthropic: Claude Sonnet 4.5
- DeepSeek: DeepSeek V3.2 Exp (retires Sep 28)
- TheDrummer: Cydonia 24B V4.1
- Relace: Relace Apply 3
- Qwen: Qwen3 VL 235B A22B Thinking (retires Oct 9)
- Qwen: Qwen3 VL 235B A22B Instruct
- Qwen: Qwen3 Max (retires Oct 9)
- Qwen: Qwen3 Coder Plus (retires Oct 9)
- DeepSeek: DeepSeek V3.1 Terminus (retires Sep 28)
- Qwen: Qwen3 Coder Flash
- Qwen: Qwen3 Next 80B A3B Thinking
- Qwen: Qwen3 Next 80B A3B Instruct
- Qwen: Qwen Plus 0728 (retires Oct 9)
- MoonshotAI: Kimi K2 0905
- Qwen: Qwen3 30B A3B Thinking 2507 (retires Oct 9)
- Nous: Hermes 4 405B
- DeepSeek: DeepSeek V3.1
- Mistral: Mistral Medium 3.1
- Z.ai: GLM 4.5V
- OpenAI: GPT-5
- OpenAI: GPT-5 Mini
- OpenAI: GPT-5 Nano
- OpenAI: gpt-oss-120b
- OpenAI: gpt-oss-20b
- Anthropic: Claude Opus 4.1 ($$$$)
- Mistral: Codestral 2508
- Qwen: Qwen3 Coder 30B A3B Instruct
- Qwen: Qwen3 30B A3B Instruct 2507
- Z.ai: GLM 4.5
- Z.ai: GLM 4.5 Air
- Qwen: Qwen3 235B A22B Thinking 2507 (retires Oct 9)
- Qwen: Qwen3 Coder 480B A35B
- ByteDance: UI-TARS 7B
- Google: Gemini 2.5 Flash Lite (retires Oct 20)
- Qwen: Qwen3 235B A22B Instruct 2507
- MoonshotAI: Kimi K2 0711
- Venice: Uncensored
- Tencent: Hunyuan A13B Instruct
- Morph: Morph V3 Large
- Morph: Morph V3 Fast
- Baidu: ERNIE 4.5 VL 424B A47B (retires Oct 8)
- Mistral: Mistral Small 3.2 24B
- MiniMax: MiniMax M1
- Google: Gemini 2.5 Flash (retires Oct 20)
- Google: Gemini 2.5 Pro (retires Oct 20)
- OpenAI: o3 Pro ($$$$)
- Google: Gemini 2.5 Pro Preview 06-05
- DeepSeek: R1 0528
- Anthropic: Claude Sonnet 4
- Mistral: Mistral Medium 3
- Meta: Llama Guard 4 12B
- Qwen: Qwen3 30B A3B
- Qwen: Qwen3 8B (retires Oct 9)
- Qwen: Qwen3 14B
- Qwen: Qwen3 32B
- Qwen: Qwen3 235B A22B (retires Oct 9)
- OpenAI: o4 Mini High
- OpenAI: o3
- OpenAI: o4 Mini
- OpenAI: GPT-4.1
- OpenAI: GPT-4.1 Mini
- OpenAI: GPT-4.1 Nano
- Meta: Llama 4 Maverick
- Meta: Llama 4 Scout
- DeepSeek: DeepSeek V3 0324
- OpenAI: o1-pro ($$$$)
- Mistral: Mistral Small 3.1 24B
- Google: Gemma 3 4B
- Google: Gemma 3 12B
- Cohere: Command A
- Reka Flash 3
- Google: Gemma 3 27B
- TheDrummer: Skyfall 36B V2
- Perplexity: Sonar Reasoning Pro
- Perplexity: Sonar Pro
- Perplexity: Sonar Deep Research
- Mistral: Saba
- OpenAI: o3 Mini High
- AionLabs: Aion-RP 1.0 (8B)
- Qwen: Qwen2.5 VL 72B Instruct
- Qwen: Qwen-Plus
- OpenAI: o3 Mini
- Mistral: Mistral Small 3
- Perplexity: Sonar
- DeepSeek: R1 Distill Llama 70B (retires Sep 28)
- DeepSeek: R1
- MiniMax: MiniMax-01
- Microsoft: Phi 4
- DeepSeek: DeepSeek V3
- Sao10K: Llama 3.3 Euryale 70B
- OpenAI: o1 ($$$$)
- Cohere: Command R7B (12-2024)
- Meta: Llama 3.3 70B Instruct
- Amazon: Nova Lite 1.0
- Amazon: Nova Micro 1.0
- Amazon: Nova Pro 1.0
- OpenAI: GPT-4o (2024-11-20)
- Mistral Large 2407
- Qwen2.5 Coder 32B Instruct
- TheDrummer: UnslopNemo 12B
- Magnum v4 72B
- Qwen: Qwen2.5 7B Instruct
- Meta: Llama 3.2 1B Instruct
- Meta: Llama 3.2 3B Instruct
- Qwen2.5 72B Instruct
- Cohere: Command R (08-2024)
- Cohere: Command R+ (08-2024)
- Sao10K: Llama 3.1 Euryale 70B v2.2
- Nous: Hermes 3 70B Instruct
- Nous: Hermes 3 405B Instruct
- Sao10K: Llama 3 8B Lunaris
- OpenAI: GPT-4o (2024-08-06)
- Meta: Llama 3.1 70B Instruct
- Meta: Llama 3.1 8B Instruct
- Mistral: Mistral Nemo
- OpenAI: GPT-4o-mini
- OpenAI: GPT-4o-mini (2024-07-18)
- Google: Gemma 2 27B
- OpenAI: GPT-4o
- OpenAI: GPT-4o (2024-05-13)
- Mistral: Mixtral 8x22B Instruct
- WizardLM-2 8x22B
- OpenAI: GPT-4 Turbo ($$$$)
- Anthropic: Claude 3 Haiku
- Mistral Large
- OpenAI: GPT-3.5 Turbo (older v0613)
- OpenRouter Auto Router
- OpenAI: GPT-3.5 Turbo Instruct
- OpenAI: GPT-3.5 Turbo 16k
- Mancer: Weaver (alpha)
- ReMM SLERP 13B
- MythoMax 13B
- OpenAI: GPT-3.5 Turbo
- OpenAI: GPT-4 ($$$$)
- Stealth: Qwen3.6 Plus (50% off)
- Stealth: Claude Opus 4.8 (20% off)
- Stealth: Claude Opus 4.7 (20% off)
- Stealth: Claude Sonnet 4.6 (20% off)
- Stealth: Claude Opus 4.6 (20% off)
- StepFun: Step 3.7 Flash (free)
- Auto Frontier
- Auto Balanced
- Auto Small
ChatGPT (Codex)9 recorded models / aliasesGo $8 / Plus $20 / Pro from $100 / month
Message estimates per five hours are published per model and tier; they are ranges, not a convertible token allowance. The Pro 20x price is not stated on the Codex pricing page and chatgpt.com/pricing blocks automated retrieval.
GPT-6 AstraPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-6 SolPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-6 LunaPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.6 SolPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.6 TerraPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.6 LunaPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.4Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.4 miniPricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Claude (Claude Code)3 recorded models / aliasesPro $20 / Max 5x $100 / Max 20x $200 / month
Tier multiples are published; the base allowance is not, so no token capacity is calculated.
Claude Fable 5.1Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Claude Opus 5.5Pricing not established
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Claude Sonnet 5$70.31 OpenRouter
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- Not established
Cursor13 recorded models / aliasesPro $20 / Pro+ $60 / Ultra $200 / month
Per-model rates are published, but the included amount per plan is not, so no token capacity is calculated.
Grok 4.7Pricing not established
- Tier / conditions
- Cursor Models pool
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Grok 4.6Pricing not established
- Tier / conditions
- Cursor Models pool
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Grok 4.5Pricing not established
- Tier / conditions
- Cursor Models pool
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Composer 2.5Pricing not established
- Tier / conditions
- Cursor Models pool
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Claude Fable 5.1Pricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Claude Opus 5.5Pricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Claude Sonnet 5$70.31 OpenRouter
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- Not established
Gemini 3.1 ProPricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Gemini 3.8 FlashPricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.6 SolPricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.6 TerraPricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-5.6 LunaPricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Muse Spark 1.3Pricing not established
- Tier / conditions
- Other Models pool, at API price
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GitHub Copilot3 recorded models / aliasesPro $10 / Pro+ $39 / Max $100 / month
Only three models are calculated; every listed model draws on the same credit pool at its own rates. Paid plans get 10% off model costs under auto model selection, which is excluded here.
Claude Sonnet 53 scenarios
- Tier / conditions
- Pro
- Processed / 30 days (output)
- 48.21M (0.23M output)
- Subscription $ / 1M output + input
- $44.43
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 36.8% less
- Tier / conditions
- Pro+
- Processed / 30 days (output)
- 224.97M (1.05M output)
- Subscription $ / 1M output + input
- $37.13
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 47.2% less
- Tier / conditions
- Max
- Processed / 30 days (output)
- 642.76M (3.00M output)
- Subscription $ / 1M output + input
- $33.32
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 52.6% less
Claude Opus 5.53 scenarios
- Tier / conditions
- Pro
- Processed / 30 days (output)
- 34.87M (0.16M output)
- Subscription $ / 1M output + input
- $61.42
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Pro+
- Processed / 30 days (output)
- 162.73M (0.76M output)
- Subscription $ / 1M output + input
- $51.33
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Max
- Processed / 30 days (output)
- 464.95M (2.17M output)
- Subscription $ / 1M output + input
- $46.06
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GPT-6 Luna3 scenarios
- Tier / conditions
- Pro
- Processed / 30 days (output)
- 964.14M (4.50M output)
- Subscription $ / 1M output + input
- $2.22
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Pro+
- Processed / 30 days (output)
- 4.50B (21.01M output)
- Subscription $ / 1M output + input
- $1.86
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Max
- Processed / 30 days (output)
- 12.86B (60.02M output)
- Subscription $ / 1M output + input
- $1.67
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Google AI Pro / Ultra3 recorded models / aliasesPro $19.99 / Ultra $99.99 (5x) or $199.99 (20x) / month
Request, task and rate-limit allowances only. No tokens-per-request figure is published, so no token capacity is calculated.
Gemini CLI (Gemini model family)3 scenarios
- Tier / conditions
- Requests per user per day
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Requests per user per day
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Requests per user per day
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Antigravity agent modelsPricing not established
- Tier / conditions
- Pro: more generous rate limits; Ultra 5x / 20x: higher and highest rate limits
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
JulesPricing not established
- Tier / conditions
- Pro 100 / Ultra 300 tasks per rolling 24 hours; 15 / 60 concurrent
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Devin Pro / Max2 recorded models / aliasesPro $20 / Max $200 / month
Quota amounts are not published; the on-demand credit rates are also not published as per-token prices.
SWE-22 scenarios
- Tier / conditions
- Free in Devin Desktop and CLI through 2026-10-10
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Free in Devin Desktop and CLI through 2026-10-10
- Processed / 30 days (output)
- Unmetered; throughput unknown
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Frontier and open-source models2 scenarios
- Tier / conditions
- Draws on the plan quota, then on-demand credits
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Draws on the plan quota, then on-demand credits
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
SuperGrok2 recorded models / aliasesSuperGrok $30 / Plus $100; Lite and Heavy prices not shown on the pricing page / month
No numerical usage limits are published for any tier.
Grok 4.6Pricing not established
- Tier / conditions
- Paid tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Grok BuildPricing not established
- Tier / conditions
- All plans; usage scales by tier
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Ollama Cloud Pro / Max6 recorded models / aliasesPro $20 ($200/yr) / Max $100 / month
Dollar credits at Ollama's own rates; MiniMax M3 is priced at twice MiniMax's list rate.
GLM-5.32 scenarios
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 188.28M (0.88M output)
- Subscription $ / 1M output + input
- $22.75
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 16.6% less
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 941.41M (4.40M output)
- Subscription $ / 1M output + input
- $22.75
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 16.6% less
GLM-5.3-Flash2 scenarios
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 1.65B (7.71M output)
- Subscription $ / 1M output + input
- $2.59
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 2.8% less
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 8.26B (38.55M output)
- Subscription $ / 1M output + input
- $2.59
- OpenRouter $ / 1M output + input
- $2.67InferenceNet, FP4
- Subscription saving
- 2.8% less
Kimi K32 scenarios
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 129.92M (0.61M output)
- Subscription $ / 1M output + input
- $32.97
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- 54.9% less
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 649.61M (3.03M output)
- Subscription $ / 1M output + input
- $32.97
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- 54.9% less
MiniMax M32 scenarios
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 407.58M (1.90M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- 21.0% less
- Tier / conditions
- Pro and Max
- Processed / 30 days (output)
- 2.04B (9.52M output)
- Subscription $ / 1M output + input
- $10.51
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- 21.0% less
DeepSeek V4.1 Flash2 scenarios
- Tier / conditions
- Off-peak rate; off-peak
- Processed / 30 days (output)
- 5.52B (25.80M output)
- Subscription $ / 1M output + input
- $0.78
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 36.8% less
- Tier / conditions
- Off-peak rate; off-peak
- Processed / 30 days (output)
- 27.62B (128.98M output)
- Subscription $ / 1M output + input
- $0.78
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 36.8% less
DeepSeek V4 Pro2 scenarios
- Tier / conditions
- Off-peak rate; off-peak
- Processed / 30 days (output)
- 1.13B (5.27M output)
- Subscription $ / 1M output + input
- $3.80
- OpenRouter $ / 1M output + input
- $6.25Baidu, FP8
- Subscription saving
- 39.3% less
- Tier / conditions
- Off-peak rate; off-peak
- Processed / 30 days (output)
- 5.64B (26.35M output)
- Subscription $ / 1M output + input
- $3.80
- OpenRouter $ / 1M output + input
- $6.25Baidu, FP8
- Subscription saving
- 39.3% less
Command Code8 recorded models / aliasesGo $1 / GOAT $10 / Pro $20 / Max 10x $100 / Max 20x $200 / month
Allowances are alternatives within one pool. Temporary deals (MiniMax M3 2x, Grok 4.7 40% off to Sep 27, DeepSeek V4.1 Flash boost to Sep 28, MiMo V2.5 up to 99% off) are excluded.
DeepSeek V4 Flash5 scenarios
- Tier / conditions
- Off-peak rates; peak doubles; off-peak
- Processed / 30 days (output)
- 920.77M (4.30M output)
- Subscription $ / 1M output + input
- $0.23
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 81.0% less
- Tier / conditions
- Off-peak rates; peak doubles; off-peak
- Processed / 30 days (output)
- 5.52B (25.80M output)
- Subscription $ / 1M output + input
- $0.39
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 68.4% less
- Tier / conditions
- Off-peak rates; peak doubles; off-peak
- Processed / 30 days (output)
- 6.45B (30.10M output)
- Subscription $ / 1M output + input
- $0.66
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 45.8% less
- Tier / conditions
- Off-peak rates; peak doubles; off-peak
- Processed / 30 days (output)
- 13.81B (64.49M output)
- Subscription $ / 1M output + input
- $1.55
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 26.4% more expensive
- Tier / conditions
- Off-peak rates; peak doubles; off-peak
- Processed / 30 days (output)
- 27.62B (128.98M output)
- Subscription $ / 1M output + input
- $1.55
- OpenRouter $ / 1M output + input
- $1.23Morph
- Subscription saving
- 26.4% more expensive
GLM-5.2$9.75 subscription
- Tier / conditions
- GOAT
- Processed / 30 days (output)
- 219.66M (1.03M output)
- Subscription $ / 1M output + input
- $9.75
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Kimi K2.7 Code$8.35 subscription
- Tier / conditions
- GOAT
- Processed / 30 days (output)
- 256.39M (1.20M output)
- Subscription $ / 1M output + input
- $8.35
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiMo V2.6 Flash$0.95 subscription
- Tier / conditions
- GOAT
- Processed / 30 days (output)
- 2.27B (10.58M output)
- Subscription $ / 1M output + input
- $0.95
- OpenRouter $ / 1M output + input
- $1.99Xiaomi, FP8
- Subscription saving
- 52.6% less
Kimi K3$49.45 subscription
- Tier / conditions
- GOAT
- Processed / 30 days (output)
- 43.31M (0.20M output)
- Subscription $ / 1M output + input
- $49.45
- OpenRouter $ / 1M output + input
- $73.04Relace, FP4
- Subscription saving
- 32.3% less
GLM-5.3$34.12 subscription
- Tier / conditions
- GOAT
- Processed / 30 days (output)
- 62.76M (0.29M output)
- Subscription $ / 1M output + input
- $34.12
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 25.1% more expensive
Claude Sonnet 5$66.64 subscription
- Tier / conditions
- Pro and above; $20 per premium model on Pro
- Processed / 30 days (output)
- 64.28M (0.30M output)
- Subscription $ / 1M output + input
- $66.64
- OpenRouter $ / 1M output + input
- $70.31Anthropic
- Subscription saving
- 5.2% less
Claude Opus 5.52 scenarios
- Tier / conditions
- Max plans only; premium-model limit
- Processed / 30 days (output)
- 232.47M (1.09M output)
- Subscription $ / 1M output + input
- $92.12
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Max plans only; premium-model limit
- Processed / 30 days (output)
- 464.95M (2.17M output)
- Subscription $ / 1M output + input
- $92.12
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
StepFun Step Plan4 recorded models / aliasesFlash Mini $6.99 / Plus $9.99 / Pro $29 / Max $99 / month
The provider states the USD conversion as approximate. Credits reset monthly and do not roll over.
Step 3.5 Flash4 scenarios
- Tier / conditions
- Flash Mini
- Processed / 30 days (output)
- 2.37B (11.09M output)
- Subscription $ / 1M output + input
- $0.63
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Plus
- Processed / 30 days (output)
- 9.50B (44.34M output)
- Subscription $ / 1M output + input
- $0.23
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Pro
- Processed / 30 days (output)
- 47.48B (221.72M output)
- Subscription $ / 1M output + input
- $0.13
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Max
- Processed / 30 days (output)
- 237.42B (1.11B output)
- Subscription $ / 1M output + input
- $0.09
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Step 5 Preview4 scenarios
- Tier / conditions
- Flash Mini
- Processed / 30 days (output)
- 600.51M (2.80M output)
- Subscription $ / 1M output + input
- $2.49
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Plus
- Processed / 30 days (output)
- 2.40B (11.22M output)
- Subscription $ / 1M output + input
- $0.89
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Pro
- Processed / 30 days (output)
- 12.01B (56.08M output)
- Subscription $ / 1M output + input
- $0.52
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Max
- Processed / 30 days (output)
- 60.05B (280.39M output)
- Subscription $ / 1M output + input
- $0.35
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Step 3.7 FlashPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
step-router-v1Pricing not established
- Tier / conditions
- Automatic routing
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
StepFun Step Plan (China)2 recorded models / aliasesFlash Mini ¥49 / Plus ¥99 / Pro ¥199 / Max ¥699 / month
China-region edition; CNY prices converted for comparison only.
Step 3.5 Flash4 scenarios
- Tier / conditions
- Flash Mini
- Processed / 30 days (output)
- 2.37B (11.09M output)
- Subscription $ / 1M output + input
- $0.65
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Plus
- Processed / 30 days (output)
- 9.50B (44.34M output)
- Subscription $ / 1M output + input
- $0.33
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Pro
- Processed / 30 days (output)
- 47.48B (221.72M output)
- Subscription $ / 1M output + input
- $0.13
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
- Tier / conditions
- Flash Max
- Processed / 30 days (output)
- 237.42B (1.11B output)
- Subscription $ / 1M output + input
- $0.09
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Step 5 PreviewPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
GLM Coding Plan (China)2 recorded models / aliasesLite ¥118 / Pro ¥538 / Max ¥1,078 / month
Sep 25 - Oct 7 all-day off-peak promotion excluded.
GLM-5.36 scenarios
- Tier / conditions
- Lite; off-peak
- Processed / 30 days (output)
- 432.12M (2.02M output)
- Subscription $ / 1M output + input
- $8.66
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 68.2% less
- Tier / conditions
- Lite; peak
- Processed / 30 days (output)
- 216.06M (1.01M output)
- Subscription $ / 1M output + input
- $17.33
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 36.5% less
- Tier / conditions
- Pro; off-peak
- Processed / 30 days (output)
- 2.59B (12.11M output)
- Subscription $ / 1M output + input
- $6.59
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 75.9% less
- Tier / conditions
- Pro; peak
- Processed / 30 days (output)
- 1.30B (6.05M output)
- Subscription $ / 1M output + input
- $13.17
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 51.7% less
- Tier / conditions
- Max; off-peak
- Processed / 30 days (output)
- 6.05B (28.25M output)
- Subscription $ / 1M output + input
- $5.65
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 79.3% less
- Tier / conditions
- Max; peak
- Processed / 30 days (output)
- 3.02B (14.12M output)
- Subscription $ / 1M output + input
- $11.31
- OpenRouter $ / 1M output + input
- $27.28Morph
- Subscription saving
- 58.5% less
GLM-5.3-FlashPricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
MiniMax Token Plan (China)2 recorded models / aliasesPlus ¥49 / Max ¥119 / Ultra ¥469 / month
China-region edition.
MiniMax M3$13.30 OpenRouter
- Tier / conditions
- All listed tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- $13.30GMICloud, FP8
- Subscription saving
- Not established
MiniMax M2.7Pricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Alibaba Bailian Coding Plan (China)1 recorded model / aliasPro ¥200 (first month ¥39.90) / month
Lite closed to new purchase 2026-03-20; Pro sold in limited daily slots.
qwen3.7-plusPricing not established
- Tier / conditions
- Pro
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
Alibaba Bailian Token Plan (China)3 recorded models / aliasesLite ¥39 / Essential ¥79 / Standard ¥139 / Pro ¥499 (limited-time; list ¥60 / ¥120 / ¥180 / ¥600) / month
Team edition seats: ¥150 / ¥550 / ¥1,398 per seat (25k / 100k / 250k Credits).
qwen3.8-maxPricing not established
- Tier / conditions
- All personal tiers
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
deepseek-v4.1-flashPricing not established
- Tier / conditions
- Half Credits 22:00-08:00 (limited time)
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
glm-5.3Pricing not established
- Tier / conditions
- See plan terms
- Processed / 30 days (output)
- Not established
- Subscription $ / 1M output + input
- Not established
- OpenRouter $ / 1M output + input
- Not quoted
- Subscription saving
- Not established
The selected route is the lowest modeled cost in the captured endpoints with status 0, advertised tool support and explicit cache-read pricing; when two eligible routes cost the same, the model author’s route is kept. Nonzero-status routes are excluded from selection. On September 25 every author route was eligible, and for MiMo and Sonnet it was also the cheapest. These are dated quotes, not a promise of availability or default routing. The snapshot records provider tags, context limits, quantization, rate overrides and source URLs in the price dataset.
Headline input prices can pick the wrong provider. For DeepSeek V4 Pro, Ionstream quotes $0.2528/M fresh input, below Baidu’s $0.3498/M. But its cache-read price is $0.08832/M versus Baidu’s $0.01113/M. Including their output rates and the purchase fee, this measured workload costs $23.21 through Ionstream versus $6.25 through Baidu. The quote, provider tag and context limit for that example are retained in the dataset. DeepSeek V4 Pro provider quotes
Reproducing the measured 96.53% input cache-hit rate is an assumption. OpenRouter uses sticky routing to help preserve caches, but fallback can move requests to a different provider. Verify actual billed cache reads in the coding harness. Pinning a provider or quantization changes the available routes; an FP4 quote is not evidence of equal application quality to an FP8 or author-hosted route. Each request must also fit that endpoint’s context and pricing tier. Caching, provider routing
Z.ai’s GLM-5.3-Flash author quote is back at its list price; the half-price promotion ended September 9. DeepSeek comparisons use OpenRouter’s V4.1 Flash and Pro 0813 entries; Chutes still serves the pinned V4 Flash 0731 checkpoint, and a subscription’s generic alias may not guarantee a specific checkpoint. The author-route DeepSeek quotes are off-peak. Several selected routes are FP4 or do not disclose quantization, including GLM-5.3, GLM-5.3-Flash, Kimi K3 and DeepSeek V4.1 Flash. Sonnet uses standard service and five-minute cache creation. Recheck these conditions before extending the estimates beyond this snapshot.
At full modeled utilization, OpenCode Go’s MiMo allocation buys about 6.33× the tokens per dollar of the selected OpenRouter route, and its MiniMax M3 allocation about 5.06×. Z.ai Max’s off-peak GLM-5.3 allocation buys about 4.59×, down from 7.73× on September 7 because a much cheaper GLM-5.3 route appeared. Two subscription allocations now lose to OpenRouter outright: OpenCode Go’s GLM-5.3 allowance costs 66.8% more than the selected route, and a Synthetic pack spent on GLM Flash costs about 7.5% more than InferenceNet’s FP4 route. Chutes Pro’s conditional DeepSeek Flash ceiling also remains more expensive than the selected route. These are different serving configurations, so accepted-work testing still decides whether the cheaper tokens help.
OpenRouter is therefore a useful baseline and overflow source. A subscription only beats it when enough relevant allowance is consumed: Go’s MiMo allocation breaks even at 15.8% utilization and Z.ai Max’s GLM-5.3 allocation at 21.8%, while Synthetic’s Kimi K3 example needs about 51.8% against the current selected quote. The OpenRouter comparison CSV includes cash-budget token capacity; the subscription comparison CSV includes savings and break-even utilization for every mapped, quantifiable offer.
Convert the allowance before comparing it#
Z.ai illustrates why a flat subscription cannot be reduced to one universal token price. Credits depend on the selected model, uncached input, cached input, output and time of day. Off-peak model usage consumes half the normal credits.
Applying its credit formula to the example workload gives these fully utilized, off-peak, steady-state 30-day equivalents:
Lite7.7× API value
- Monthly fee
- $18
- GLM-5.3 output + input
- 2.0M + 430M
- Direct API equivalent
- $138
- Value multiple
- 7.7×
Pro10.3× API value
- Monthly fee
- $80
- GLM-5.3 output + input
- 12.1M + 2581M
- Direct API equivalent
- $826
- Value multiple
- 10.3×
Max11.5× API value
- Monthly fee
- $168
- GLM-5.3 output + input
- 28.2M + 6021M
- Direct API equivalent
- $1,928
- Value multiple
- 11.5×
The weekly allowance is prorated for comparison; it is not a guaranteed calendar-month entitlement. Five-hour limits, concurrency, tool consumption and uneven demand can reduce realized value. All-peak operation halves the token allowances. The same subscriptions produce approximately 2.6–4.0× ordinary GLM Flash API value in this scenario, because Flash’s direct API is already much cheaper. Z.ai’s own documentation describes the monthly quota as roughly 15–30× the fee at API prices; this measured, cache-heavy workload converts to 7.7–11.5× for GLM-5.3. The provider does not state the workload behind its figure. Its subscription page renders prices in the browser, so the $18 / $80 / $168 fees come from the credits-based price set in the official site’s own code, consistent with the documentation’s “starting at just 18 USD per month”. Plan update notice
OpenCode Go has a different mechanism. Some models receive six times the subscription fee in API allowance; others receive three or one-and-a-half times. Allocating its entire qualifying allowance to MiMo V2.6 Flash supports approximately 31.7M output plus 6.76B input tokens for $10 in this scenario. Allocating it to MiniMax M3 supports approximately 3.8M output plus 811M input. Those are alternative uses of one pool, still subject to shorter limits. Each model is also capped at 20% of its monthly limit per five hours and 50% per week. Allowance rules
MiMo’s own plan shows the danger of reading credits as tokens. Each uncached V2.6 Flash input token costs 100 credits; each output token costs 200. V2.6 deducts exactly as V2.5 did, and V2.5 is deprecated on October 21. Its 4.1B-credit Lite plan therefore represents approximately $5.74 of daytime direct API usage for a $6 fee, or $7.18 with the off-peak credit discount. Large numbers need units. Conversion rules, API prices
Synthetic’s $24 weekly allowance represents about $103 per 30 days at Synthetic’s rates. Always compare the serving provider’s cache prices and model configuration with the original API before calling that a direct-provider discount. Limits, model metadata and prices
NanoGPT illustrates the effect of counting cache hits at full input weight. At the measured ratio, its 60M weekly input-unit allowance yields approximately 258M total processed tokens, including 1.21M output, per 30 days on a 1× model. A 2× model halves those estimates. A request-count plan needs another measurement—tokens per billable request—before it can be converted. The local message count is not assumed to equal provider API calls. NanoGPT quota definitions
Utilization is the other half of the arithmetic. A $10 plan containing $60 of relevant API usage breaks even once it replaces more than $10 of API spending. The unused $50 has no value. More packs can buy parallelism, but idle parallelism is still a bill.
Native coding subscriptions belong in the trial too#
Codex through ChatGPT, Claude Code, Cursor, Copilot, Gemini CLI, Devin and Grok Build can be economical when their native workflows produce useful results. Each now has its own row in the comparison above, with the provider’s own prices, models and limits. Only Copilot converts to token capacity from published numbers. Its credits are spent at listed per-token model prices. The others publish request counts, message estimates, relative multiples or undisclosed quotas. Those do not establish a cost per token, and the rows leave the capacity as not established rather than borrowing someone else’s measurement.
Temporary offers, kept separate#
As checked September 25, 2026:
- Z.ai charges all-day usage at the off-peak rate from September 25 to October 7. Its GLM-5.3-Flash campaign, extended to October 7, gives unlimited Flash use through ZCode and AutoClaw daily 15:00–01:00 UTC and doubled quota through other supported agents in that window. Plan notice, Flash campaign
- OpenCode Go gives DeepSeek V4.1 Flash four times its $15 allowance, $60, until September 27, and lists Space Bunny as free for a limited time. Go usage
- deepseekv4pro.com advertises a 2× quota on its Coding tier with no stated end date. Pricing
These offers can improve a trial; the standing calculations exclude all of them. Z.ai’s half-price GLM-5.3-Flash API offer ended September 9 and has been removed from the price data.
The next analysis should measure accepted work#
Run 20–50 representative tasks through the same harness, tests and retry budget. Record input, cache hits, billed output, subscription consumption, cash spend, retries, queue time, review time and whether the change was accepted. Include abandoned attempts in the bill.
The metric is total inference and retry cost divided by changes that pass tests and review. Report elapsed time alongside it: a cheap queue can become expensive when it delays everything else. Allocate the full subscription fee across the measured period, including unused quota, rather than claiming the theoretical maximum discount.
That experiment may favor a cheap default with occasional stronger calls. It may favor one reliable model that finishes sooner. This initial research does not establish which outcome the fleet will produce.
Reproduce and update the comparison#
Download the dated price dataset, measured usage profile, model catalogs, calculator, API comparison CSV, plan capacity CSV, Z.ai comparison CSV, OpenRouter quote CSV, and OpenRouter versus subscriptions CSV. The dataset records source URLs and check dates; the artifact’s comparisons and CSVs are generated from it.
Put calculate.mjs and prices.json in the same directory. With Node 22 or newer:
node calculate.mjs --format api > api.csv
node calculate.mjs --format plans > capacity.csv
node calculate.mjs --format zai > zai.csv
node calculate.mjs --format openrouter > openrouter.csv
node calculate.mjs --format openrouter-plans > openrouter-plans.csv
# Sensitivity scenario; defaults above use the measured profile.
node calculate.mjs --ratio 50 --cache-hit 0.95 --cache-write-share 0.2 --format plans
--cache-write-share is the fraction of non-cache-read input spent creating caches. It preserves separate write pricing while --cache-hit controls reads as a fraction of all input. The calculator does not convert a message count into requests.
For OpenRouter, --topup-credits 10 models the fee on a $10 inference-credit purchase; --budget remains a cash budget. The default amortizes a $100 credit purchase. Ratio overrides reprice the captured provider routes; they do not select a new cheapest provider. Use the recorded endpoint API URLs to refresh route selection for a different workload.
Future revisions can add measured task outcomes, new providers, different cache-hit scenarios and throughput observations without moving this artifact. The page-level verification date advances only after the comparison has been checked as a whole; partial refreshes retain individual source dates. Corrections and changed recommendations belong in the changelog, newest first.
Changelog#
2026-09-25 — Every tracked subscription, first-party only#
- Added fifteen subscriptions, bringing the comparison to twenty-seven: ChatGPT (Codex), Claude (Claude Code), Cursor, GitHub Copilot, Google AI Pro / Ultra, Devin, SuperGrok, Ollama Cloud, Command Code, StepFun Step Plan (international and China), and China-region editions of the GLM Coding Plan, MiniMax Token Plan, Alibaba Bailian Coding Plan and Alibaba Bailian Token Plan.
- Calculated capacity only where the provider publishes both an allowance and the rates it is spent at: Copilot, Ollama Cloud, Command Code, StepFun and the GLM China edition. Every other new row quotes the provider’s own limits and leaves capacity unestablished; no third-party usage measurement was adopted.
- Converted CNY prices at the CFETS central parity of 6.7489 CNY/USD for September 24, recorded in the dataset. Kimi’s China-region price list is served only inside China, so it is noted on the existing Kimi row rather than added.
- Recorded two first-party discrepancies with other trackers. Devin’s own pricing page ends the SWE-2 unmetered promotion on October 10. SuperGrok Lite and Heavy prices are not published on xAI’s pricing page, so they are not shown.
- The native-subscription section now points to these rows instead of summarizing them.
2026-09-25 — First-party re-verification#
- Rechecked every plan, direct API price, OpenRouter route and native subscription against the providers’ own pages and APIs, and advanced the page-level verification date. The September 7 usage measurement is unchanged.
- Every adopted value is now first-party evidence: the provider’s own published page or API, or my own measurement. An independent third-party dataset was used only to find what to recheck; none of its values are adopted. Where a provider publishes no convertible allowance (MiniMax, Kimi, Alibaba, deepseekv4pro.com), the plan stays unquantified and its note records what the provider does publish.
- MiMo: priced V2.6 Flash and V2.6 Pro, which bill and deduct exactly as V2.5 did; V2.5 is deprecated October 21.
- DeepSeek: the V4 Flash names are retired and billed at V4.1 Flash off-peak rates of $0.15 / $0.003 / $0.60, down from $0.22 / $0.007 / $0.66. OpenCode Go’s Flash allocation was recalculated at those rates.
- OpenCode Go: added the GPT 6 Luna allocation and the current per-model limits and catalog (Grok 4.7, MiMo V2.6, DeepSeek V4.1 Flash, MiniMax M2.5, Space Bunny; Omen Alpha removed).
- Synthetic withdrew GLM-5.2 and added DeepSeek V4.1 Flash. Updated the camelStream fleet, Kimi tiers and quota rules, Alibaba’s limits, Chutes’ overage discounts, the NanoGPT and Kilo catalogs, and deepseekv4pro.com’s prices and plan-token quotas.
- Re-captured OpenRouter with the same selection rule, adding an author-route tie-break. Z.ai Max’s GLM-5.3 advantage fell from 7.73× to 4.59×, and OpenCode Go’s GLM-5.3 and Synthetic’s GLM Flash allocations now cost more than the cheapest eligible route. The MiMo cache-price example no longer held and was replaced with a DeepSeek V4 Pro example.
- Removed the expired GLM-5.3-Flash promotional price and rewrote the temporary-offers section.
2026-09-07 — Published as a data artifact#
- Retired the private Notes draft and published the maintained comparison at
/data/inference-plans/. - Kept the report, Markdown twin, calculator, source data and generated CSVs together under one public artifact path.
- This changes the publication surface only; the analysis, measurement and source-verification dates are unchanged.
2026-09-07 — Display the analysis date#
- Added a prominent analysis date near the top of the artifact, separate from publication, measurement and source-verification dates.
2026-09-07 — Index by subscription and model#
- Reorganized the OpenRouter comparison into all twelve researched coding subscriptions, then their models, with tier and time-window rows underneath.
- Included the saved model catalogs; large unpriced catalogs expand within the relevant subscription instead of disappearing from the comparison.
- Added token capacity beside each calculated model/tier cost and retained explicit unknowns for missing allowances, rates and model multipliers.
- Preserved the existing measured mix and dated prices; this is a coverage and presentation update, not a source refresh.
2026-09-07 — OpenRouter route comparison#
- Added nine model comparisons using coherent provider quotes, measured cache composition and the standard credit-purchase fee.
- Added selected provider, quantization, context and status metadata; separated flagged author quotes and current promotional rates.
- Added OpenRouter cash-budget capacity and subscription savings/break-even CSVs, plus a configurable top-up size.
- Found that provider cache pricing can reverse headline-price rankings, and that the selected DeepSeek Flash route undercuts Chutes Pro’s modeled ceiling.
- Retained the earlier usage measurement and subscription snapshot; this revision refreshes OpenRouter quotes only.
2026-09-07 — Normalize subscription costs#
- Converted the measured mix into effective costs per million processed tokens and per million output tokens including associated input for every quantifiable offer.
- Added a main normalization table, expandable model/time-window calculations and matching CSV fields.
- Added explicit same-model direct API savings and break-even utilization, preserving negative savings and unknown comparisons.
- Kept the existing measurement and dated price snapshot; this revision adds arithmetic rather than claiming a new measurement or source refresh.
2026-09-07 — Measured usage, supported models and capacity#
- Replaced the illustrative 20:1 input/output ratio and 80% cache-hit assumption with a fresh local measurement: 213.2:1 and 96.53% of input from cache.
- Normalized separately reported reasoning into billable output and retained distinct cache-read and cache-write buckets.
- Added model coverage for every plan in the main table, dated catalog snapshots and reproducible token-capacity estimates using serving-provider rates.
- Kept unknown quotas, unmetered service and conditional ceilings explicit; added downloadable usage and capacity data.
- Recalculated API examples. The shortlist remains a set of trial candidates; no accepted-work benchmark has yet been run.
2026-09-07 — Initial research draft#
- Compared twelve subscription offerings with direct API costs and native coding alternatives.
- Added reproducible API and Z.ai calculations, source-check dates, and explicit workload assumptions.
- Recorded automation restrictions, conflicting reseller documentation and expiring promotions.
- Established the cost-per-accepted-change experiment as the next analysis; no paid performance benchmark has been run for this note.