# Pricing

> Release history is kept in the reference notes below. Use the current table and September 16 refresh note for present scores and estimates. The September 11 roster cleanup removes two retired models.

> **Latest refresh (September 16, 2026):** OpenRouter's live benchmark card puts the unversioned Qwen3.8 Max entry at **#2 with 53.4**. The public model API no longer exposes that unversioned ID and canonicalizes it to 0902; the exact `qwen/qwen3.8-max-0902` metadata row now reports **45.4**. The table preserves both values as separate records rather than silently merging them. Qwen Cloud's current PAYG API rates remain **$2 / $0.25 implicit cache / $6** for Qwen3.8 Max. DeepSeek now serves V4.1 Flash as `deepseek-flash` at peak rates of **$0.30 / $0.006 cache hit / $1.20**; the legacy V4 Flash alias is retired. Anthropic kept Sonnet 5 at **$2 / $0.20 / $10**, and OpenAI's Sol rate is **$4 / $0.40 / $20** through at least November 21. GLM 5.3 Flash's 50% launch promotion ended September 9, so list rates are now the applicable baseline.

> **⏰ June 1, 2026 — GitHub Copilot switched to usage-based billing (AI Credits) today.**
>
> Before this change, Copilot used **premium request-based billing** — each model had its own multiplier, and every request consumed `multiplier × 1` from your monthly premium-request allowance. Now **every interaction burns AI credits** based on actual token consumption. Agent mode and complex multi-file tasks consume significantly more tokens than simple Q&A, which means your 7,000 Pro+ credits can disappear fast if you're using frontier models.
>
> **The practical workaround:** use cheaper alternative models (DeepSeek V4 Flash, MiMo, Qwen) that are still powerful enough for coding — often at **5–55× less cost** than the Copilot defaults. The tables below show the exact comparison.
>
> 1 AI credit = $0.01 USD. All paid plans include a monthly credit allowance:
>
> | Plan | Price/mo | Base credits | Flex allotment | Total monthly |
> | ---- | -------- | ------------ | -------------- | ------------- |
> | Pro  | $10      | 1,000        | 500            | **1,500**     |
> | Pro+ | $39      | 3,900        | 3,100          | **7,000**     |
> | Max  | $100     | 10,000       | 10,000         | **20,000**    |
>
> Code completions remain unlimited and **not** billed. Auto model selection gets a 10% discount.

All prices below are in **USD per 1M tokens** (non-cached). To convert to AI credits, multiply by 100 (e.g., $5.00/1M = 500 credits/1M). DeepSeek rows use peak rates for the comparison; its official off-peak rates are half. Provider promotions are shown separately in the notes; OpenRouter's effective, batch, or free-provider prices are excluded.

All models are listed together below, sorted by Cost per intelligence ascending (lower is better). Models without a Cost per intelligence are ordered by estimated session cost ascending. Session cost assumes ~10K input + ~2K output tokens per turn, 50 turns.

| Model                     | Provider  | Cost per intelligence | Intelligence Score | Input (per 1M)                | Cached input                  | Output (per 1M)               | Est. session | Context window |
| ------------------------- | --------- | --------------------- | ------------------ | ----------------------------- | ----------------------------- | ----------------------------- | ------------ | -------------- |
| **GLM 5.3 Flash**         | Z.ai      | **~$0.0030**          | **41.9**           | $0.15                         | $0.03                         | $0.50                         | ~$0.13       | 1M             |
| **MiMo V2.5**             | Xiaomi    | **~$0.0045**          | **22.3**           | $0.14                         | $0.0028                       | $0.28                         | ~$0.10       | 1M             |
| **GPT-5.6 Luna**          | OpenAI    | **~$0.0059**          | **37.5**           | $0.20                         | $0.02                         | $1.20                         | ~$0.22       | 1M             |
| **DeepSeek V4.1 Flash**   | DeepSeek  | **~$0.0068**          | **39.5**           | $0.30                         | $0.006                        | $1.20                         | ~$0.27       | 1M             |
| **MiniMax M3**            | MiniMax   | **~$0.0091**          | **29.6**           | $0.60 (≤512K) / $1.20 (>512K) | $0.12 (≤512K) / $0.24 (>512K) | $2.40 (≤512K) / $4.80 (>512K) | ~$0.27       | 1M             |
| **MiMo V2.5 Pro**         | Xiaomi    | **~$0.0114**          | **26.4**           | $0.435                        | $0.0036                       | $0.87                         | ~$0.30       | 1M             |
| **MiniMax M3 Priority**   | MiniMax   | **~$0.0139**          | **29.6**           | $0.90 (≤512K) / $1.80 (>512K) | $0.18 (≤512K) / $0.36 (>512K) | $3.60 (≤512K) / $7.20 (>512K) | ~$0.41       | 1M             |
| **Gemini 3.8 Flash**      | Google    | **~$0.0182**          | **41.2**           | $0.75                         | $0.075                        | $3.75                         | ~$0.75       | 1M             |
| **GLM 5.3**               | Z.ai      | **~$0.0254**          | **44.9**           | $1.40                         | $0.26                         | $4.40                         | ~$1.14       | 1M             |
| **DeepSeek V4 Pro 0813**  | DeepSeek  | **~$0.0292**          | **36.3**           | $1.32                         | $0.044                        | $3.96                         | ~$1.06       | 1M             |
| **Qwen 3.8 Max (0803)²⁵** | DashScope | **~$0.0300**          | **53.4**           | $2.00                         | $0.25                         | $6.00                         | ~$1.60       | 1M             |
| **Qwen 3.8 Max (0902)²⁶** | DashScope | **~$0.0352**          | **45.4**           | $2.00                         | $0.25                         | $6.00                         | ~$1.60       | 1M             |
| **Grok 4.6**              | xAI       | **~$0.0360**          | **44.4**           | $2.00                         | $0.50                         | $6.00                         | ~$1.60       | 500K           |
| **GPT-5.6 Terra**         | OpenAI    | **~$0.0520**          | **42.3**           | $2.00                         | $0.20                         | $12.00                        | ~$2.20       | 1M             |
| **Claude Sonnet 5**       | Anthropic | **~$0.0521**          | **38.4**           | $2.00                         | $0.20                         | $10.00                        | ~$2.00       | 1M             |
| **Kimi K3**               | Moonshot  | **~$0.0685**          | **43.8**           | $3.00                         | $0.30                         | $15.00                        | ~$3.00       | 1M             |
| **GPT-5.6 Sol**           | OpenAI    | **~$0.0849**          | **47.1**           | $4.00                         | $0.40                         | $20.00                        | ~$4.00       | 1M             |
| **Claude Opus 5**         | Anthropic | **~$0.0986**          | **50.7**           | $5.00                         | $0.50                         | $25.00                        | ~$5.00       | 1M             |
| **Claude Fable 5.1**      | Anthropic | **~$0.1873**          | **53.4**           | $10.00                        | $0.25                         | $50.00                        | ~$10.00      | 1M             |
| **GPT-6 Astra**           | OpenAI    | **~$0.1894**          | **52.8**           | $10.00                        | $1.00                         | $50.00                        | ~$10.00      | 1M             |
| **Claude Fable 5**        | Anthropic | **~$0.2012**          | **49.7**           | $10.00                        | $1.00                         | $50.00                        | ~$10.00      | 1M             |

> **Reference notes:** The details below preserve release and pricing context. Use the current table and September 10 refresh note for present values.

⁵ **MiniMax M3 Priority** is not a separate model — it is the same `MiniMax-M3` weights invoked with `"service_tier": "priority"` in the request body. Priority costs **1.5× Standard** across input, output, and cache reads (list prices shown above; effective rates after the standing 50% off are $0.45/$1.80/$0.09 ≤512K and $0.90/$3.60/$0.18 >512K), in exchange for **priority admission** (faster responses, fewer failures during MiniMax peak hours — typically 15:00–17:30 weekdays). Capabilities, context window (1M, guaranteed 512K), vision, tool calling, rate limits (200 RPM / 10M TPM), and thinking modes are identical to Standard. To enable it on the custom-endpoint entry, add `"service_tier": "priority"` to the `requestBody` of the single `MiniMax-M3` block (and remove it to go back to Standard). See [docs/models/minimax.md](models/minimax.md#4-m3-priority-tier-optional) and [docs/research/minimax-m3-priority.md](research/minimax-m3-priority.md) for the full breakdown.

⁶ **Claude Sonnet 5** launched at **$2.00 / $0.20 / $10.00** (input / cached / output). Anthropic's current pricing makes those rates standard; the previously announced September 1 increase to $3.00 / $0.30 / $15.00 will not occur. See [Anthropic's pricing page](https://platform.claude.com/docs/en/about-claude/pricing) for the latest.

⁷ **Kimi K3** was released on **July 16, 2026**. 2.8T parameters (open-source, weights by July 27, 2026). 1M context window. Always-thinking reasoning model — uses `reasoning_effort` (not the K2.x `thinking` parameter). Fixed sampling: `temperature=1`, `top_p=0.95`. Pricing is flat (no tiering by context length). See [Kimi K3 pricing](https://platform.kimi.ai/docs/pricing/chat-k3) and the [Artificial Analysis model page](https://artificialanalysis.ai/models/kimi-k3).

⁸ **GPT-5.6** pricing is **$0.20 / $0.02 / $1.20** for Luna, **$2.00 / $0.20 / $12.00** for Terra, and currently **$4.00 / $0.40 / $20.00** for Sol (input / cached input / output per 1M tokens). The Sol rate is a promotional reduction from $5.00 / $0.50 / $30.00 and is available at least through **November 21, 2026**. OpenAI reduced Luna by 80% and Terra by 20% on July 30, 2026. OpenAI bills cache writes at 1.25x the uncached input rate for GPT-5.6 and later. The Batch API provides an additional 50% discount for asynchronous jobs. All three tiers have a 1M-token context window and support text + image input. See [OpenAI's API pricing](https://developers.openai.com/api/docs/pricing) and [GPT-5.6 announcement](https://openai.com/index/gpt-5-6/).

¹⁰ **Claude Opus 5** was released on **July 24, 2026**. AA Intelligence Index score (**63.1**) confirmed by [Artificial Analysis](https://artificialanalysis.ai/models/claude-opus-5). Pricing is $5.00 / $0.50 / $25.00 per MTok input/cached/output. 1M context window, text + image input, adaptive reasoning. Also available in Fast mode ($10/$50 per MTok input/output) and Batch ($2.50/$12.50 per MTok input/output). Uses the newer Claude tokenizer (~30% more tokens than pre-4.7 models). See [Anthropic's Opus 5 announcement](https://www.anthropic.com/news/claude-opus-5).

¹¹ **Historical Qwen 3.8 Max launch record:** the August 3, 2026 launch package used the unversioned `qwen3.8-max` name. An earlier Artificial Analysis snapshot reported **58.1**; the current OpenRouter ranking now distinguishes the unversioned **0803** row (**53.4**) from the separate **0902** row (**45.4**). Use the current rows and footnotes ²⁵–²⁶ for comparisons rather than carrying the earlier score forward. See the [Qwen Cloud model page](https://www.qwencloud.com/models/qwen3.8-max), [AA model page](https://artificialanalysis.ai/models/qwen3-8-max), and [OpenRouter model page](https://openrouter.ai/qwen/qwen3.8-max).

¹⁴ **Grok 4.6** (xAI/SpaceXAI, released **August 12, 2026**) is now a **GitHub Copilot native** model (GA). AA Intelligence Index **61** (high reasoning preset, #6/188). 500K context window, text + image input, text output. Priced at $2.00 / $0.50 / $6.00 per 1M input/cached/output (75% cache discount). See the [AA model page](https://artificialanalysis.ai/models/grok-4-6), the [xAI announcement](https://x.ai/news/grok-4-6), and the [GitHub Copilot models & pricing page](https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing).

¹⁶ **Claude Fable 5** (Anthropic, released **June 9, 2026**) is now a **GitHub Copilot native** model (GA, Powerful; Anthropic's new flagship tier, above Opus). AA Intelligence Index **62.1** (**#3/188**; AA page lists a rounded **62** on its [model page](https://artificialanalysis.ai/models/claude-fable-5)). Currently **#1 in Text, Agent, Code, and Overall Arena**. 1M context window, text + image input, adaptive reasoning. Priced at **$10.00 / $1.00 / $50.00** per 1M input/cached/output (90% cache discount; cache write $12.50). Uses the newer Claude tokenizer (~30% more tokens than pre-4.7 models). **Superseded as Anthropic's flagship by Claude Fable 5.1 (September 1, 2026 — see footnote ²³)**. See the [GitHub Copilot models & pricing page](https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing).

¹⁸ **GLM 5.3** is Z.ai's newest flagship (released **August 18, 2026**) with post-training improvements over the previous GLM flagship (~50% coding gain on Z.ai Code Bench). AA Intelligence Index **59.5** (OpenRouter's AA-sourced benchmark; AA's [model page](https://artificialanalysis.ai/models/glm-5-3) lists a rounded **60**, **#8/182**), Coding **74.8**, Agentic **59.1**. Pricing is **$1.40** input, **$0.26** cached, and **$4.40** output per 1M tokens (modeled session cost ~$1.14). 1M context window, 753B params, text-only. Always-thinking: `thinking.type` only supports `enabled` and `reasoning_effort` accepts `low` / `high` / `max` (default `max`). See the [Z.ai model page](https://docs.z.ai/guides/llm/glm-5.3).

¹⁹ **DeepSeek V4.1 Flash** and **DeepSeek V4 Pro 0813** are the current API builds. Flash uses canonical model ID `deepseek-flash`, supports text + image input, and has AA Intelligence **39.5** in the current OpenRouter metadata. Pro remains text-only with AA Intelligence **36.3**. DeepSeek's current peak/off-peak pricing is Flash **$0.30 / $0.006 / $1.20** and Pro **$1.32 / $0.044 / $3.96** per 1M input/cache-hit/output at peak, with off-peak rates at half those values. The legacy `deepseek-v4-flash` name remains accepted but is retired. See the [DeepSeek pricing page](https://api-docs.deepseek.com/quick_start/pricing), [update log](https://api-docs.deepseek.com/updates/), and [canonical model record](models/deepseek.md).

²⁰ **GLM 5.3 Flash** (`glm-5.3-flash`) is Z.ai's first native multimodal model in the GLM-5 series (released **August 26, 2026**; tested pre-release as "ox-alpha"). Hybrid sparse + linear attention architecture with 320B total / 18B active parameters (open weights, MIT license) and a 1M-token context window; text + image input, text output. AA Intelligence Index **57.5** (OpenRouter's AA-sourced benchmark; AA's [model page](https://artificialanalysis.ai/models/glm-5-3-flash) lists a rounded **57** and ranks it **#3/110** within its open-weights size class), Coding **71.5**, Agentic **58.2**. Launch pricing is **$0.15** input / **$0.03** cached / **$0.50** output per 1M (list), with a **50% off** launch promo (effective **$0.075 / $0.015 / $0.25**) through **September 9, 2026** (24:00 UTC+8). At list price the modeled session cost is ~$0.13 and the cost per intelligence is ~$0.0022 — the cheapest row in this table. Always-thinking: `thinking.type` only supports `enabled`, `reasoning_effort` accepts `low` / `high` / `max` (default `max`). See the [Z.ai model page](https://docs.z.ai/guides/llm/glm-5.3-flash) and [AA model page](https://artificialanalysis.ai/models/glm-5-3-flash).

²¹ **Gemini 3.8 Flash** (`gemini-3.8-flash`) was released on **September 2, 2026** — Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and enterprise workflows. AA Intelligence Index **58.7** (OpenRouter's AA-sourced benchmark; AA's [model page](https://artificialanalysis.ai/models/gemini-3-8-flash) rounds to **59**, **#17/196**), Coding **76.3**, Agentic **50.0**. 1M context window (1,048,576 input / 65,536 output), text + image + speech + video + PDF input, text output; configurable thinking (`low`/`medium`/`high` — `minimal` is not supported). Fast at **305.5 t/s** per AA but **very verbose** (120M output tokens on the Intelligence Index vs a 71M median) with a high TTFT (13.39 s), so expect output-token costs above what the score alone suggests. **Not yet GitHub Copilot native** (GitHub's Google lineup still ends at 3.7 Flash) — priced from the Gemini API at the **same promotional rates as 3.7 Flash**: $0.75 / $0.075 / $3.75 per 1M input/cached/output through **December 31, 2026**, then $1.50 / $0.15 / $7.50 from January 1, 2027 (90% cache discount; Batch API 50% off; cache storage $0.50/1M per hour during the promo). Vendor-reported: HLE-Verified **54.9**. See the [Google announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/), the [Gemini API model page](https://ai.google.dev/gemini-api/docs/models/gemini-3.8-flash), and the [AA model page](https://artificialanalysis.ai/models/gemini-3-8-flash).

²³ **Claude Fable 5.1** was released on **September 1, 2026** — Anthropic's new flagship, improving on Fable 5 across the board with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work (long code refactors, front-end/visual code generation, finance and analysis), and tending to be more concise in plans and summaries. AA Intelligence Index **65.7** (OpenRouter's AA-sourced benchmark; AA's [model page](https://artificialanalysis.ai/models/claude-fable-5-1) rounds to **66** and ranks it **#1/196**) — the **highest score in this table** — with Coding **81.6** and Agentic **61.3**. GitHub Copilot native (GA, Powerful). 1M context window, text + image input, text output, adaptive reasoning; uses the newer Claude tokenizer (~30% more tokens than pre-4.7 models). Priced at **$10.00 / $0.25 / $50.00** per 1M input/cached/output — the same rates as Fable 5 except **cached input drops from $1.00 to $0.25** (97.5% cache discount); cache write $12.50. Modeled session cost stays ~$10.00, and CPI improves only to ~$0.15 (vs Fable 5's ~$0.161) — still the second-most expensive scored row. At max effort it is slow (66.2 t/s per AA) and very verbose (140M output tokens on the Intelligence Index vs a 71M median; TTFT 283.83 s). Like Fable 5, Anthropic retains request data by default, with zero-data-retention available through the end of 2026 under a time-bound exemption. See the [GitHub Copilot models & pricing page](https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing) and the [AA model page](https://artificialanalysis.ai/models/claude-fable-5-1).

²⁴ **GPT-6 Astra** was released on **September 3, 2026** as `gpt-6-astra`, an OpenAI API model and **GitHub Copilot native** model (GA, Powerful). It supports reasoning, text + image input, text output, and a 1M-token context window. Standard pricing is **$10.00 / $1.00 / $50.00** per 1M input/cached/output tokens; cache writes are $12.50/M, and prompts above 272K input tokens use long-context rates. The modeled session cost is ~$10.00 and CPI is ~$0.163. See the [GPT-6 Astra announcement](https://openai.com/index/gpt-6-astra/), [OpenAI API pricing](https://developers.openai.com/api/docs/pricing), and [GitHub Copilot models & pricing](https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing).

²⁵ **Qwen 3.8 Max (0803)** is the August 3, 2026 launch checkpoint represented by OpenRouter's unversioned ranking card [`qwen/qwen3.8-max`](https://openrouter.ai/qwen/qwen3.8-max). The September 10 ranking card places it at **#2 with 53.4**. It shares Qwen Cloud's **$2.00 / $0.25 implicit cache / $6.00** PAYG rates, 1M context, 131K maximum output, and text/image/video input with the 0902 snapshot. The validated DashScope setup continues to use provider ID `qwen3.8-max` and the proxy forwards that ID unchanged.

²⁶ **Qwen 3.8 Max (0902)** is a separate OpenRouter snapshot, [`qwen/qwen3.8-max-0902`](https://openrouter.ai/qwen/qwen3.8-max-0902), released September 4, 2026. The current public model API reports an AA Intelligence Index of **45.4** for this exact ID. It has the same published **$2.00 / $0.25 implicit cache / $6.00** rates, 1M context, 131K maximum output, and text/image/video input. The OpenRouter slug is not automatically a valid DashScope model ID; see [the Qwen setup guide](models/qwen.md) before changing the upstream URL.

Cost per intelligence = estimated session cost ÷ Intelligence Index score. Session cost assumes ~10K input + ~2K output tokens per turn, 50 turns.

> **Notes:**
>
> - **DeepSeek** input pricing shown is the **peak cache miss** price. Peak cache hits are $0.006/M for V4.1 Flash and $0.044/M for Pro; off-peak rates are half. See footnote ¹⁹ for the full schedule.
> - **MiMo** input pricing shown is the **cache miss** price. Cache hits are essentially free for V2.5 Pro ($0.0036/M, ~120× cheaper) and V2.5 ($0.0028/M, ~50× cheaper). A Xiaomi price cut took effect on 2026-05-27.
> - **Anthropic (Claude)** models also have a cache write cost ($6.25/MTok for Opus, $3.75/MTok for Sonnet, $1.25/MTok for Haiku). Opus 5, Sonnet 5, and Fable 5 / 5.1 use a new tokenizer that produces approximately 30% more tokens for the same text.
> - **OpenAI** models support cached input at 0.1× base input rate.
> - **Qwen** models use **tiered pricing** — determined by total input tokens per request. Prices above are for non-thinking mode.
> - **Kimi** official tables list **Cache Hit before Cache Miss** (opposite order to our table). The rows below transpose them so "Input" = cache miss and "Cached input" = cache hit.
> - **DashScope** offers a **free quota** of 1M input + 1M output tokens per model, valid for 90 days.
> - **MiniMax M3** uses **tiered pricing** — input price doubles above 512K input tokens. Cache hits are priced at 20% of the input rate ($0.12/M ≤512K, $0.24/M >512K). A **permanent 50% off** discount applies to all MiniMax-M3 pay-as-you-go usage (Standard and Priority tiers), making the effective rates half the list prices above.
> - **MiniMax M3 Priority** is not a separate model — it is the same `MiniMax-M3` weights invoked with `"service_tier": "priority"` in the request body. Priority costs **1.5× Standard** across input, output, and cache reads, in exchange for **priority admission** (faster responses, fewer failures during MiniMax peak hours — typically 15:00–17:30 weekdays). Capabilities, context window (1M, guaranteed 512K), vision, tool calling, rate limits (200 RPM / 10M TPM), and thinking modes are identical to Standard. See [docs/research/minimax-m3-priority.md](research/minimax-m3-priority.md) for the full breakdown.
> - **GLM** models support prompt caching — cache hits are priced at $0.26/M for 5.3 and $0.03/M for 5.3 Flash (list; $0.015/M during the launch promo).
> - **MiMo** offers a **Token Plan** subscription model with discounted rates and a free cache-writing promotion.
> - For typical Copilot chat usage (short-to-medium prompts), you'll almost always fall in the lowest pricing tier.

> **How long does 7,000 credits last?** A Pro+ subscriber running 50-turn sessions can estimate monthly capacity by multiplying each modeled session cost by 100 to convert it to AI credits, then mixing models as needed.

> Prices last verified: September 16, 2026. Always check the official pages for the latest rates:
>
> - [GitHub Copilot models & pricing](https://docs.github.com/en/copilot/reference/copilot-billing/models-and-pricing)
> - [OpenAI pricing](https://openai.com/api/pricing/)
> - [OpenAI GPT-5.6 announcement](https://openai.com/index/gpt-5-6/)
> - [OpenAI GPT-5.6 price update](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/)
> - [Anthropic (Claude) pricing](https://platform.claude.com/docs/en/about-claude/pricing)
> - [Google Gemini pricing](https://ai.google.dev/pricing)
> - [DashScope pricing](https://www.alibabacloud.com/help/en/model-studio/billing-for-model-studio)
> - [DeepSeek pricing](https://api-docs.deepseek.com/quick_start/pricing)
> - [MiMo pricing](https://mimo.mi.com/docs/en-US/pricing)
> - [MiniMax pricing](https://platform.minimax.io/docs/pricing/overview)
> - [Z.ai (GLM) pricing](https://docs.z.ai/guides/overview/pricing)
> - [Kimi pricing](https://platform.kimi.ai/docs/pricing)
> - [xAI model pricing](https://docs.x.ai/docs/models)
