Education

Kimi K2.6 and K3 API Pricing: Every Provider Compared

Kimi K2.6 costs $0.28 per million input tokens on SayGm and $0.95 direct from Moonshot. Here's what every major provider charges for K2.6 and K3, what 100 million tokens actually costs, and which host fits your workload.

BSBrittany Seales· Marketing5 min read
Kimi API Pricing Compared

SayGm serves Kimi K2.6 inside a hardware-attested TEE at $0.28 per million input tokens and $1.67 per million output, and Kimi K3 at $1.80 in and $8.99 out. Both are below Moonshot's own list price, and the K2.6 rate is the lowest of the 20 or so providers we checked for this guide.

Table of Contents

That matters more than it did a year ago: since Moonshot AI shipped Kimi K3 in July, a 2.8-trillion-parameter open-weight model that Fortune reports triggered a sell-off in AI and chip stocks, Kimi has gone from "the cheap coding model" to a serious frontier option, and Kimi K2 API pricing has become one of the first things builders check before picking an open-weight model.

This guide covers Moonshot's official Kimi API pricing, a deep dive on the six providers that matter, what each charges for K2.6 and K3, what a realistic month of usage actually costs, whether a free Kimi K2.6 API still exists, and how to pick a provider based on what you are building. If you are already weighing routers against each other, our OpenRouter alternatives comparison is the companion piece.

What does the Kimi API cost in 2026?

Moonshot publishes its own rates on the Kimi developer platform. These are the list prices every other provider is measured against, per million tokens:

ModelInput (cache miss)Input (cache hit)OutputContext
Kimi K2.6$0.95$0.16$4.00262K
Kimi K2.7 Code$0.95$0.19$4.00262K
Kimi K2.7 Code Highspeed$1.90$0.38$8.00262K
Kimi K3$3.00$0.30$15.001M

Two things stand out. First, K3 is roughly three times the price of K2.6 on both sides of the ledger, so the version you pick is the biggest single lever on your bill. Second, Moonshot's cache-hit rate is very aggressive: a repeated system prompt or a long shared document costs about a sixth of a fresh read on K2.6 and a tenth on K3. If your workload re-sends the same context on every turn, as most coding agents do, the cache-hit column is closer to your real input cost than the headline number.

The rest of this guide is about the third-party market, because Moonshot's direct API is neither the cheapest nor the fastest way to run these models. Artificial Analysis clocks Moonshot's own K2.6 endpoint at around 36 tokens per second, against 160 to 220 tokens per second on the faster hosts.

Kimi API providers compared

One thing to know before comparing hosts: the hyperscalers are not in this race yet. As of August 2026, Amazon Bedrock, Azure AI Foundry and Google Vertex AI have not added Kimi K3, and Bedrock's Kimi catalog stops at K2.5. So the realistic options for K2.6 and K3 are Moonshot itself, the big inference specialists, a router, or a confidential gateway. Here is what each of the six that matter is actually good at.

SayGm

SayGm Homepage (SayGm)[https://saygm.com] runs kimi-k2.6-tee at $0.28 in and $1.67 out, and kimi-k3 at $1.80 and $8.99, with a (kimi-k3-tee)[https://saygm.com/models/kimi-k3] route at $2.22 and $11.09. The TEE routes execute the model weights inside an Intel TDX confidential VM, so the prompt is sealed from the host, the network operators and SayGm itself; the attestation report is the proof, and you can fetch it (our guide to private inference walks through how). The K2.6 rate is the lowest of any host we found, and it is the standing price in our machine-readable catalog: no subscription, no top-up fee and no markup on the model maker's rate. What SayGm does not offer: dedicated deployments, fine-tuning, or the raw throughput of a GPU-dense host. If you need 200 tokens per second on a single stream, look at Fireworks. If you need the cheapest and most private way to run a Kimi agent loop, this is it. (Phala also serves Kimi in confidential hardware, at $1.09 and $4.60 for K2.6; the guarantee is comparable, the price is not.)

Moonshot AI (direct)

Moonshot Kimi Platform Homepage

Moonshot'shttps://platform.kimi.ai/) own platform at platform.kimi.ai is the reference: $0.95 in, $4.00 out for K2.6, $3.00 and $15.00 for K3, with the most generous cache-hit pricing of anyone ($0.16 and $0.30). It is the only place to get K2.7 Code and its Highspeed variant, and it is where new checkpoints land first. The trade-offs are speed and jurisdiction. Artificial Analysis measures Moonshot's K2.6 endpoint at about 36 tokens per second, the slowest in its provider set, and OpenRouter's K3 data shows it at over seven seconds to first token. Requests are processed by Moonshot, a Beijing company, which some security teams will not sign off on regardless of the privacy policy. Go direct if you want day-one access to every checkpoint and heavy prefix caching; go elsewhere if you need speed, a specific region, or a hardware privacy guarantee.

OpenRouter

OpenRouter Homepage

OpenRouter is where most developers first meet a new open-weight model, and for Kimi it lists around 20 hosts for K2.6 and 17 for K3 behind one endpoint. Inference is passed through at each host's price; the cost is a 5.5 percent fee on card credit purchases (5 percent by crypto). What you get for that is routing: Balanced, Nitro (fastest) and Exacto (best tool-calling accuracy) sort modes, automatic fallback to a healthy host when one errors, per-host uptime and throughput data, and data-policy filters so you can exclude hosts that train on prompts. For K3 that spread runs from inference.net at $2.10 and $10.95 to Fireworks Fast at $4.50 and $22.50. The things OpenRouter cannot do are the things a router structurally cannot: it sees every prompt in the clear, its privacy is whatever the downstream host's policy says, and the cheapest route on any given day may be a host you have never heard of. We compared it to SayGm in detail in SayGm vs OpenRouter.

Fireworks

Fireworks Homepage

Fireworks charges list price ($0.95 and $4.00 for K2.6, $3.00 and $15.00 for K3) and sells speed and enterprise posture on top. A Fast tier at plus 50 percent and a Priority tier at plus 25 percent cover latency-sensitive and congestion-sensitive workloads, on-demand dedicated deployments remove rate limits entirely, and LoRA fine-tuning is available for K3. Zero data retention is on by default and there are US-only serverless endpoints for regulated industries. That is a policy-level guarantee rather than a hardware one, but for teams whose procurement checklist says "US region, no retention, SLA", Fireworks is the easiest box to tick. Fireworks Fast is also the most expensive way to run K3 on this list.

Together AI

Together AI Homepage

Together mirrors Moonshot's K3 pricing at $3.00 and $15.00 (with $0.30 cached input) and serves the model at FP4 with the full 1M context. What you pay for is the platform around it: serverless and dedicated endpoints under a 99.9 percent SLA, provisioned throughput for predictable load, and a 200-plus model catalog that makes it a sensible place to consolidate if Kimi is one of several open-weight models you run. Together is a K3 shop for Kimi purposes; it does not currently appear on OpenRouter's K2.6 provider list.

DeepInfra

Deepinfra Homepage

DeepInfra is the volume play among the mainstream hosts. K2.6 is $0.75 in and $3.50 out, served at native INT4, the same quantisation Moonshot used for K2-Thinking, so quality is close to reference while the memory footprint drops. It also offers a Flex tier at 0.8x for latency-tolerant batch work and a Priority tier at 1.5x when you need queue position, a useful dial for agent fleets. K3 is $2.85 and $14.25, a small discount to list. DeepInfra says it raised a $107M Series B this year to scale its inference cloud, so capacity is not the concern. Throughput on K2.6 is modest (Artificial Analysis clocks it at roughly 45 tokens per second on the FP4 route), and there is no confidential compute option.

Kimi K2.6 API pricing by provider

Best Priced Kimi K2.6 Providers in 2026. SayGm is the cheapest for both input and output tokens

K2.6 is the model most people mean when they search for Kimi K2 API pricing. Moonshot's Hugging Face model card describes it as a 1-trillion-parameter mixture-of-experts model with 32 billion active parameters under a modified MIT licence, and OpenRouter lists its release as April 2026; and it is the one that put Moonshot on the map for long-horizon coding and multi-agent work. Because the weights are open, more than a dozen hosts run it, and the spread between them is wide.

Prices below are per million tokens for the providers above plus the cheapest and the confidential alternatives, as listed on each provider's page or on OpenRouter's K2.6 listing at the time of writing. SayGm's rates come straight from our live catalog.

ProviderInputOutputCache readRuns in a TEE?
SayGm (kimi-k2.6-tee)$0.28$1.67see catalogYes
OpenRouter (cheapest route + 5.5% fee)from $0.61from $3.59variesNo
Chutes$0.58$3.40$0.06No
DeepInfra (FP4)$0.75$3.50$0.15No
Venice$0.75$3.50$0.16No
Moonshot AI (direct)$0.95$4.00$0.16No
Fireworks$0.95$4.00$0.16No
Phala$1.09$4.60$0.37Yes

A few notes on reading this. OpenRouter's row is a floor. It routes to whichever host is cheapest or fastest depending on your sort mode, so on a Balanced sort you will often land on a $0.75 to $0.95 host, and the 5.5 percent applies to every credit purchase regardless of route. DeepInfra's lowest tier is an FP4 quantisation, which is part of how it gets to $0.75; if you need the full-precision weights, check which quant a host is actually serving before comparing on price alone. Phala is the other host on this list running K2.6 inside confidential hardware, and it shows the usual trade: privacy on most platforms costs a premium of roughly 15 percent over list. SayGm's confidential tier goes the other way, because our open-weight models run on decentralised Bittensor miners inside Intel TDX rather than on a single company's cloud bill. The cheapest tier is also the private one. If "runs in a TEE" is new to you, What Is Private Inference? explains what the enclave does and does not protect.

Kimi K3 API pricing by provider

Kimi K3 API Pricing by provider. SayGm is the cheapest in both input and output tokens

K3 is a different animal: 2.8 trillion total parameters, 104 billion active, a native 1-million-token context, and a licence Moonshot calls the Kimi K3 License rather than MIT. Moonshot's own benchmark table, published with the weights on GitHub, puts it competitive with Anthropic's Claude Fable 5 on reasoning and agentic tasks; treat that as the vendor's claim until independent evals catch up. The pricing reflects that ambition, and the third-party spread is narrower than K2.6's because fewer hosts can serve a model this large at full context.

ProviderInputOutputContextRuns in a TEE?
SayGm (kimi-k3)$1.80$8.991MNo
SayGm (kimi-k3-tee)$2.22$11.091MYes
OpenRouter (cheapest route + 5.5% fee)from $2.22from $11.551MNo
inference.net$2.10$10.951MNo
DeepInfra$2.85$14.251MNo
Phala$2.85$14.251MYes
Moonshot AI (direct)$3.00$15.001MNo
Fireworks$3.00$15.001MNo
Together$3.00$15.001MNo

Most of the big hosts simply mirror Moonshot's $3 and $15. The sub-list-price options are the smaller specialist hosts and SayGm. Note that SayGm lists K3 twice: the standard route and a TEE route. K3 in the enclave costs about 23 percent more than K3 outside it, which is the honest price of running a 2.8T model on attested hardware, and it is still below what Moonshot charges directly.

What 100 million Kimi tokens actually costs

Per-token rates are hard to feel. A more useful number is what a month of real usage comes to. The table below assumes 100 million tokens at a 4:1 input-to-output ratio (80M in, 20M out), which is typical for an agentic coding workload where the model reads far more than it writes. No cache discounts are applied, so treat these as ceilings.

ProviderK2.6, per 100M tokensK3, per 100M tokens
SayGm$56 (TEE)$323 standard, $399 TEE
OpenRouter (cheapest route + 5.5% fee)$121$408
Chutes$114$540
DeepInfra$130$513
Moonshot AI (direct) / Fireworks$156$540
Phala (TEE)$179$513

At a billion tokens a month, multiply everything by ten: K2.6 on SayGm lands around $560, against roughly $1,560 direct from Moonshot. The gap on K3 is smaller in percentage terms but larger in dollars, about $2,170 a month at that volume.

[IMAGE: After the 100M-token table. Stacked bar chart, K2.6 vs K3 monthly cost at 100M and 1B tokens, SayGm vs Moonshot direct. Alt text: "Chart of monthly Kimi API cost at 100M and 1B tokens on SayGm versus Moonshot direct." Suggested filename: kimi-api-monthly-cost-100m-1b-tokens.png]

Is there a free Kimi K2.6 API?

Sometimes. OpenRouter has periodically listed a kimi-k2.6:free endpoint, and at the time of writing it is not on the model page. Free variants on routers tend to appear at launch, run with tight rate limits, and disappear once the sponsoring provider's promotional budget runs out. They are fine for a weekend experiment. They are not something to build a product on, and the prompts you send to a free endpoint are generally routed to whichever host is subsidising it, on that host's data terms.

Moonshot itself offers a consumer chat app with a free tier, but the developer API is pay-as-you-go from the first token. The practical answer for most builders is a cheap paid endpoint rather than a free one: at SayGm's rate, a million input tokens of K2.6 costs 28 cents. We ran the same comparison for OpenRouter's free models against SayGm's paid rates if you want the detail.

Which Kimi provider fits your workload?

Cheapest possible K2.6: SayGm, then Chutes. If you are running a high-volume agent loop and every request re-reads the same repository, also compare cache-read rates, since that column will dominate your bill.

Fastest K2.6: Nebius and Azure top Artificial Analysis's throughput chart at over 220 tokens per second, with CoreWeave and Parasail close behind. You pay for it. Nebius's blended rate is more than double SayGm's.

Longest context on K3: every serious host offers the full 1M window, so context is not a differentiator. Pick on price and privacy.

Prompts you cannot afford to leak: this is the case for a confidential route. Kimi is popular with teams pointing it at proprietary codebases, and a codebase is exactly the kind of prompt a security review will ask about. On SayGm's TEE tier the model weights and your prompt sit inside the same Intel TDX enclave, so the host, the network operators and SayGm itself structurally cannot read the request. The attestation report lets you verify that yourself. Our guide to private inference on TEE hardware explains how to check the attestation yourself. Phala offers a comparable guarantee at a higher price.

Failover across many hosts: OpenRouter, with its fallback routing and 20 Kimi hosts, is the strongest option if uptime across providers matters more than who can read the prompt. SayGm's cascade mode covers the same need across the models in our catalog.

Drop-in replacement for an existing OpenAI or Anthropic integration: any OpenAI-compatible host works, but check that tool calling and streaming behave the same on the quant being served. SayGm supports Chat Completions, the Responses API and Anthropic's Messages API against the same base URL.

How to call Kimi through SayGm

Switching is a base URL and a model ID. If you already have OpenAI-style code, it looks like this:

python
from openai import OpenAI

client = OpenAI(
    base_url="https://api.saygm.com/v1",
    api_key="YOUR_SAYGM_KEY",
)

response = client.chat.completions.create(
    model="kimi-k2.6-tee",
    messages=[{"role": "user", "content": "Refactor this function to be async."}],
)
print(response.choices[0].message.content)

Swap kimi-k2.6-tee for kimi-k3 or kimi-k3-tee as needed. The quickstart covers key creation and the integrations page has the config for Cursor, Cline, Claude Code and Codex CLI. Every rate in this post is also available programmatically at api.saygm.com/v1/models, which is the source of truth if prices move after publication.

Common questions about Kimi pricing

What is the difference between Kimi K2, K2.6 and K3?

Kimi K2 is the original 1T-parameter model from 2025. K2.6 is the current release of that line: same architecture, updated training, native vision, and much stronger agentic coding scores (80.2 percent on SWE-Bench Verified per Moonshot's model card). K3 is a new, larger family at 2.8T parameters with a 1M context and a different licence. When people search "kimi k2 api pricing" today they almost always mean K2.6.

Is Kimi cheaper than Claude or GPT?

At list price, yes, by a wide margin. Moonshot's K3 output rate is $15 per million tokens, and Fortune notes that undercuts the equivalent Anthropic frontier tier several times over. K2.6 at $4 output direct, or $1.67 through SayGm, is cheaper still. Whether it is cheaper for your task depends on how many tokens each model needs to finish the job, so compare on a real workload rather than on rates alone.

How does cache-hit pricing work on Kimi?

Moonshot, and most hosts that mirror its pricing, charge a reduced rate when the prefix of your prompt matches something the server has recently processed. On K2.6 that is $0.16 instead of $0.95 for every million tokens you send. Coding agents that resend the same system prompt and file context on every turn can see the majority of their input billed at the cache rate.

Should I use the Kimi coding plan or the API?

Moonshot's coding plan is a subscription for its own Kimi Code tool, with a usage allowance rather than per-token billing. It suits individuals using one tool heavily. The API, through Moonshot or a third-party host, suits anything you are building for other people, anything that needs to run unattended, and any workload where you want to route between models. SayGm has no subscription and no top-up fee; you prepay credit and it is drawn down per request.

TL;DR

Kimi K2 API pricing in 2026 comes down to three numbers. Moonshot's list price for K2.6 is $0.95 in and $4.00 out; the cheapest third-party route we found is SayGm's TEE tier at $0.28 and $1.67; and K3 costs about three times K2.6 wherever you run it, with SayGm at $1.80 and $8.99 against Moonshot's $3.00 and $15.00. At 100 million tokens a month that is roughly $56 versus $156 for K2.6. Pick a fast host if latency is your constraint, a confidential host if your prompts are, and check the quantisation before you trust any sub-list price.

If your Kimi workload touches code or data you would rather no one else could read, grab a SayGm API key, point your existing client at api.saygm.com/v1, and say gm to AI.

About SayGm

SayGm is a drop-in inference gateway for teams who don't want to just take a company's word that their prompts are private. Every request is routed through a hardware-verified confidential environment - not even SayGm can see what's inside it. That's not a policy, it's provable. Swap in your existing OpenAI, Anthropic, or Gemini code and you're covered in minutes, at transparent, published rates with no hidden markup.

Say gm to AI at saygm.com.

Website | Twitter | Discord | Blog | Medium | Docs

  • Kimi
  • Pricing
BS

Brittany SealesMarketing

Saying gm to marketing (and AI)

X ↗