DeepSeek API Pricing 2026: Peak, Off-Peak and Providers
DeepSeek is the only major lab that charges by time of day, and most traffic is still running on a version that's been superseded. Here's what peak and off-peak actually cost, which model you should be on, and the cheapest provider for each.

On this page
Table of Contents
- Peak and off-peak: how DeepSeek pricing works
- Which DeepSeek model should you use?
- Who serves DeepSeek
- DeepSeek V4.1 Flash pricing by provider
- DeepSeek V4 Flash 0731 pricing by provider
- What 100 million tokens costs
- Running DeepSeek privately
- Switching to SayGm
- Common questions
- TL;DR
- About SayGm
DeepSeek is the only major lab that charges by time of day. Its API costs half as much outside a 35-hour peak window each week, which means the same request can cost $0.30 or $0.15 per million input tokens depending on when you send it. That is the first thing to understand about DeepSeek API pricing, and it is why comparing DeepSeek to other providers on a single number gets you the wrong answer.
There are four DeepSeek models in circulation right now. This guide covers what DeepSeek charges, which version you should actually be on, who providers it, and what each of them charges.
Peak and off-peak: how DeepSeek pricing works
DeepSeek's published rates have two columns. Peak runs 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Everything else, including all weekend, is off-peak at half price.
| Model | Input, cache hit | Input, cache miss | Output |
|---|---|---|---|
| DeepSeek-Flash, off-peak | $0.003 | $0.15 | $0.60 |
| DeepSeek-Flash, peak | $0.006 | $0.30 | $1.20 |
| DeepSeek-V4-Pro, off-peak | $0.022 | $0.66 | $1.98 |
| DeepSeek-V4-Pro, peak | $0.044 | $1.32 | $3.96 |
Peak is 35 hours out of 168, so most of the week is discounted. For a batch job you control the timing of, that is a real 50% saving for no work. For anything user-facing, you pay peak rates during your users' working day and there is nothing to optimise.
The cache-hit rate is the other number worth staring at. At $0.003 per million tokens against $0.15 fresh, DeepSeek discounts repeated context by 98%. No other lab comes close. If your agent resends the same repository or system prompt on every turn, your real input cost is nearer the cache column than the headline.
Which DeepSeek model should you use?
V4.1 Flash, unless the work is hard reasoning or a long coding session, in which case V4 Pro.
Most DeepSeek traffic is not on V4.1 Flash. It shipped on 10 September and the market is still quoting the July checkpoint, which is smaller, older and on most providers more expensive.
| Model | What it is for | Released | Size |
|---|---|---|---|
| V4.1 Flash | High-volume work where tokens are the constraint: extraction, classification, routing, summarising, the cheap steps inside an agent loop. Takes images in and out, so it also handles documents and screenshots | 10 Sep 2026 | 552B total, 8B active in / 16B out |
| V4 Flash 0731 | The same jobs on a smaller, older model. Worth keeping only if you have pinned to this checkpoint and validated against it | 31 Jul 2026 | 284B total, 13B active |
| V4 Pro | Work where a wrong answer costs more than the tokens do: hard debugging, multi-step reasoning, long refactors that have to hold context | GA 13 Aug 2026 | 1.6T total, 49B active |
| V3.2 | Sealed work on a budget. Its confidential route is a flat rate both ways, which makes it the cheapest private DeepSeek option, and the only other reason to run it is that you validated against it | Previous generation | - |
Pro costs roughly four times Flash. Flash still scores 91.6% on LiveCodeBench and 86.2% on MMLU-Pro, per the model card.
All four are MIT licensed. That is why two dozen providers serve them and why third-party prices sit below DeepSeek's own.
SayGm carries V4.1 Flash, V4 Flash 0731, a sealed route for 0731, and a sealed V3.2. Not V4 Pro. Each page carries the live rate, throughput and success rate for that model.
Who serves DeepSeek
Twenty-one providers serve V4.1 Flash and twenty-six serve the July checkpoint. Here are the ones worth knowing, and what each is good at.
SayGm
Good at: the lowest price on V4.1 Flash, sealed routes for the July checkpoint and V3.2, and frontier models on the same key if you want Claude or GPT next to your DeepSeek traffic. Credit is prepaid and drawn down per request, with no subscription and no top-up fee.
Missing: V4 Pro, fine-tuning, and dedicated capacity.
DeepSeek

Good at: the cache discount, which is 98% off a fresh read and the most aggressive in the market. It is also the only route to V4 Pro, and it posts 100% uptime on OpenRouter's tracking.
Missing: a flat price. You pay double during peak hours, and requests are processed in China, which some security reviews will ask about.
OpenRouter

Good at: reach. One key gets you every provider in both tables below, sorted by price or speed, with automatic failover when one goes down and filters to exclude providers that train on prompts.
Missing: a price advantage, once the 5.5% card top-up fee is counted, and any privacy of its own. A router reads every prompt that passes through it.
DeepInfra

Good at: being cheap on both checkpoints, with 99.84% uptime on V4.1 Flash. Work that can wait in a queue costs less; work that cannot can jump it.
Missing: a sealed route. They also serve compressed builds of some models, which is worth testing against your own evals.
Fireworks

Good at: what an enterprise checklist asks for. Fine-tuning, dedicated deployments with no rate limits, zero data retention by default, and US-only endpoints.
Missing: a competitive rate. They sit near the top of both tables.
Together

Good at: consolidation, if DeepSeek is one of several open-weight models you run. Dedicated endpoints and reserved capacity, across a catalog past 200 models.
Missing: any DeepSeek-specific advantage. The price mirrors DeepSeek's peak rate.
Also worth knowing: Relace and StreamLake are the cheapest routes to the July checkpoint by some distance, Cloudflare serves it at the edge inside an account you may already have, and Phala is the other provider offering a sealed enclave. Check uptime before committing to a small provider; Baseten's V4.1 Flash route sits at 77.73% on OpenRouter's tracking while most of the field is above 99%.
DeepSeek V4.1 Flash pricing by provider

SayGm is the cheapest provider for V4.1 Flash at $0.089 input and $0.356 output per million tokens. Prices below are from each provider's page or OpenRouter's V4.1 Flash listing.
| Provider | Input | Output | Cache read |
|---|---|---|---|
| SayGm | $0.089 | $0.356 | $0.002 |
| Relace | $0.130 | $0.520 | $0.003 |
| Morph (55% off) | $0.135 | $0.540 | $0.004 |
| DeepInfra (30% off) | $0.140 | $0.420 | $0.004 |
| DeepSeek, off-peak | $0.150 | $0.600 | $0.003 |
| Fireworks | $0.220 | $0.660 | $0.007 |
| Together, SiliconFlow, Modal, DigitalOcean | $0.300 | $1.200 | $0.006 |
| Phala | $0.345 | $1.380 | $0.007 |
Two of the cheaper rates are temporary. Morph is running 55% off and DeepInfra 30% off, both labelled as promotions on OpenRouter. The cluster at $0.300 and $1.200 is DeepSeek's peak rate, mirrored.
DeepSeek V4 Flash 0731 pricing by provider
This is where the older checkpoint gets interesting. Providers are clearing capacity on 0731 now that V4.1 Flash exists, and several are pricing it far below anything V4.1 sells for.
| Provider | Input | Output | Cache read |
|---|---|---|---|
| Relace | $0.040 | $0.080 | $0.016 |
| StreamLake | $0.044 | $0.132 | $0.001 |
| Baidu Qianfan | $0.048 | $0.144 | $0.002 |
| DeepInfra | $0.060 | $0.180 | $0.015 |
| Inceptron (EU) | $0.080 | $0.200 | $0.020 |
| SayGm | $0.132 | $0.395 | $0.004 |
| Together, Parasail | $0.140 | $0.280 | $0.030 |
| Fireworks, SiliconFlow | $0.220 | $0.660 | $0.007 |
| Phala, Cloudflare, AtlasCloud | $0.440 | $1.320 | $0.028 |
0731 is a deprecated checkpoint and the prices show it: providers are clearing capacity rather than competing for new traffic. V4.1 Flash is the replacement, at nearly double the parameter count, and on SayGm it costs less than 0731 does. Anything still running on 0731 should move.
What 100 million tokens costs
At 80 million input and 20 million output, which is a typical agentic mix, with no cache discount applied.
| Provider | V4.1 Flash | V4 Flash 0731 |
|---|---|---|
| SayGm | $14 | $18 |
| DeepInfra | $20 | $8 |
| Relace | $21 | $5 |
| DeepSeek direct, off-peak | $24 | not served |
| Fireworks | $31 | $21 |
| DeepSeek direct, peak | $48 | not served |
| Together | $48 | $17 |
Read the two columns together. The cheapest 0731 route beats the cheapest V4.1 route on raw price, and you are paying for a smaller, older model to get there. Against DeepSeek's own API, V4.1 Flash through SayGm is $14 against $24 off-peak and $48 at peak.
Running DeepSeek privately
SayGm serves deepseek-v4-flash-0731-tee and deepseek-v3.2-tee inside Intel TDX enclaves. The prompt is sealed for the whole request, so nobody reads it, including us. Each session returns an attestation report, a signed statement from the chip confirming what ran; our guide to private inference covers how to check one.
The confidential routes cost more than the standard ones: $0.343 and $1.028 for 0731, $0.431 flat both ways for V3.2. The flat rate makes V3.2 the cheaper sealed route at any input-output mix, $43 per 100 million tokens against $48. Phala is the other provider offering an enclave, at $0.440 and $1.320 for 0731.
DeepSeek is a Chinese company and requests to its own API are processed there. On a sealed route that question does not arise, because the prompt is unreadable to everyone in the chain. There is no confidential route for V4.1 Flash yet.
Switching to SayGm
A base URL and a model ID:
from openai import OpenAI
client = OpenAI(
base_url="https://api.saygm.com/v1",
api_key="YOUR_SAYGM_KEY",
)
response = client.chat.completions.create(
model="deepseek-v4.1-flash",
messages=[{"role": "user", "content": "Summarise this changelog for a release note."}],
)
print(response.choices[0].message.content)Swap in deepseek-v4-flash-0731 or deepseek-v4-flash-0731-tee as needed. The quickstart covers keys, and every rate in this post comes from the live catalog at api.saygm.com/v1/models, which is the source of truth if prices move.
Common questions
How much does the DeepSeek API cost?
DeepSeek charges $0.30 per million input tokens and $1.20 output for Flash at peak, halving to $0.15 and $0.60 off-peak. V4 Pro is $1.32 and $3.96 at peak. Third-party providers are cheaper: V4.1 Flash is $0.089 and $0.356 on SayGm, which works out at $14 per 100 million tokens against $24 buying off-peak from DeepSeek directly.
When are DeepSeek's peak hours?
01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Everything else is half price, including weekends.
What is the difference between V4 Pro and V4 Flash?
Pro is 1.6 trillion parameters with 49 billion active, built for hard reasoning and long coding work. Flash is 284 billion with 13 billion active in the July checkpoint, or 552 billion with 8 to 16 active in V4.1. Pro costs about four times more. Most production traffic runs fine on Flash.
Is DeepSeek cheaper than Kimi or GLM?
Yes, by a wide margin, at every tier. V4.1 Flash through SayGm is a fraction of what Kimi K2.6 or GLM 5.3 cost, and Flash-class models are priced to be used in volume rather than reserved for hard problems. Our Kimi pricing guide and GLM pricing guide have the full numbers.
Is there a free DeepSeek API?
Not from DeepSeek, and not from any provider on OpenRouter's V4.1 Flash or 0731 listings at the time of writing. At $0.089 per million input tokens, a million tokens of V4.1 Flash costs nine cents, which is close enough to free for testing.
TL;DR
DeepSeek API pricing has two moving parts: the time of day, and the version. Peak rates are double off-peak, and peak is only 35 hours a week. V4.1 Flash is the current model and SayGm is the cheapest provider for it at $0.089 and $0.356. The older 0731 checkpoint is being discounted hard by several providers, but V4.1 Flash is newer, bigger and still cheaper on our platform than 0731 is.
Point your OpenAI client at api.saygm.com/v1, set the model to deepseek-v4.1-flash, and say gm to AI.
About SayGm
SayGm is a drop-in inference gateway for teams who don't want to just take a company's word that their prompts are private. Every request is routed through a hardware-verified confidential environment - not even SayGm can see what's inside it. That's not a policy, it's provable. Swap in your existing OpenAI, Anthropic, or Gemini code and you're covered in minutes, at transparent, published rates with no hidden markup.
Say gm to AI at saygm.com.
- DeepSeek
- Pricing
- TEE


