Open weights, run inside a TEE. Inference happens sealed, so the prompt stays inside the enclave.
Released
Sep 2026Modalities
Text · Tool calling · Image inputBest price on SayGM
-10.0%In
Cached
Out
per Mtok*
Create an API key and set it as SAYGM_API_KEY in your terminal.
curl "https://api.saygm.com/v1/chat/completions" \
-H "Authorization: Bearer $SAYGM_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "deepseek-v4.1-flash-tee",
"messages": [
{
"role": "user",
"content": "Hello!"
}
]
}'SayGM's best available rate for deepseek-v4.1-flash-tee next to the published list price — 10.0% off.
| Dimension | SayGM price | List price |
|---|---|---|
| Input | $0.27per Mtok | $0.30 |
| Output | $1.08per Mtok | $1.20 |
| Cache read | $0.0054per Mtok | $0.006 |
What buyers have actually paid on this model recently, across every provider and cache hits.
Input
$0.27
$0.30Cached input
$0.005
$0.006Output
$1.08
$1.20Requests are load balanced across every provider serving this model, each offering its own rate, so what you pay is a blend rather than a single number. These are the blended rates actually paid — priced at what a provider offered, capped at list.
Measured over 168 windows: 52 requests and 68,477 input-side tokens, short of the 2,000,000 tokens the window walks for. Most recent window closed 2026-09-28 12:04 UTC.
Every provider currently serving this model, ranked by discount.
| Provider | Discount | In | Out |
|---|---|---|---|
| Provider 1 | 10.0% off |
Explore the DeepSeek lab and compare related models.
The best available rate for deepseek-v4.1-flash-tee on SayGM right now is $0.27 per million input tokens and $1.08 per million output tokens.
SayGM's best available rate for deepseek-v4.1-flash-tee is 10.0% off list price.
Yes. deepseek-v4.1-flash-tee is served on SayGM's OpenAI-compatible chat/completions surface, so an unmodified OpenAI SDK pointed at SayGM's base URL works with it.
| $0.27 |
| $1.08 |
Measured on the requests SayGM routed to this model.
Fixed-seed suites SayGM runs against this model.
No benchmark run has been published for this model.
What is private AI inference? →