//01 SayGM vs direct

SayGM or direct from
the model maker?

SayGM serves Claude, GPT and Gemini through each maker’s own API shape, priced at or below the maker’s list price, on one key and one prepaid balance. Buying direct gives you batch discounts and terms in your own name. Pick a maker to see live rates model by model.

[03] At a glance

Real-time on SayGM, batch direct.

Model makerSayGM vs list, typicalBuying direct adds
Anthropic12% below50% off batch jobs, zero data retention for eligible models, HIPAA BAAAs of September 2026
OpenAI18% below50% off Batch and Flex, zero data retention by approval, regional processingAs of September 2026
Google15% below50% off batch jobs, Flex inference, zero data retention for eligible featuresAs of September 2026

The typical saving is the middle model on each maker’s page, at the best current rate a provider offers. You pay the rate of the provider that serves each request, never above list.

[04] Questions

Common questions about switching.

How can SayGM price at or below buying direct?

Independent providers compete to serve each request, and the winning bid sets your rate. Every token line is capped at the model maker’s own list price, so the price you pay is at or below what the maker charges for the same model.

When is buying direct the better choice?

For work that can wait: as of September 2026, batch APIs at Anthropic, OpenAI and Google are 50% off list. Also when you need data terms, compliance reports or an enterprise contract in your own name. Each maker page lists the specifics.

Do I need to change my code?

For standard requests, only the base URL and the key. SayGM serves the Anthropic Messages, OpenAI Chat Completions and Responses, and Gemini generateContent APIs, so each maker’s own SDK keeps working, and one key covers all three. Batch jobs stay with the maker, as do OpenAI’s background Responses and hosted tools; each maker page lists the specifics.

Who can read my prompt?

The SayGM gateway runs inside an Intel TDX enclave, sealed from its operators and the host machine, with an attestation anyone can verify. The model provider serving the request still receives the prompt and processes it under its API terms. Optional guardrails can redact personal data and secrets inside the enclave first.