SwitchLLM

Build without limits.

Cut 90% of wasted tokens.

Get a free API key See pricing
No Chinese ID needed Live in a minute Pay as you go · no monthly fee
Measured, not claimed

Same prompt. A tenth of the cost.

Identical prompt, identical model. The only difference is switching off reasoning you never read.

Default call
811
output tokens
736
of which "thinking"
$0.0036
cost per call
→
With SwitchLLM
41
output tokens
0
"thinking" tokens
$0.0002
cost per call

Measured 2026-10-09 directly against the provider's official API. Model: glm-5.3.

⚡

Kill the thinking tax

By default these models spend 85–97% of their output tokens on hidden reasoning. One switch turns it off. Same answers, a fraction of the cost.

🔑

One key

Every model behind a single OpenAI-compatible endpoint. Switching models is changing one string — no new SDK, no new account.

🌏

Ready in a minute

Registering upstream requires a Chinese phone number and ID verification. We have done all of that. You need an email and one minute.

Pricing

Pay for answers, not for thinking

Top up once, spend across every model. No subscription, credit never expires.

Starter

$5
Pay as you go · minimum top-up
  • Every model
  • Reasoning switch included
  • OpenAI-compatible endpoint
  • Email support
Start free
Most popular

Pro

$25
Top-up · +5% bonus credit
  • Everything in Starter
  • +5% bonus credit
  • Higher rate limits
  • Priority email support
Top up $25

Scale

$100
Top-up · +12% bonus credit
  • Everything in Pro
  • +12% bonus credit
  • Highest rate limits
  • Direct technical contact
Top up $100
FAQ

Questions we get a lot

What exactly is "reasoning off"?
Modern models write a long private chain of thought before answering. You never see it, but you are billed for every token. Our switch tells the model to skip it, on the models that support it. For most tasks — classification, extraction, summarisation, chat, code — the answer is the same at a fraction of the cost. For genuinely hard reasoning, leave it on.
Do I need a Chinese phone number or ID?
No. Registering directly with the upstream providers requires Chinese identity verification. We have already done that. You only need an email address.
Is the API really OpenAI-compatible?
Yes. The endpoint is /v1/chat/completions and works with the official openai Python and Node SDKs, LangChain, LlamaIndex and anything else that speaks the OpenAI protocol. Just change base_url.
Which models are available?
We currently serve the flagship models from DeepSeek, Zhipu GLM and Moonshot Kimi, and we add more regularly. The live list in your console is always authoritative.
How do I pay?
Card payments are being onboarded now. Until then we accept USDT and can issue credit manually — contact [email protected] and we will set you up within a few hours.
What if a provider goes down?
We run multiple independent providers with automatic health checks. If one degrades, traffic moves to the others. Same key, same endpoint — nothing changes on your side.

Get your key in under a minute

No credit card needed to start.

Create free account