Open-Weight Hosting Cost Compare

Pick an open-weight model family and your monthly workload. See what it costs across inference hosts, and the volume where renting a dedicated GPU beats serverless.

What this calculator does

Prices an open-weight model (GLM, DeepSeek, Kimi, Qwen, MiniMax, Llama, GPT-OSS, Nemotron) across inference hosts - DeepInfra, Fireworks, Together, Novita, Groq - at your monthly workload, then computes the volume where renting a dedicated GPU 24/7 beats serverless.

Why use it
  • The same open-weight model costs different amounts on different hosts - sometimes a 10x+ spread - and most teams never re-check after picking one.
  • Above a certain volume, a dedicated GPU is cheaper than per-token serverless; this finds that crossover so you know which regime you are in.
Model family
100 = 100M tokens/mo
rest are output
default ~500k/min (Fireworks)

Cost updates instantly. DeepInfra + Fireworks are exact curated prices; Together / Novita / Groq come from the nightly feed.

Cheapest host$11.23/mo · Novita cheapest for GLM (Z.ai)
GLM (Z.ai) — serverless cost / month 32.5× spread across hosts
HostCheapest model$/1M in$/1M out$ / month
Novita cheapest~ autoglm-phone-9b-multilingual $0.04 $0.14 $11.23
Together~ GLM-4.5-Air-FP8 $0.20 $1.10 $87.50
Fireworks GLM 5.2 $1.40 $4.40 $365.00
Hosts list different model variants in this family, so this is the cheapest option per host, not a strict same-model match. Check the model column.
Self-host crossover Serverless wins at your volume
Cheapest on-demand GPU
B200 180GB · DeepInfra
$2.79/hr → $2,037/mo (24/7)
Self-host $/1M*
$0.093
vs serverless $0.112/1M
Breakeven volume
~18,136M tok/mo
above this, a dedicated GPU is cheaper
Assumes 8,333 tokens/sec* on a saturated GPU (~499,980 tok/min). A single GPU at this throughput serves ~21,899M tokens/month if fully utilized. Throughput is the one assumption you control, edit it to match your model and hardware.

DeepInfra and Fireworks serverless prices and all GPU-hour rates are curated by aicost.ai* (verified 2026-06-28) from each provider's pricing page. Hosts marked ~ (Together, Novita, Groq) come from the LiteLLM nightly feed and may lag new releases. Self-host figures assume full GPU utilization at the stated throughput and exclude ops/setup overhead, treat the crossover as a planning estimate, not a quote. GPU rental ignores scale-to-zero idle savings and reserved-tier discounts.

Go deeper

Our playbooks on cutting this number.

🌏
Chinese vs Western Cost
Open-weight vs frontier price spread
⚖️
Jurisdiction Risk
Is this host acceptable for your data?
🖥️
Inference Serving Cost
Self-host serving math in detail
⚖️
Self-Host Breakeven
When owning GPUs pays off

Need help using this calculator for your workloads?

AICost.ai has 50+ calculators and playbooks. Schedule an AvatarVA meeting and we'll work through your real cost scenarios across AI & Cloud: visibility, cost reduction, optimization, forecasting and capacity planning, without sacrificing accuracy or performance.

📅 Schedule an AvatarVA meeting →