Local vs Cloud Inference Cost

Pick an open-weight model and your monthly workload. Compare the three ways to serve it — serverless host, dedicated GPU rental, and self-hosting on your own cloud — and see which wins at your volume.

What this calculator does

Compares the three ways to serve an open-weight model - a serverless host (per token), a dedicated GPU rental (per hour), and self-hosting on your own cloud (AWS GPU per hour) - at your monthly output volume, then names the cheapest and shows where the crossovers fall.

Why use it
  • These are three different pricing models (per-token, per-GPU-hour rented, per-GPU-hour owned) and the cheapest one flips with volume - most teams never compare all three.
  • Cheap per-token models stay on serverless to very high volume; expensive ones cross to dedicated GPUs early. The verdict tells you which regime you are in.
Model family
500 = 500M output tok/mo
for serverless input cost
output tok/s per GPU
idle GPUs still bill
your own cloud box
on-demand / reserved / spot

Updates instantly. The cheapest of the three at your volume is ranked first.

Cheapest serving$74.83/mo · Serverless wins · GLM (Z.ai)
GLM (Z.ai) — cheapest way to serve 500M output tok/mo
#1
Serverless host cheapest
Novita
$74.83/mo
$0.112/1M
#2
Dedicated GPU rental
DeepInfra B200 180GB
$2,037/mo
$4.073/1M
#3
Self-host (your cloud)
AWS g5.48xlarge
$11,892/mo
$23.783/1M
At 2,500 tok/sec* and 45% utilization, one GPU/instance serves ~2,957M tokens/month, so your 500M-token workload needs 1 GPU for the rental and self-host legs. Lower utilization or throughput raises that count and the GPU-leg cost.
Serverless host
Novita · autoglm-phone-9b-multilingual · $0.112/1M blended ~
Dedicated GPU rental
DeepInfra B200 180GB · $2.79/hr × 730 × 1 GPU
Self-host (your cloud)
AWS g5.48xlarge · 8×A10G · onDemand $16.29/hr

Serverless and rental prices come from the curated hosting table (DeepInfra, Fireworks; verified 2026-06-28) and the LiteLLM nightly feed (~ Together / Novita / Groq). Self-host GPU $/hr comes from the resource-pricing source of truth. All three GPU legs assume full utilization at the stated throughput* and exclude setup, ops, networking, and reliability overhead, treat the verdict as a planning estimate, not a quote. Cheap per-token models favor serverless to much higher volumes; expensive ones cross over to dedicated GPUs sooner.

Go deeper

Our playbooks on cutting this number.

🖧
Hosting Cost Compare
Serverless host spread for this model
🖥️
Inference Serving Cost
Deep self-host serving economics
⚖️
Self-Host Breakeven
When owning GPUs pays off
⚖️
Jurisdiction Risk
Is this host OK for your data?

Need help using this calculator for your workloads?

AICost.ai has 50+ calculators and playbooks. Schedule an AvatarVA meeting and we'll work through your real cost scenarios across AI & Cloud: visibility, cost reduction, optimization, forecasting and capacity planning, without sacrificing accuracy or performance.

📅 Schedule an AvatarVA meeting →