Pick an open-weight model family and your monthly workload. See what it costs across inference hosts, and the volume where renting a dedicated GPU beats serverless.
Prices an open-weight model (GLM, DeepSeek, Kimi, Qwen, MiniMax, Llama, GPT-OSS, Nemotron) across inference hosts - DeepInfra, Fireworks, Together, Novita, Groq - at your monthly workload, then computes the volume where renting a dedicated GPU 24/7 beats serverless.
| Host | Cheapest model | $/1M in | $/1M out | $ / month |
|---|---|---|---|---|
| Novita cheapest~ | autoglm-phone-9b-multilingual | $0.04 | $0.14 | $11.23 |
| Together~ | GLM-4.5-Air-FP8 | $0.20 | $1.10 | $87.50 |
| Fireworks | GLM 5.2 | $1.40 | $4.40 | $365.00 |
DeepInfra and Fireworks serverless prices and all GPU-hour rates are curated by aicost.ai* (verified 2026-06-28) from each provider's pricing page. Hosts marked ~ (Together, Novita, Groq) come from the LiteLLM nightly feed and may lag new releases. Self-host figures assume full GPU utilization at the stated throughput and exclude ops/setup overhead, treat the crossover as a planning estimate, not a quote. GPU rental ignores scale-to-zero idle savings and reserved-tier discounts.
AICost.ai has 50+ calculators and playbooks. Schedule an AvatarVA meeting and we'll work through your real cost scenarios across AI & Cloud: visibility, cost reduction, optimization, forecasting and capacity planning, without sacrificing accuracy or performance.
📅 Schedule an AvatarVA meeting →