aicost.ai
VC/PE Diligence Memo
2026-07-22
Pricing as of 2026-07-20
aicost.gpu-inference-diligence
GPU Cost-per-Inference & Self-Host Break-even
Verdict
SELF-HOST WINS
Self-hosting is 64.1% cheaper at this volume, a material saving that can justify owning the serving stack.
Key figures
| API cost / inference |
$0.018 |
| Self-host cost / inference |
$0.00646 |
| API monthly |
$90,000 |
| Self-host monthly |
$32,300 |
| Monthly delta |
$57,700 |
| Saving |
64.1% |
| Break-even volume |
1,794,444 inf/mo |
| Break-even vs today |
0.36x |
| Fleet capacity ratio |
0.57x |
| GPUs actually needed |
3 |
Assessment
At 5,000,000 inferences/month, the API path costs $0.018 per inference ($90,000/mo, vendor-exact on claude-sonnet-4-6) against a self-hosted $0.00646 per inference ($32,300/mo including $25,000 of loaded MLOps). Verdict: self-host wins, saving $57,700/mo (64.1%). Break-even is around 1,794,444 inferences/month (0.36x today's volume). Fleet runs at 50% utilization; the self-host case is highly sensitive to that number and to whether the MLOps headcount is fully loaded. Ask for serving benchmarks and the real utilization curve before crediting any self-host saving in the model.
Questions for the founder
- What is the measured sustained throughput (output tokens/sec per GPU) at production batch sizes, not the vendor benchmark?
- What is real average GPU utilization over a month, including nights and weekends?
- Which engineers are counted in the MLOps cost, and is that fully loaded?
- What happens to the break-even if traffic drops 30%, and who absorbs the idle GPU cost?
- Are reserved or committed GPU contracts in place, and what is the term and exit cost?
- What is the plan when the frontier model gets cheaper again, and does the self-host case survive that?
Assumptions & method
- API side priced vendor-exact on claude-sonnet-4-6 at 2026-07-20 list rates from the aicost pricing SSOT.
- Self-host = $2.5/GPU-hour x 4 GPUs x 730h + $25,000 loaded MLOps.
- Sustained throughput 1000 output tokens/sec/GPU at 50% utilization. *
- Excludes model licensing, data-centre egress, and the cost of falling behind frontier model quality.
What moves the answer
- gpuUtilizationPct: The single biggest lever. Idle GPUs are billed anyway, so low utilization destroys the self-host case.
- mlOpsMonthlyUsd: Usually understated. Fully loaded serving engineers often exceed the GPU bill itself.
- unitsPerMonth: Self-host is a fixed-cost bet. Volume below break-even makes it strictly worse than the API.