aicost.ai
VC/PE Diligence · AICost.ai cost decision engine

🖥️ Does the self-host plan actually pay back?

Price the API path vendor-exact against a self-hosted fleet per inference, solve the break-even volume, and check whether the claimed GPUs can physically serve the claimed traffic.

Inputs

Enter what the company claims. Everything recomputes live.

Priced vendor-exact from the pricing SSOT.
Drives API cost only.
Drives both API cost and GPU capacity.
Monthly served volume.
Rental or amortized owned rate per GPU.
$
Serving fleet size claimed.
Sustained production throughput, not the vendor benchmark.
Average over a month including nights and weekends, not peak.
%
Fully loaded cost of the people keeping the serving stack alive.
$
Optional.
%
Optional.
%
Verdict
SELF-HOST WINS

Self-hosting is 64.1% cheaper at this volume, a material saving that can justify owning the serving stack.

API cost / inference
$0.018
Self-host cost / inference
$0.00646
API monthly
$90,000
Self-host monthly
$32,300
Monthly delta
$57,700
Saving
64.1%
Break-even volume
1,794,444 inf/mo
Break-even vs today
0.36x
Fleet capacity ratio
0.57x
GPUs actually needed
3

Deal memo

At 5,000,000 inferences/month, the API path costs $0.018 per inference ($90,000/mo, vendor-exact on claude-sonnet-4-6) against a self-hosted $0.00646 per inference ($32,300/mo including $25,000 of loaded MLOps). Verdict: self-host wins, saving $57,700/mo (64.1%). Break-even is around 1,794,444 inferences/month (0.36x today's volume). Fleet runs at 50% utilization; the self-host case is highly sensitive to that number and to whether the MLOps headcount is fully loaded. Ask for serving benchmarks and the real utilization curve before crediting any self-host saving in the model.

Questions for the founder

  1. What is the measured sustained throughput (output tokens/sec per GPU) at production batch sizes, not the vendor benchmark?
  2. What is real average GPU utilization over a month, including nights and weekends?
  3. Which engineers are counted in the MLOps cost, and is that fully loaded?
  4. What happens to the break-even if traffic drops 30%, and who absorbs the idle GPU cost?
  5. Are reserved or committed GPU contracts in place, and what is the term and exit cost?
  6. What is the plan when the frontier model gets cheaper again, and does the self-host case survive that?

Assumptions

  • API side priced vendor-exact on claude-sonnet-4-6 at 2026-07-20 list rates from the aicost pricing SSOT.
  • Self-host = $2.5/GPU-hour x 4 GPUs x 730h + $25,000 loaded MLOps.
  • Sustained throughput 1000 output tokens/sec/GPU at 50% utilization. *
  • Excludes model licensing, data-centre egress, and the cost of falling behind frontier model quality.

Values marked * are analyst estimates rather than vendor-verified data.