{"ok":true,"engine":"aicost.gpu-inference-diligence","inputs":{"modelSlug":"claude-sonnet-4-6","inputTokensPerUnit":3000,"outputTokensPerUnit":600,"unitsPerMonth":5000000,"cachedInputPct":0,"batchPct":0,"gpuHourlyUsd":2.5,"gpuCount":4,"throughputTokensPerSec":1000,"gpuUtilizationPct":50,"mlOpsMonthlyUsd":25000},"result":{"ok":true,"verdict":"SELF-HOST WINS","verdictReason":"Self-hosting is 64.1% cheaper at this volume, a material saving that can justify owning the serving stack.","apiCostPerUnitUsd":0.018,"selfHostCostPerUnitUsd":0.00646,"perUnitDeltaUsd":0.01154,"apiMonthlyUsd":90000,"selfHostMonthlyUsd":32300,"gpuMonthlyUsd":7300,"mlOpsMonthlyUsd":25000,"monthlyDeltaUsd":57700,"savingPct":64.1,"breakEvenUnitsPerMonth":1794444,"breakEvenMultipleOfToday":0.36,"capacity":{"capacityTokensPerMonth":5256000000,"tokensNeededPerMonth":3000000000,"capacityRatio":0.57,"shortfall":false,"gpusNeeded":3,"utilizationRequiredPct":28.5},"model":{"slug":"claude-sonnet-4-6","inputPerM":3,"outputPerM":15},"monthlyUsd":32300,"cost":0.00646,"memo":"At 5,000,000 inferences/month, the API path costs $0.018 per inference ($90,000/mo, vendor-exact on claude-sonnet-4-6) against a self-hosted $0.00646 per inference ($32,300/mo including $25,000 of loaded MLOps). Verdict: self-host wins, saving $57,700/mo (64.1%). Break-even is around 1,794,444 inferences/month (0.36x today's volume). Fleet runs at 50% utilization; the self-host case is highly sensitive to that number and to whether the MLOps headcount is fully loaded. Ask for serving benchmarks and the real utilization curve before crediting any self-host saving in the model.","questions":["What is the measured sustained throughput (output tokens/sec per GPU) at production batch sizes, not the vendor benchmark?","What is real average GPU utilization over a month, including nights and weekends?","Which engineers are counted in the MLOps cost, and is that fully loaded?","What happens to the break-even if traffic drops 30%, and who absorbs the idle GPU cost?","Are reserved or committed GPU contracts in place, and what is the term and exit cost?","What is the plan when the frontier model gets cheaper again, and does the self-host case survive that?"],"assumptions":["API side priced vendor-exact on claude-sonnet-4-6 at 2026-07-20 list rates from the aicost pricing SSOT.","Self-host = $2.5/GPU-hour x 4 GPUs x 730h + $25,000 loaded MLOps.","Sustained throughput 1000 output tokens/sec/GPU at 50% utilization. *","Excludes model licensing, data-centre egress, and the cost of falling behind frontier model quality."],"sensitivities":[{"driver":"gpuUtilizationPct","note":"The single biggest lever. Idle GPUs are billed anyway, so low utilization destroys the self-host case."},{"driver":"mlOpsMonthlyUsd","note":"Usually understated. Fully loaded serving engineers often exceed the GPU bill itself."},{"driver":"unitsPerMonth","note":"Self-host is a fixed-cost bet. Volume below break-even makes it strictly worse than the API."}],"_model":"gpu-inference-diligence"}}