Price the API path vendor-exact against a self-hosted fleet per inference, solve the break-even volume, and check whether the claimed GPUs can physically serve the claimed traffic.
Self-hosting is 64.1% cheaper at this volume, a material saving that can justify owning the serving stack.
At 5,000,000 inferences/month, the API path costs $0.018 per inference ($90,000/mo, vendor-exact on claude-sonnet-4-6) against a self-hosted $0.00646 per inference ($32,300/mo including $25,000 of loaded MLOps). Verdict: self-host wins, saving $57,700/mo (64.1%). Break-even is around 1,794,444 inferences/month (0.36x today's volume). Fleet runs at 50% utilization; the self-host case is highly sensitive to that number and to whether the MLOps headcount is fully loaded. Ask for serving benchmarks and the real utilization curve before crediting any self-host saving in the model.
Values marked * are analyst estimates rather than vendor-verified data.