Guides → Playground & Guide → AI Cost Diligence for VC/PE — Stress-Test a Startup's Unit Economics

AI Cost Diligence for VC/PE — Stress-Test a Startup's Unit Economics

Meet Deal team analyst. VC/PE associate running technical + financial diligence on an AI startup. "The founder claims $0.02/query COGS and 83% gross margin — does that survive vendor-exact pricing, and where does the margin trend?"

🔥 Unit economics is the 2026 gatekeeper; subsidised inference is often disguised as growth spend, and inconsistent deck-vs-data-room numbers stall term sheets.

The story

Second and third partner meetings drill into gross margin and inference cost, not vision. This tool prices the target's stated workload at vendor-exact list rates from the aicost pricing SSOT and compares it to the claim.

It returns a PLAUSIBLE / OPTIMISTIC / IMPLAUSIBLE verdict, the real gross margin, and — because investors price the trajectory, not the level — a 12-24 month margin path using the documented inference-cost decline plus the startup's own caching and routing levers.

The output includes an auto-written deal-memo paragraph and the exact founder questions to ask, downloadable as a PDF for the data room.

🎮 Playground

AI Cost Diligence for VC/PE Playground

Here are the inputs that move the result the most. Play with the sliders and check it out. The number updates live.

Does the startup's AI cost claim survive?

Stress-test a claimed COGS against vendor-exact pricing. Verdict, real margin vs the 2026 AI median, trajectory, and a downloadable deal-memo.

Vendor-exact floor

💡The vendor-exact floor is set by model + tokens; the gap and real margin follow. Investors price the trajectory, not the level.

Ready to run the numbers?

Open the full calculator. Pick a model, enter your tokens, see per-call, daily, monthly, and annual cost.

🚀 Open the full calculator →

Top 3 right now

Verified 23 hours ago

About this calculator: AI Cost Diligence for VC/PE — Stress-Test a Startup's Unit Economics

Stress-test a startup's claimed AI COGS against vendor-exact pricing. Verdict, real gross margin vs the 2026 AI median, 12-24mo margin trajectory, model-routing lever, and a downloadable deal-memo.

🎛 Inputs you control

Each input shapes the cost. Click an input on the calculator to set it. The explanations below match the live calculator field by field.

Claimed COGS / unit: The founder's stated cost per unit (per query/task/etc).
How to choose: Take it from the deck or data room; this is the number under test.
Claimed model: The model the startup says it runs on.
How to choose: Verify against actual billing - founders quote a cheap model but often ship on a pricier tier.
Input tokens / unit: Estimated input tokens per unit.
How to choose: The most-gamed number; a '2K token' claim that is really 8K flips most verdicts. Ask for a traced sample.
Output tokens / unit: Estimated output tokens per unit.
How to choose: Output is priced 4-6x input; small under-statements matter.
Price / unit: Revenue per unit (what they charge).
How to choose: Enables the real gross-margin and trajectory analysis.
Volume / month: Monthly unit volume.
How to choose: Drives the annualized COGS understatement figure.
Claimed cache-hit %: Share of input tokens the founder claims are cache-served.
How to choose: Require production evidence of a durable cache-hit rate, not a target.
Claimed batch %: Share of volume via batch API.
How to choose: Batch trades latency for ~50% price; not all workloads tolerate it.
Inference decline %/yr: Assumed per-year inference cost decline for the trajectory.
How to choose: 2026 trend is ~5-10x per 12-18mo (aggressive); 40-50%/yr is a conservative default.
Most popular

Everything above is the 80% case. The last 20% is where the money is.

The gaps we just listed are real, and they are the expensive ones: your actual prompts, your switching cost, your MLOps overhead. An AICost expert spends the hour on your AI and cloud costs, not a generic playbook. You leave with a written report: the way forward, in 30 days of concrete steps.

  • An hour with the people who built the engines
  • Report the same day
  • The fee credits toward any AICost plan
  • Two slots a week
Book a Solution Session: $299 → or $99 for small business →

Not sure yet? The $39 AICost Blueprint credits toward a Session, and the Session fee credits toward any plan. You never pay twice for the same ground. See all pricing →

Ready to run your own numbers?

You have seen the shape of it. Open the calculator with your model, your tokens, your volume.

🚀 Open the full calculator →

Methodology

Editorial gate
8-layer defense, see aicost.ai/ai-cost-economics
Last verified
7/21/2026, 8:00:00 PM

Author: Subu Vdaygiri, Founder & CEO of CloudIntelligence.ai. 17 years Fortune 100 (Ingram Micro, Siemens). Wharton CTO program · Kellogg CPO program · 10× AWS+Azure certified.

📖 Data sources & methodology 158 text models · 9 embeddings · 37 vision · 55 audio · 8 vector DBs across 10 vendor pages · last verified 2026-07-22

Methodology

  • All prices are USD per 1 million tokens, current as of 2026-07-22.
  • Vendor-published values have no mark. Inferred/extrapolated values are marked with * and listed below.
  • Batch API discounts are 50% off standard rates across providers that offer Batch mode.
  • Prompt caching discounts vary by provider (typically 80-90% off cached input tokens).
  • Regional data-residency surcharges (Anthropic 1.1x, OpenAI 1.1x, Google regional tiers) are NOT included in base rates.
  • Long-context pricing tiers apply when input exceeds model threshold.
  • Embedding prices are input-only (no output tokens generated).

Primary sources

Last-verified date is the most recent successful daily snapshot (aicost_pricing_snapshots) or, when no snapshot exists yet, the latest successful crawler run (aicost_crawler_runs). 10 of 10 vendors are currently verified. Aggregator services (TokenCost, AI Pricing Guru, etc.) are not listed.

Anthropic
2026-07-22
https://www.anthropic.com/pricing
Daily snapshot since Sep 2023 · 625 days captured
Anthropic Docs
2026-07-22
https://platform.claude.com/docs/en/about-claude/pricing
Daily snapshot since Sep 2023 · 625 days captured
OpenAI
2026-07-22
https://openai.com/api/pricing/
Daily snapshot since Sep 2023 · 626 days captured
Google AI
2026-07-22
https://ai.google.dev/gemini-api/docs/pricing
Daily snapshot since Dec 2023 · 601 days captured
Google Vertex
2026-07-22
https://cloud.google.com/vertex-ai/generative-ai/pricing
Daily snapshot since Dec 2023 · 601 days captured
DeepSeek
2026-07-22
https://api-docs.deepseek.com/quick_start/pricing
Daily snapshot since May 2024 · 540 days captured
xAI
2026-07-22
https://x.ai/api
Daily snapshot since Nov 2024 · 458 days captured
Mistral
2026-07-22
https://mistral.ai/pricing
Daily snapshot since Dec 2023 · 599 days captured
Cohere
2026-07-22
https://cohere.com/pricing
Daily snapshot since Sep 2023 · 625 days captured

Inferred values (marked with * in calculator tables)

Derived from industry conventions, not directly published by the vendor. Typical conventions: cached input = 10% of base (90% off), Batch API = 50% of base (50% off).

Vendor / Model Field Why it’s inferred
Anthropic — Claude Sonnet 4.6 cachedInput Derived at 10% of input rate — Anthropic publishes 90% cache-hit discount on this tier.
Anthropic — Claude Sonnet 4.5 cachedInput Derived at 10% of input rate; same 90% cache-hit convention as Sonnet 4.6.
Anthropic — Claude Sonnet 4.5 batchInput Derived at 50% of standard input — Anthropic documents uniform 50% Batch discount.
Anthropic — Claude Sonnet 4.5 batchOutput Derived at 50% of standard output — Anthropic documents uniform 50% Batch discount.
Anthropic — Claude Haiku 4.5 cachedInput Derived at 10% of input rate — Anthropic 90% cache-hit discount convention.
OpenAI — GPT-5.4 Mini cachedInput Derived at 10% of input — OpenAI documents automatic 90% discount on cache hits across GPT-5.x tier.
OpenAI — GPT-5.4 Nano cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
OpenAI — GPT-5.4 Nano batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Nano batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Pro cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
OpenAI — GPT-5.4 Pro batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Pro batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.2 cachedInput Derived at 10% of input; no residency uplift.
OpenAI — GPT-5.2 batchInput Derived at 50% of input.
OpenAI — GPT-5.2 batchOutput Derived at 50% of output.
OpenAI — GPT-5 cachedInput Derived at 10% of input.
OpenAI — GPT-5 batchInput Derived at 50% of input.
OpenAI — GPT-5 batchOutput Derived at 50% of output.
OpenAI — GPT-5.5 Pro cachedInput Derived at 10% of input — OpenAI does not publish a cached rate for *-pro models; using the family convention.
OpenAI — GPT-5.5 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5.5 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5.2 Pro cachedInput Derived at 10% of input — pro-tier convention.
OpenAI — GPT-5.2 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5.2 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5.1 batchInput Derived at 50% of input.
OpenAI — GPT-5.1 batchOutput Derived at 50% of output.
OpenAI — GPT-5 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5 Nano cachedInput Derived at 10% of input.
OpenAI — GPT-5 Nano batchInput Derived at 50% of input.
OpenAI — GPT-5 Nano batchOutput Derived at 50% of output.
Google — Gemini 3 Flash cachedInput Derived at 10% of input — Google caching discount convention ~90%.
Google — Gemini 3.1 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 3.1 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 3.1 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.5 Pro cachedInput Derived at 10% of input.
Google — Gemini 2.5 Flash cachedInput Derived at 10% of input.
Google — Gemini 2.5 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 2.5 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.5 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash cachedInput Derived at 25% of input per Google 2.0 family caching rates.
Google — Gemini 2.0 Flash batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 2.0 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
xAI — Grok 4 (legacy) cachedInput Extrapolated at 25% of base.

Pricing is cross-verified against the LiteLLM community registry when available. Daily snapshots are kept in aicost_pricing_snapshots; every change is logged to aicost_price_changelog with old & new values for full audit trail. Read the full methodology →