Margin Calculator · for SaaS founders

Is your AI feature actually profitable?

See AI cost as % of revenue, gross margin after AI, and what cheaper models do to your bottom line.

Pricing verified: 2026-07-28 For SaaS / per-seat products
What this calculator does

Unit economics of your AI feature: revenue per request vs AI cost per request.

Why use it
  • "Our AI feature is losing money per user" — find out before accounting tells you

New to this calculator? Start with the ⚡ Playground — a few sliders, instant ballpark. Then switch to the 🧮 Calculator for your exact number.

Two ways to use this: visualize in the Playground, then get your number in the Calculator.

Margin Calculator Playground
Drag the sliders for an instant estimate.
What’s your gross margin per user?

Revenue per user minus blended inference + overhead. Power-user multiplier and overhead fixed at defaults.

Gross margin
Change the input sliders below to see new estimates.

How we got this estimate

💡Margin % = (revenue − blended cost) ÷ revenue. Power users cost more, dragging blended cost up.

Revenue minus AI cost is your real margin. Enter your pricing and usage below to see whether the feature actually makes money.

Margin Calculator Calculator

Enter your exact numbers for a precise result.
📊 Not sure of a value? Fields with a ▾ Typical pill offer broad industry ballparks (sourced typical ranges; AI product gross margins 2026: ICONIQ average 52% (up from 41% in 2024); LLM-native ~65% (Bessemer); classic SaaS ceiling 80-90%. An $80 seat with an AI assistant commonly carries ~$15 direct AI COGS.) so you can move forward now — your result gets more accurate as you replace them with your own measured numbers. Values marked * are rough estimates.
🎛 CALCULATOR
🏢 Your unit economics

Per-user, per-month. Honest numbers beat optimistic ones.

Hint: Paying users per month.
Hint: Average revenue per user per month, in dollars.
Hint: Non-AI cost per user (hosting, support). A rough estimate is fine.
🤖 Their AI usage

Fills the per-request usage below. Open ⚙ Advanced to fine-tune tokens, cache, and model.

Hint: AI calls the average user makes per day.
Hint: Words sent in per request.
Hint: Words sent back per request.
Hint: Which model you run. Drives the per-request cost.
Hint: How often the same setup text is reused. 30-50% typical; 90% off on cached parts.

Results

📈 RESULTS
- Calculating…
Gross margin after AI costs
-
-
AI cost - Other COGS - Gross margin -
AI cost / user / mo
-
AI cost / revenue
-
Monthly AI bill
-
Annual AI bill
-
Monthly gross profit
-
Break-even price/user
-
to cover AI + other COGS
💡 Recommendations
    📊 What if you switched models?

    Your margin at different model choices, same workload.

    Model AI cost / user % of revenue Gross margin Monthly AI bill
    Full cost calculator → What happens at 10x users? → Get a profitability audit →
    🎯 Use this result to
    • 💹 See your real margin — AI features look cheap until you do the math. See actual gross margin.
    • 🎯 Hit margin targets — Pricing decisions defended by math. Justify $X per user with unit economics.
    • 📊 AI as share of COGS — When AI becomes 30%+ of COGS, model swap or caching saves the business.
    • 🔌 Integrate with your AI agents — MCP available for agentic workflow integration. Margin-aware feature routing.
    📅 Schedule a call to apply this to your workload

    Go deeper

    Our playbooks on cutting this number.

    🧮
    AI Unit Economics
    The foundational topic for this tool
    💾
    Prompt Caching
    50-90% off - biggest margin lever
    💸
    Cheapest Model Finder
    Find cheaper alternatives
    📉
    Token Volatility
    How prices change over time

    Need help using this calculator for your workloads?

    AICost.ai has 50+ calculators and playbooks. Schedule an AvatarVA meeting and we'll work through your real cost scenarios across AI & Cloud: visibility, cost reduction, optimization, forecasting and capacity planning, without sacrificing accuracy or performance.

    📅 Schedule an AvatarVA meeting →
    📖 Data sources & methodology 163 text models · 9 embeddings · 37 vision · 55 audio · 8 vector DBs across 10 vendor pages · last verified 2026-07-28

    Methodology

    • All prices are USD per 1 million tokens, current as of 2026-07-28.
    • Vendor-published values have no mark. Inferred/extrapolated values are marked with * and listed below.
    • Batch API discounts are 50% off standard rates across providers that offer Batch mode.
    • Prompt caching discounts vary by provider (typically 80-90% off cached input tokens).
    • Regional data-residency surcharges (Anthropic 1.1x, OpenAI 1.1x, Google regional tiers) are NOT included in base rates.
    • Long-context pricing tiers apply when input exceeds model threshold.
    • Embedding prices are input-only (no output tokens generated).

    Primary sources

    Last-verified date is the most recent successful daily snapshot (aicost_pricing_snapshots) or, when no snapshot exists yet, the latest successful crawler run (aicost_crawler_runs). 10 of 10 vendors are currently verified. Aggregator services (TokenCost, AI Pricing Guru, etc.) are not listed.

    Anthropic
    2026-07-28
    https://www.anthropic.com/pricing
    Daily snapshot since Sep 2023 · 631 days captured
    Anthropic Docs
    2026-07-28
    https://platform.claude.com/docs/en/about-claude/pricing
    Daily snapshot since Sep 2023 · 631 days captured
    OpenAI
    2026-07-28
    https://openai.com/api/pricing/
    Daily snapshot since Sep 2023 · 632 days captured
    Google AI
    2026-07-28
    https://ai.google.dev/gemini-api/docs/pricing
    Daily snapshot since Dec 2023 · 607 days captured
    Google Vertex
    2026-07-28
    https://cloud.google.com/vertex-ai/generative-ai/pricing
    Daily snapshot since Dec 2023 · 607 days captured
    DeepSeek
    2026-07-28
    https://api-docs.deepseek.com/quick_start/pricing
    Daily snapshot since May 2024 · 546 days captured
    xAI
    2026-07-28
    https://x.ai/api
    Daily snapshot since Nov 2024 · 464 days captured
    Mistral
    2026-07-28
    https://mistral.ai/pricing
    Daily snapshot since Dec 2023 · 605 days captured
    Cohere
    2026-07-28
    https://cohere.com/pricing
    Daily snapshot since Sep 2023 · 631 days captured

    Inferred values (marked with * in calculator tables)

    Derived from industry conventions, not directly published by the vendor. Typical conventions: cached input = 10% of base (90% off), Batch API = 50% of base (50% off).

    Vendor / Model Field Why it’s inferred
    Anthropic — Claude Sonnet 4.6 cachedInput Derived at 10% of input rate — Anthropic publishes 90% cache-hit discount on this tier.
    Anthropic — Claude Sonnet 4.5 cachedInput Derived at 10% of input rate; same 90% cache-hit convention as Sonnet 4.6.
    Anthropic — Claude Sonnet 4.5 batchInput Derived at 50% of standard input — Anthropic documents uniform 50% Batch discount.
    Anthropic — Claude Sonnet 4.5 batchOutput Derived at 50% of standard output — Anthropic documents uniform 50% Batch discount.
    Anthropic — Claude Haiku 4.5 cachedInput Derived at 10% of input rate — Anthropic 90% cache-hit discount convention.
    OpenAI — GPT-5.4 Mini cachedInput Derived at 10% of input — OpenAI documents automatic 90% discount on cache hits across GPT-5.x tier.
    OpenAI — GPT-5.4 Nano cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
    OpenAI — GPT-5.4 Nano batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
    OpenAI — GPT-5.4 Nano batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
    OpenAI — GPT-5.4 Pro cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
    OpenAI — GPT-5.4 Pro batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
    OpenAI — GPT-5.4 Pro batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
    OpenAI — GPT-5.2 cachedInput Derived at 10% of input; no residency uplift.
    OpenAI — GPT-5.2 batchInput Derived at 50% of input.
    OpenAI — GPT-5.2 batchOutput Derived at 50% of output.
    OpenAI — GPT-5 cachedInput Derived at 10% of input.
    OpenAI — GPT-5 batchInput Derived at 50% of input.
    OpenAI — GPT-5 batchOutput Derived at 50% of output.
    OpenAI — GPT-5.5 Pro cachedInput Derived at 10% of input — OpenAI does not publish a cached rate for *-pro models; using the family convention.
    OpenAI — GPT-5.5 Pro batchInput Derived at 50% of input.
    OpenAI — GPT-5.5 Pro batchOutput Derived at 50% of output.
    OpenAI — GPT-5.2 Pro cachedInput Derived at 10% of input — pro-tier convention.
    OpenAI — GPT-5.2 Pro batchInput Derived at 50% of input.
    OpenAI — GPT-5.2 Pro batchOutput Derived at 50% of output.
    OpenAI — GPT-5.1 batchInput Derived at 50% of input.
    OpenAI — GPT-5.1 batchOutput Derived at 50% of output.
    OpenAI — GPT-5 Pro batchInput Derived at 50% of input.
    OpenAI — GPT-5 Pro batchOutput Derived at 50% of output.
    OpenAI — GPT-5 Nano cachedInput Derived at 10% of input.
    OpenAI — GPT-5 Nano batchInput Derived at 50% of input.
    OpenAI — GPT-5 Nano batchOutput Derived at 50% of output.
    Google — Gemini 3 Flash cachedInput Derived at 10% of input — Google caching discount convention ~90%.
    Google — Gemini 3.1 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
    Google — Gemini 3.1 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
    Google — Gemini 3.1 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
    Google — Gemini 2.5 Pro cachedInput Derived at 10% of input.
    Google — Gemini 2.5 Flash cachedInput Derived at 10% of input.
    Google — Gemini 2.5 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
    Google — Gemini 2.5 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
    Google — Gemini 2.5 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
    Google — Gemini 2.0 Flash cachedInput Derived at 25% of input per Google 2.0 family caching rates.
    Google — Gemini 2.0 Flash batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
    Google — Gemini 2.0 Flash batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
    Google — Gemini 2.0 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
    Google — Gemini 2.0 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
    Google — Gemini 2.0 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
    xAI — Grok 4 (legacy) cachedInput Extrapolated at 25% of base.

    Pricing is cross-verified against the LiteLLM community registry when available. Daily snapshots are kept in aicost_pricing_snapshots; every change is logged to aicost_price_changelog with old & new values for full audit trail. Read the full methodology →