Control runaway AI costs with AICost CostWall.
Plan your AI with confidence with AICost CostProof.

A firewall blocks the bad traffic and lets the good traffic through. AICost CostWall does that for AI spend.

Your teams keep shipping. The waste does not get through. Budgets, daily caps, spike detection and a kill switch, compiled into the gateway you already run. Governance, control and security sit on the money, and the AI innovation carries on.
AICost CostProof turns modeled AI savings into concrete evidence, so your next AI plan is built on what actually happened rather than what someone hoped.

Most popular

Solution Session

  • One hour on YOUR AI + cloud costs - bill spiking, cutting without losing accuracy, or planning a new workload
  • Same-day written report - the way forward in 30 days of concrete steps
  • You leave with a toolchain - free AICost calculators to run + MCP decision engines to deploy (or a third-party tool if it fits better)
  • Fee credits toward any AICost plan - you never pay twice for the same ground
  • An hour with the people who built the engines
  • Report the same day
  • The fee credits toward any AICost plan
  • Two slots a week
$299one hour
or $99 for small business
Book a Session → Concrete, expert, same day.
The $39 Blueprint credits toward it.

Every way in, in order

Start at $0. Pay only when your AI workload needs holding.

Most teams stop at the first card. Each rung credits into the next: the $39 Blueprint comes off a Session, the Session fee comes off any plan. You never pay twice for the same ground.

Most people stop here

Self-service

Everything you need to see it, cut it and plan it.

$0

No signup for the calculators.

  • 84 ad-hoc AI cost calculators
  • 98 AI cost decision engines via MCP
  • All 5 AICost product lines
  • CostOptimization.ai knowledge base, 37 cloud categories
  • ToolsInfo.com, incl. open source
  • 1 workspace, 1 workload, fair use
Start free Browse the knowledge base ↗
Before you pay a cent

MCP beta key

Run the AICost decision engines on your own AI workload.

$0

Two steps. Then decide.

  • 1. Request a key. AICost issues by hand
  • 2. Add mcp.aicost.ai to Claude, ChatGPT or Cursor
  • Nothing to deploy, no cloud credentials
  • The same engines the paid plans run
  • Not a demo, not a trial
Request a key See the integration →

Workload Blueprint

Your workload, priced and planned. Automated.

$39one-time

Credits in full toward a Session.

  • Short wizard, no call, no scheduling
  • Grounded 30-day plan as a PDF
  • Every number dated and sourced
  • Built on your workloads, not a template
Get the Blueprint
Founding price

AICost Control

CostWall stops it. CostProof proves it.

$99/month$199/mo

Founding price. Locked for as long as you stay.

  • AICost CostWall. Budgets, daily caps, spike anomaly, kill switch
  • AICost CostProof. Persistent ledger, modeled vs actual
  • Policy compiled for LiteLLM, Cloudflare, Portkey, Kong
  • 5 workloads
  • Cloud billing ingest
Start with Control
Founding price

AICost Scale

The cloud bill under the model bill, too.

$499/month$799/mo

Everything in Control, plus:

  • Cloud billing ingest: AWS CUR, Azure, BigQuery
  • Reconcile & attribution by feature, customer, team
  • CI/CD cost gate. Fail the build over budget
  • Unlimited workloads
Start with Scale

Agents that will not sit still. An AI bill nobody can attribute. A key that could end a quarter. AICost CostWall and AICost CostProof live in AICost Control and AICost Scale.

Talk to us

Bigger than a credit card

Three things AICost does not sell through a checkout page. Each one starts with a conversation about your AI estate rather than a plan picker.

MSP Partner Beta

$999/mo · 5 tenants · $99 each additional

  • Multi-tenant with hard isolation. Client A cannot see Client B
  • Your brand, your price. Set a markup % per tenant
  • Issue and revoke client keys yourself. Admin API, no tickets
  • Portfolio roll-up and a CostProof report per client
  • Limited founding partners. Co-delivery and feedback expected
Apply to the partner beta →

Enterprise

from $2,500/mo

  • SSO / SAML, RBAC
  • Private MCP endpoint
  • DPA, data retention terms, SLA
  • Security review and procurement support
  • Custom pricing overrides per tenant
Contact sales →

Implementation

from $5,000 one-off

  • Gateway integration and custom policies
  • Billing ingest wiring (CUR / Azure / BigQuery)
  • Private deployment
  • Not the normal path. Most teams are live in a day on their own
Talk it through →

The line

What you get for nothing, and what you pay for

Five AICost product lines are free to run on your own AI workloads. Two products are paid, because they hold your gateway in production and persist the evidence. All seven work across the same six AI cost pillars.

Model your AI costs for nothing. Pay when you want AICost to hold the brake in production and show the receipts afterwards.

FamilyFreeControlScale
AICost ClaritySo you understand why your bill is what it is. Bill diagnose, attribution, unit economics
AICost OptimizeSo you pay less, measurably. Routing, caching, batching, agent loops, RAG
AICost ForecastSo you know what you will pay before you ship. TCO, ROI, scale projection
AICost RiskSo you do not pay later for what goes wrong. Hallucination liability, compliance exposure, PII leakage
AICost LeverageSo you negotiate like the Fortune 500 does. Commitment discounts across Anthropic, OpenAI, Bedrock, Azure OpenAI, Vertex
AICost CostWallBlocks the waste, lets the work through. Budgets, caps, anomaly, kill switch, enforced in your gateway
AICost CostProofProve it. Persistent ledger, modeled vs actual, realized savings
Cloud billing ingestAWS CUR, Azure Cost Management, GCP BigQuery
WorkloadsConcurrent workloads under management 15Unlimited

Questions people actually ask

Why are the engines free? What is the catch?

There is no catch, and AICost is not going to insult you by saying “free forever”. Nothing is.

Here is what is actually true. The AI landscape moves every week. Teams are running AI workloads nobody has had to manage before: MLOps pipelines feeding models that retrain themselves, RAG over a corpus that changes daily, agent loops that call themselves and bill you for the privilege. AICost does not know that its AI cost decision engines work on your specific AI workload until you run them on it. So AICost would much rather you found that out for nothing.

MIT’s Media Lab studied 300 public enterprise AI deployments and found that 95% of GenAI pilots delivered no measurable P&L impact. The cause was not model quality. It was the learning gap: generic tools that never adapt to a real workflow. Tools brought in from outside succeeded about twice as often as internal builds. The GenAI Divide: State of AI in Business 2025. MIT Media Lab, NANDA initiative, July 2025.

AICost is not interested in being one of that 95%. So the order matters: your AI workload gets to production first. Once the AICost decision engines have it live and under control, you will want AICost CostWall holding your gateway and AICost CostProof persisting the evidence. That is when AICost makes money. If the free AI cost tools solve it and you never pay AICost a cent, that is still a good outcome. You won.

What is a “founding price”?

The real AICost price for the first customers. $99/mo instead of $199, $499/mo instead of $799. It is locked for as long as you stay subscribed. Not a trial, and not a discount that expires. When founding closes, new customers pay standard and you do not.

How does the credit actually work?

Buy an AICost Blueprint, the $39 comes off a Solution Session. Buy a Session, the fee comes back as a one-use code toward any AICost plan. AICost emails it to you and you enter it at checkout. Each rung credits into the next, so you never pay twice for the same ground.

Is AICost CostWall going to slow my teams down?

Actually the opposite. AICost wants your teams as productive as they can possibly be: the right model for the right capability, without paying a surcharge for the privilege.

Runaway AI spend is not only a finance problem. It is the thing that gets an AI project cancelled before it ever reaches production. Every dollar burned on a retry loop nobody bounded, or on a frontier model doing work a small one does just as well, is a dollar not spent on the part of your AI that is actually innovative. Teams that lose control of the AI bill do not just overspend. They fail to make it to market.

There is a plainer reason to trust this. AICost does not earn a cent until your AI workload is live and under control. AICost CostWall and AICost CostProof only sell to teams who got to production. Your project shipping is the AICost business model, so AICost is aligned with you rather than with your invoice.

Mechanically: AICost CostWall caps what a workload can burn in a day, catches a spike the moment it starts, and gives you a kill switch. It does not queue your requests, sit in your hot path, or ask an engineer for permission. Under the cap nobody notices it is there. Over the cap you find out in seconds rather than on the invoice.

Will AICost route my traffic to a cheaper model behind my back?

No. The AICost router takes a quality floor and a latency ceiling that you set. A cheaper model that fails your floor is never selected. AICost simply does not claim the saving. Nothing is auto-applied either: AICost emits a policy, a human reviews it, your gateway runs it. AICost never touches your AI traffic.

Do you need our cloud credentials?

No. The AICost decision engines model from facts you supply. No cloud credentials, no API keys, no customer data. On AICost Scale you can send a billing export if you want reconciliation. That is a file you choose to send, not access you grant.

What is a workload?

One AI workload with its own budget: a support bot, an agent pipeline, a summarization endpoint. Not one model, and not one API key. It is the unit you would want AICost CostWall to cap independently.

Can I cancel?

Any time, from your account. No call, no retention script. Your AICost CostProof history stays exportable for 30 days after.

All prices USD. How the integration works · For MSPs · Try 84 calculators free · [email protected]

📖 Data sources & methodology 158 text models · 9 embeddings · 37 vision · 55 audio · 8 vector DBs across 10 vendor pages · last verified 2026-07-21

Methodology

  • All prices are USD per 1 million tokens, current as of 2026-07-21.
  • Vendor-published values have no mark. Inferred/extrapolated values are marked with * and listed below.
  • Batch API discounts are 50% off standard rates across providers that offer Batch mode.
  • Prompt caching discounts vary by provider (typically 80-90% off cached input tokens).
  • Regional data-residency surcharges (Anthropic 1.1x, OpenAI 1.1x, Google regional tiers) are NOT included in base rates.
  • Long-context pricing tiers apply when input exceeds model threshold.
  • Embedding prices are input-only (no output tokens generated).

Primary sources

Last-verified date is the most recent successful daily snapshot (aicost_pricing_snapshots) or, when no snapshot exists yet, the latest successful crawler run (aicost_crawler_runs). 10 of 10 vendors are currently verified. Aggregator services (TokenCost, AI Pricing Guru, etc.) are not listed.

Anthropic
2026-07-21
https://www.anthropic.com/pricing
Daily snapshot since Sep 2023 · 624 days captured
Anthropic Docs
2026-07-21
https://platform.claude.com/docs/en/about-claude/pricing
Daily snapshot since Sep 2023 · 624 days captured
OpenAI
2026-07-21
https://openai.com/api/pricing/
Daily snapshot since Sep 2023 · 625 days captured
Google AI
2026-07-21
https://ai.google.dev/gemini-api/docs/pricing
Daily snapshot since Dec 2023 · 600 days captured
Google Vertex
2026-07-21
https://cloud.google.com/vertex-ai/generative-ai/pricing
Daily snapshot since Dec 2023 · 600 days captured
DeepSeek
2026-07-21
https://api-docs.deepseek.com/quick_start/pricing
Daily snapshot since May 2024 · 539 days captured
xAI
2026-07-21
https://x.ai/api
Daily snapshot since Nov 2024 · 457 days captured
Mistral
2026-07-21
https://mistral.ai/pricing
Daily snapshot since Dec 2023 · 598 days captured
Cohere
2026-07-21
https://cohere.com/pricing
Daily snapshot since Sep 2023 · 624 days captured

Inferred values (marked with * in calculator tables)

Derived from industry conventions, not directly published by the vendor. Typical conventions: cached input = 10% of base (90% off), Batch API = 50% of base (50% off).

Vendor / Model Field Why it’s inferred
Anthropic — Claude Sonnet 4.6 cachedInput Derived at 10% of input rate — Anthropic publishes 90% cache-hit discount on this tier.
Anthropic — Claude Sonnet 4.5 cachedInput Derived at 10% of input rate; same 90% cache-hit convention as Sonnet 4.6.
Anthropic — Claude Sonnet 4.5 batchInput Derived at 50% of standard input — Anthropic documents uniform 50% Batch discount.
Anthropic — Claude Sonnet 4.5 batchOutput Derived at 50% of standard output — Anthropic documents uniform 50% Batch discount.
Anthropic — Claude Haiku 4.5 cachedInput Derived at 10% of input rate — Anthropic 90% cache-hit discount convention.
OpenAI — GPT-5.4 Mini cachedInput Derived at 10% of input — OpenAI documents automatic 90% discount on cache hits across GPT-5.x tier.
OpenAI — GPT-5.4 Nano cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
OpenAI — GPT-5.4 Nano batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Nano batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Pro cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
OpenAI — GPT-5.4 Pro batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Pro batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.2 cachedInput Derived at 10% of input; no residency uplift.
OpenAI — GPT-5.2 batchInput Derived at 50% of input.
OpenAI — GPT-5.2 batchOutput Derived at 50% of output.
OpenAI — GPT-5 cachedInput Derived at 10% of input.
OpenAI — GPT-5 batchInput Derived at 50% of input.
OpenAI — GPT-5 batchOutput Derived at 50% of output.
OpenAI — GPT-5.5 Pro cachedInput Derived at 10% of input — OpenAI does not publish a cached rate for *-pro models; using the family convention.
OpenAI — GPT-5.5 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5.5 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5.2 Pro cachedInput Derived at 10% of input — pro-tier convention.
OpenAI — GPT-5.2 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5.2 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5.1 batchInput Derived at 50% of input.
OpenAI — GPT-5.1 batchOutput Derived at 50% of output.
OpenAI — GPT-5 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5 Nano cachedInput Derived at 10% of input.
OpenAI — GPT-5 Nano batchInput Derived at 50% of input.
OpenAI — GPT-5 Nano batchOutput Derived at 50% of output.
Google — Gemini 3 Flash cachedInput Derived at 10% of input — Google caching discount convention ~90%.
Google — Gemini 3.1 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 3.1 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 3.1 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.5 Pro cachedInput Derived at 10% of input.
Google — Gemini 2.5 Flash cachedInput Derived at 10% of input.
Google — Gemini 2.5 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 2.5 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.5 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash cachedInput Derived at 25% of input per Google 2.0 family caching rates.
Google — Gemini 2.0 Flash batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 2.0 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
xAI — Grok 4 (legacy) cachedInput Extrapolated at 25% of base.

Pricing is cross-verified against the LiteLLM community registry when available. Daily snapshots are kept in aicost_pricing_snapshots; every change is logged to aicost_price_changelog with old & new values for full audit trail. Read the full methodology →