Talk to an AI cost expert

Get your AI & Cloud Costs under control today.

AI & Cloud Cost Relief: immediate steps you can take.At AICost.ai and our sister site CostOptimization.ai we provide the largest free toolset and knowledge base for resolving your cost issues: bill spiking & cost visibility, optimizing & reducing cost without losing quality, or planning a new AI workload accurately.

$299 Standard One hour, one workload
$99 Small business Same hour, same output

Most people who take this call go away and fix it themselves with the free tools on this site. That is the intended outcome, and that is fine with us.

What the hour actually contains

Five things, in this order

It is one hour, so it is structured. We are not going to spend it establishing rapport.

  1. We capture the AI workload that is actually hurting

    Not your whole estate. The one thing: the agent that tripled last month, the RAG bot nobody can price, the Bedrock bill that came in at twice the estimate, the GPU cluster that never scales to zero.

    • What it does, how it is called, and roughly how often
    • Which models, which cloud, which gateway (if any)
    • What you already tried, and why it did not hold
  2. We show you where MCP plugs into what you already run

    The 116 AI Cost decision engines are callable from Claude, ChatGPT or Cursor, or from your CI. Nothing to deploy, no cloud credentials, no data pipeline. We map them onto the three things you actually need:

    • See it - where the money goes, by team, feature and customer
    • Cut it - routing, caching, batching, agent loops, from tokens to RAG
    • Plan and forecast it - TCO and ROI before you build, caps that hold after
  3. Free

    We name the calculators that solve your case, and the order to run them

    There are 96 of them and that is genuinely too many to browse. The value here is the chain: which three or four, in which sequence, feeding which number into the next.

    • An agent problem is usually agent-loop-costretry-costagentic-envelope
    • A RAG bill is usually chunking-optimizerembedding-costvector-db-costrag-pipeline
    • A "which model" argument is usually multi-model-router with a quality floor, then prompt-cache-roi
  4. Cloud + tooling

    We look at the bill underneath the model bill

    Your model spend is rarely the whole story. The cloud beneath it - GPUs, Kubernetes, vector stores, egress, the observability you turned all the way up - is usually bigger.

    • CostOptimization.ai - 37 cloud cost categories, each with its modeled range, the specific spend it applies to, and the provider action
    • Its vendor-neutral tools catalog - we do not sell any of them, so the recommendation is just a recommendation
    • ToolsInfo.com for the wider category, including open source, when buying is not the answer
  5. You leave with a way forward to bring your AI costs under control and develop AI with confidence

    Written down, in your inbox, the same day. Which lever first, what it is worth, what it costs to do, and who has to do it.

    • Do it yourself - the calculators are free and the plan is yours. This is what most people do.
    • Do it with us - if it is bigger than an hour, we will say so and quote it. If it is not, we will say that too.

Being straight with you

Most customers take this call and then fix it themselves with the free tools. We are fine with that.

We would rather be the people who told you the truth in an hour than the people who sold you a six-week engagement to reach the same answer. The calculators stay free. The engines stay free to try. If the hour saves you a month of arguing internally, it paid for itself, and if you never speak to us again that is a perfectly good outcome.

Bring the workload that is bothering you

One hour. One workload. A written path forward the same day.

$299 Standard One hour, one workload
$99 Small business Same hour, same output

Not ready? Start with the 96 free calculators - no signup.
Managing AI spend for clients? The MSP program · Questions: [email protected]

📖 Data sources & methodology 163 text models · 9 embeddings · 37 vision · 55 audio · 8 vector DBs across 10 vendor pages · last verified 2026-07-28

Methodology

  • All prices are USD per 1 million tokens, current as of 2026-07-28.
  • Vendor-published values have no mark. Inferred/extrapolated values are marked with * and listed below.
  • Batch API discounts are 50% off standard rates across providers that offer Batch mode.
  • Prompt caching discounts vary by provider (typically 80-90% off cached input tokens).
  • Regional data-residency surcharges (Anthropic 1.1x, OpenAI 1.1x, Google regional tiers) are NOT included in base rates.
  • Long-context pricing tiers apply when input exceeds model threshold.
  • Embedding prices are input-only (no output tokens generated).

Primary sources

Last-verified date is the most recent successful daily snapshot (aicost_pricing_snapshots) or, when no snapshot exists yet, the latest successful crawler run (aicost_crawler_runs). 10 of 10 vendors are currently verified. Aggregator services (TokenCost, AI Pricing Guru, etc.) are not listed.

Anthropic
2026-07-28
https://www.anthropic.com/pricing
Daily snapshot since Sep 2023 · 631 days captured
Anthropic Docs
2026-07-28
https://platform.claude.com/docs/en/about-claude/pricing
Daily snapshot since Sep 2023 · 631 days captured
OpenAI
2026-07-28
https://openai.com/api/pricing/
Daily snapshot since Sep 2023 · 632 days captured
Google AI
2026-07-28
https://ai.google.dev/gemini-api/docs/pricing
Daily snapshot since Dec 2023 · 607 days captured
Google Vertex
2026-07-28
https://cloud.google.com/vertex-ai/generative-ai/pricing
Daily snapshot since Dec 2023 · 607 days captured
DeepSeek
2026-07-28
https://api-docs.deepseek.com/quick_start/pricing
Daily snapshot since May 2024 · 546 days captured
xAI
2026-07-28
https://x.ai/api
Daily snapshot since Nov 2024 · 464 days captured
Mistral
2026-07-28
https://mistral.ai/pricing
Daily snapshot since Dec 2023 · 605 days captured
Cohere
2026-07-28
https://cohere.com/pricing
Daily snapshot since Sep 2023 · 631 days captured

Inferred values (marked with * in calculator tables)

Derived from industry conventions, not directly published by the vendor. Typical conventions: cached input = 10% of base (90% off), Batch API = 50% of base (50% off).

Vendor / Model Field Why it’s inferred
Anthropic — Claude Sonnet 4.6 cachedInput Derived at 10% of input rate — Anthropic publishes 90% cache-hit discount on this tier.
Anthropic — Claude Sonnet 4.5 cachedInput Derived at 10% of input rate; same 90% cache-hit convention as Sonnet 4.6.
Anthropic — Claude Sonnet 4.5 batchInput Derived at 50% of standard input — Anthropic documents uniform 50% Batch discount.
Anthropic — Claude Sonnet 4.5 batchOutput Derived at 50% of standard output — Anthropic documents uniform 50% Batch discount.
Anthropic — Claude Haiku 4.5 cachedInput Derived at 10% of input rate — Anthropic 90% cache-hit discount convention.
OpenAI — GPT-5.4 Mini cachedInput Derived at 10% of input — OpenAI documents automatic 90% discount on cache hits across GPT-5.x tier.
OpenAI — GPT-5.4 Nano cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
OpenAI — GPT-5.4 Nano batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Nano batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Pro cachedInput Derived at 10% of input — OpenAI 90% cache-hit convention.
OpenAI — GPT-5.4 Pro batchInput Derived at 50% of input — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.4 Pro batchOutput Derived at 50% of output — OpenAI Batch API uniform 50% discount.
OpenAI — GPT-5.2 cachedInput Derived at 10% of input; no residency uplift.
OpenAI — GPT-5.2 batchInput Derived at 50% of input.
OpenAI — GPT-5.2 batchOutput Derived at 50% of output.
OpenAI — GPT-5 cachedInput Derived at 10% of input.
OpenAI — GPT-5 batchInput Derived at 50% of input.
OpenAI — GPT-5 batchOutput Derived at 50% of output.
OpenAI — GPT-5.5 Pro cachedInput Derived at 10% of input — OpenAI does not publish a cached rate for *-pro models; using the family convention.
OpenAI — GPT-5.5 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5.5 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5.2 Pro cachedInput Derived at 10% of input — pro-tier convention.
OpenAI — GPT-5.2 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5.2 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5.1 batchInput Derived at 50% of input.
OpenAI — GPT-5.1 batchOutput Derived at 50% of output.
OpenAI — GPT-5 Pro batchInput Derived at 50% of input.
OpenAI — GPT-5 Pro batchOutput Derived at 50% of output.
OpenAI — GPT-5 Nano cachedInput Derived at 10% of input.
OpenAI — GPT-5 Nano batchInput Derived at 50% of input.
OpenAI — GPT-5 Nano batchOutput Derived at 50% of output.
Google — Gemini 3 Flash cachedInput Derived at 10% of input — Google caching discount convention ~90%.
Google — Gemini 3.1 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 3.1 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 3.1 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.5 Pro cachedInput Derived at 10% of input.
Google — Gemini 2.5 Flash cachedInput Derived at 10% of input.
Google — Gemini 2.5 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 2.5 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.5 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash cachedInput Derived at 25% of input per Google 2.0 family caching rates.
Google — Gemini 2.0 Flash batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash-Lite cachedInput Derived at 10% of input — Google caching convention.
Google — Gemini 2.0 Flash-Lite batchInput Derived at 50% of input — Google Batch API uniform 50% discount.
Google — Gemini 2.0 Flash-Lite batchOutput Derived at 50% of output — Google Batch API uniform 50% discount.
xAI — Grok 4 (legacy) cachedInput Extrapolated at 25% of base.

Pricing is cross-verified against the LiteLLM community registry when available. Daily snapshots are kept in aicost_pricing_snapshots; every change is logged to aicost_price_changelog with old & new values for full audit trail. Read the full methodology →