Find the cheapest AI model for your workload
Compare every major LLM side-by-side. Sorted by price. Filter by context, modality, or compliance.
Filter the full catalog of 150+ AI models by your role, workload, capability needs, and price ceiling — output is a ranked shortlist you can ship to other calculators.
- Stop reading vendor blog posts — every model is in one table, freshness-stamped
- Filter by what you actually need (vision, long-context, tier-1 only) instead of skimming marketing pages
- Price slider + min-context slider narrow 150+ models to the 5-10 worth comparing
- Compare checkbox lets you hold 2-4 candidates side-by-side before exporting to Cost Calculator
New to this calculator? Start with the ⚡ Playground — a few sliders, instant ballpark. Then switch to the 🧮 Calculator for your exact number.
Two ways to use this: visualize in the Playground, then get your number in the Calculator.
Quality and complexity choose the tier; volume scales it. One estimate from all three.
💡Complexity sets the tier (1-3 budget, 4-6 mid-tier, 7+ premium). Quality is a floor: 7+ rules out budget, 9+ forces premium. Monthly volume multiplies the tier rate.
- Your roleTunes filter defaults to your priorities
- Your workload scenarioPre-tunes capability filters
- Capability chipsAND-filters by feature
- Max input price per 1M tokensCuts everything above your ceiling
- Minimum context windowCuts small-window models
- Provider filtersVendor allowlist
- Ranked model listModels matching all filters
- Side-by-side compareCheck 2-4 to compare in detail
- Pricing freshnessDays since last verification
- Capability tagsWhat this model can do
Cut 150+ models to 5-10 candidates without reading any vendor blog post
New providers launch monthly — this surfaces them instead of you finding out on Twitter
Compare table is screenshot-ready for finance and security review
Export shortlist directly to Cost Calculator or Multi-Model Router for the actual decision
👇 Now try the calculator below with your own AI workloads
| Model | Provider | Input $/M ↓ | Output $/M | Cached $/M | Batch | Context | Modalities | Tags |
|---|
- Shortlist your exact-fit models — filter by role, workload, price, and context
- Compare 2–3 finalists — side-by-side price, context, and capabilities
- Price your real workload — hit "Calc →" on any row to drop it into the cost calculator
What this means + what to do next
- Real workload cost — sticker price tells you nothing about your token shape
- Quality at YOUR task — same model is great on one workload, weak on another
- Cache-aware effective pricing — caching changes the effective $/M dramatically for repeat-context workloads
- Vendor lock-in cost — switching prompts between vendors often takes weeks of eval
- Get exact $/month at your token shape per candidate model Cost Calculator
- Quantifies lock-in risk before you commit Vendor Concentration Risk
- Some models support prompt caching — drops effective $/M by 50-80% above ~22% hit rate Prompt Cache Roi
This is a discovery tool. ROI conversations happen downstream:
- Is the cheapest model in the shortlist quality-acceptable for my task? (Run eval.)
- Is the most-capable model in my shortlist 2× better, or 10× better? (The price gap usually reflects 2×, the value gap often reflects 1.3×.)
- How locked in am I to my current vendor — what would migration cost?
- Mixed workloads benefit from routing — different models for different query types Multi Model Router
- For some open-weight models, self-hosting becomes cheaper above a usage threshold Self Host Breakeven
- Translate $/request into $/customer or $/feature margin Margin Calculator
If you already have specific candidates in mind, skip discovery:
- You know which model you want; you just need the dollar number Cost Calculator
- You want the cheapest model meeting a quality floor for a specific workload Cheapest Model
- You've decided you want multiple models for different traffic patterns Multi Model Router
Need help using this calculator for your workloads?
AICost.ai has 50+ calculators and playbooks. Schedule an AvatarVA meeting and we'll work through your real cost scenarios across AI & Cloud: visibility, cost reduction, optimization, forecasting and capacity planning, without sacrificing accuracy or performance.
📅 Schedule an AvatarVA meeting →📖 Data sources & methodology
Methodology
- All prices are USD per 1 million tokens, current as of 2026-07-28.
- Vendor-published values have no mark. Inferred/extrapolated values are marked with * and listed below.
- Batch API discounts are 50% off standard rates across providers that offer Batch mode.
- Prompt caching discounts vary by provider (typically 80-90% off cached input tokens).
- Regional data-residency surcharges (Anthropic 1.1x, OpenAI 1.1x, Google regional tiers) are NOT included in base rates.
- Long-context pricing tiers apply when input exceeds model threshold.
- Embedding prices are input-only (no output tokens generated).
Primary sources
Last-verified date is the most recent successful daily snapshot
(aicost_pricing_snapshots) or, when no snapshot exists yet,
the latest successful crawler run (aicost_crawler_runs).
10 of 10
vendors are currently verified. Aggregator services (TokenCost, AI Pricing Guru, etc.)
are not listed.
Inferred values (marked with * in calculator tables)
Derived from industry conventions, not directly published by the vendor. Typical conventions: cached input = 10% of base (90% off), Batch API = 50% of base (50% off).
| Vendor / Model | Field | Why it’s inferred |
|---|---|---|
| Anthropic — Claude Sonnet 4.6 | cachedInput |
Derived at 10% of input rate — Anthropic publishes 90% cache-hit discount on this tier. |
| Anthropic — Claude Sonnet 4.5 | cachedInput |
Derived at 10% of input rate; same 90% cache-hit convention as Sonnet 4.6. |
| Anthropic — Claude Sonnet 4.5 | batchInput |
Derived at 50% of standard input — Anthropic documents uniform 50% Batch discount. |
| Anthropic — Claude Sonnet 4.5 | batchOutput |
Derived at 50% of standard output — Anthropic documents uniform 50% Batch discount. |
| Anthropic — Claude Haiku 4.5 | cachedInput |
Derived at 10% of input rate — Anthropic 90% cache-hit discount convention. |
| OpenAI — GPT-5.4 Mini | cachedInput |
Derived at 10% of input — OpenAI documents automatic 90% discount on cache hits across GPT-5.x tier. |
| OpenAI — GPT-5.4 Nano | cachedInput |
Derived at 10% of input — OpenAI 90% cache-hit convention. |
| OpenAI — GPT-5.4 Nano | batchInput |
Derived at 50% of input — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Nano | batchOutput |
Derived at 50% of output — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Pro | cachedInput |
Derived at 10% of input — OpenAI 90% cache-hit convention. |
| OpenAI — GPT-5.4 Pro | batchInput |
Derived at 50% of input — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Pro | batchOutput |
Derived at 50% of output — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.2 | cachedInput |
Derived at 10% of input; no residency uplift. |
| OpenAI — GPT-5.2 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.2 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 | cachedInput |
Derived at 10% of input. |
| OpenAI — GPT-5 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.5 Pro | cachedInput |
Derived at 10% of input — OpenAI does not publish a cached rate for *-pro models; using the family convention. |
| OpenAI — GPT-5.5 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.5 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.2 Pro | cachedInput |
Derived at 10% of input — pro-tier convention. |
| OpenAI — GPT-5.2 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.2 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.1 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.1 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 Nano | cachedInput |
Derived at 10% of input. |
| OpenAI — GPT-5 Nano | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 Nano | batchOutput |
Derived at 50% of output. |
| Google — Gemini 3 Flash | cachedInput |
Derived at 10% of input — Google caching discount convention ~90%. |
| Google — Gemini 3.1 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 3.1 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 3.1 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.5 Pro | cachedInput |
Derived at 10% of input. |
| Google — Gemini 2.5 Flash | cachedInput |
Derived at 10% of input. |
| Google — Gemini 2.5 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 2.5 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.5 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash | cachedInput |
Derived at 25% of input per Google 2.0 family caching rates. |
| Google — Gemini 2.0 Flash | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 2.0 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| xAI — Grok 4 (legacy) | cachedInput |
Extrapolated at 25% of base. |
Pricing is cross-verified against the
LiteLLM community registry
when available. Daily snapshots are kept in aicost_pricing_snapshots;
every change is logged to aicost_price_changelog with old & new
values for full audit trail. Read the full methodology →