Guides → Playground & Guide → Cheapest Model - Best Value for Your Workload
Meet Marcus Lee. Senior Engineer told to 'use the cheap model' for a new feature. "Cheapest model is meaningless without context. Cheapest for WHAT?"
🔥 Switched to Haiku to save money. Quality dropped. Switched back. CFO unhappy.
'Cheapest model' is the wrong question - 'cheapest model that hits my quality bar' is the right one. Gemini 3 Flash at $0.50/1M output is cheap. So is DeepSeek V3 at $0.27. Both 'fail' on certain tasks where Claude Haiku or GPT-5 Mini succeed. Cheapest only matters if quality clears the threshold.
Marcus's mistake: switched to the absolute cheapest tier without testing on his workload. Customer support classification - Haiku worked, saved 60%. But for the agentic workflow with tool calls, Haiku struggled with the schema and returned malformed JSON. Quality cost outweighed price savings.
Three tiers. (1) Budget (DeepSeek, Gemini Flash, Haiku) — fine for classification, simple Q&A, narrow extraction. (2) Mid-tier (GPT-5 Mini, Sonnet 3.5) — solid for general agentic work, RAG, structured outputs. (3) Premium (Sonnet 4.6, GPT-5.5) — needed for complex reasoning, math, code generation. Pick the cheapest tier that passes your eval, not the cheapest model overall.
Here are the inputs that move the result the most. Play with the sliders and check it out. The number updates live.
Two of these pick the model; one sets the scale. The estimate below uses all three together.
💡Complexity sets the tier — 1-3 budget, 4-6 mid-tier, 7+ premium (~13× pricier per token). Required quality only raises that floor: 7+ rules out budget, 9+ forces premium. Tasks/day then multiplies the tier rate by your volume.
Same calculator, three team sizes. Click a tab to see how the numbers shift.
Sentiment classification on user reviews. Low complexity, low quality bar. Budget tier wins by 80% over mid-tier. ~$300/mo at 100K/day.
Healthy range: DeepSeek V3 / Gemini Flash, ~$300/mo
Marcus's mid-tier workload. Quality 7 = budget not enough. Mid-tier tier hits the sweet spot. Haiku 4.5 ~$1.5K/mo, GPT-5 Mini similar. 4× cheaper than Sonnet at acceptable quality.
Healthy range: Haiku 4.5 or GPT-5 Mini, ~$1.5K/mo
Same calculator, different applications. Sizes above, workloads here. Pick the one that looks like yours.
Pre-loaded scenarios for the most common applications. Click a tab to see realistic numbers, then hit "Try this scenario" to load it into the calculator above.
High-volume tool selection (which API to call?). Mid-tier with prompt caching cuts cost dramatically. Tool defs cache → 80% input discount.
Healthy range: Haiku + caching, ~$2.5K/mo
Autocomplete-style code suggestions. High volume, modest complexity. Mid-tier fine. Fine-tuned cheap (Mistral/Llama) often wins at this scale.
Healthy range: GPT-5 Mini or fine-tuned Haiku
Marketing copy, blog posts, ad variants. Output quality matters for brand voice. Mid-tier tier. Sonnet 4.6 at this complexity beats both budget (quality) and Opus (cost).
Healthy range: Sonnet 4.6 sweet spot
Long-form research synthesis. Budget tiers fail. Premium tier mandatory. Volume is small, so absolute cost is fine.
Healthy range: Opus 4.7 / GPT-5.5 Pro mandatory
Open the full calculator. Pick a model, enter your tokens, see per-call, daily, monthly, and annual cost.
🚀 Open the full calculator →Cheapest LLM for your workload - by tier, by task type, by quality threshold. Updated daily as vendors shift pricing. Beyond the per-token table.
Each input shapes the cost. Click an input on the calculator to set it. The explanations below match the live calculator field by field.
| Input | Default | Typical ballparks |
|---|---|---|
tasksPerDay
moves the needle
|
50,000 | Pilot team · ~100/day = 100 · Small production · ~1K/day = 1,000 · Mid production · ~10K/day = 10,000 |
qualityThreshold
moves the needle
|
6 | — |
complexityScore
moves the needle
|
5 | — |
inputTokensPerTask
|
2,000 | — |
outputTokensPerTask
|
400 | — |
Ballparks are broad industry starting points (sourced ranges; * = rough estimate) — your result gets more accurate as you replace them with measured numbers.
Try them live in the calculator;
API & agent users get the same data from the MCP resource aicost://input-reference/cheapest-model.
What you'll see after the calculator runs. Each card explains how to read the number.
Complexity picks the tier. 1-3 → a budget model is plenty. 4-6 → mid-tier. 7+ → premium.
Quality is a floor, not a dial. Required quality 7+ rules out the budget tier; 9+ forces premium — whatever the complexity. It can only raise the tier, never lower it.
Volume amplifies savings - and risks. At 50K/day, picking 30% cheaper-per-task = $X/mo saved. Picking 5% lower-quality = customer complaints. Test before scaling.
Honest limitations. Every model is wrong; some are useful. Where this one falls short:
For these, use: Multi-Model Router for routing strategy. Cost Calculator for full bill.
Cost isn't the only dimension. Click any constraint to see how recommendations change.
Cost ranks change weekly. Anchor on tier, not vendor. Re-check pricing quarterly because rankings shift as vendors compete.
Hallucination rate scales inversely with model capability - usually. Test on your domain because some cheap models punch above weight on specific tasks.
Compliance certifications often skip budget tiers. Check before assuming the cheapest model has the same BAA/SOC 2 as the flagship.
Some cheap-tier APIs default to data-used-for-training. Read terms before piping user content.
Bonus: budget tier usually wins on latency too. Mid-tier beats premium for sub-300ms requirements.
The cheapest model 6 months ago isn't the cheapest now. Build vendor abstraction; don't hardcode the bargain.
Switching to a cheaper model without an eval is gambling. Build the eval first, switch second.
Tradeoff analysis is where most AI projects go sideways. Talk to a CFO-grade AI cost analyst →
The gaps we just listed are real, and they are the expensive ones: your actual prompts, your switching cost, your MLOps overhead. An AICost expert spends the hour on your AI and cloud costs, not a generic playbook. You leave with a written report: the way forward, in 30 days of concrete steps.
Not sure yet? The $39 AICost Blueprint credits toward a Session, and the Session fee credits toward any plan. You never pay twice for the same ground. See all pricing →
You have seen the shape of it. Open the calculator with your model, your tokens, your volume.
🚀 Open the full calculator →Author: Subu Vdaygiri, Founder & CEO of CloudIntelligence.ai. 17 years Fortune 100 (Ingram Micro, Siemens). Wharton CTO program · Kellogg CPO program · 10× AWS+Azure certified.
Why this matters: pricing for major vendors has dropped 40-90% in the last 24 months. A budget set 12 months ago is probably wrong by 30%+.
View 3-year history for →
Last-verified date is the most recent successful daily snapshot
(aicost_pricing_snapshots) or, when no snapshot exists yet,
the latest successful crawler run (aicost_crawler_runs).
10 of 10
vendors are currently verified. Aggregator services (TokenCost, AI Pricing Guru, etc.)
are not listed.
Derived from industry conventions, not directly published by the vendor. Typical conventions: cached input = 10% of base (90% off), Batch API = 50% of base (50% off).
| Vendor / Model | Field | Why it’s inferred |
|---|---|---|
| Anthropic — Claude Sonnet 4.6 | cachedInput |
Derived at 10% of input rate — Anthropic publishes 90% cache-hit discount on this tier. |
| Anthropic — Claude Sonnet 4.5 | cachedInput |
Derived at 10% of input rate; same 90% cache-hit convention as Sonnet 4.6. |
| Anthropic — Claude Sonnet 4.5 | batchInput |
Derived at 50% of standard input — Anthropic documents uniform 50% Batch discount. |
| Anthropic — Claude Sonnet 4.5 | batchOutput |
Derived at 50% of standard output — Anthropic documents uniform 50% Batch discount. |
| Anthropic — Claude Haiku 4.5 | cachedInput |
Derived at 10% of input rate — Anthropic 90% cache-hit discount convention. |
| OpenAI — GPT-5.4 Mini | cachedInput |
Derived at 10% of input — OpenAI documents automatic 90% discount on cache hits across GPT-5.x tier. |
| OpenAI — GPT-5.4 Nano | cachedInput |
Derived at 10% of input — OpenAI 90% cache-hit convention. |
| OpenAI — GPT-5.4 Nano | batchInput |
Derived at 50% of input — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Nano | batchOutput |
Derived at 50% of output — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Pro | cachedInput |
Derived at 10% of input — OpenAI 90% cache-hit convention. |
| OpenAI — GPT-5.4 Pro | batchInput |
Derived at 50% of input — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Pro | batchOutput |
Derived at 50% of output — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.2 | cachedInput |
Derived at 10% of input; no residency uplift. |
| OpenAI — GPT-5.2 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.2 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 | cachedInput |
Derived at 10% of input. |
| OpenAI — GPT-5 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.5 Pro | cachedInput |
Derived at 10% of input — OpenAI does not publish a cached rate for *-pro models; using the family convention. |
| OpenAI — GPT-5.5 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.5 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.2 Pro | cachedInput |
Derived at 10% of input — pro-tier convention. |
| OpenAI — GPT-5.2 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.2 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.1 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.1 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 Nano | cachedInput |
Derived at 10% of input. |
| OpenAI — GPT-5 Nano | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 Nano | batchOutput |
Derived at 50% of output. |
| Google — Gemini 3 Flash | cachedInput |
Derived at 10% of input — Google caching discount convention ~90%. |
| Google — Gemini 3.1 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 3.1 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 3.1 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.5 Pro | cachedInput |
Derived at 10% of input. |
| Google — Gemini 2.5 Flash | cachedInput |
Derived at 10% of input. |
| Google — Gemini 2.5 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 2.5 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.5 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash | cachedInput |
Derived at 25% of input per Google 2.0 family caching rates. |
| Google — Gemini 2.0 Flash | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 2.0 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| xAI — Grok 4 (legacy) | cachedInput |
Extrapolated at 25% of base. |
Pricing is cross-verified against the
LiteLLM community registry
when available. Daily snapshots are kept in aicost_pricing_snapshots;
every change is logged to aicost_price_changelog with old & new
values for full audit trail. Read the full methodology →