Guides → Playground & Guide → Image Generation Cost - What You Actually Pay Per Kept Image
Meet Diego Alvarez. Founder shipping an AI product-photo feature for an e-commerce store. "I see $0.04/image on the pricing page - but I regenerate 3 times to get one good shot. What's my real monthly bill?"
🔥 Catalog has 40K SKUs, each needs a few variants. Investor wants a COGS number this week.
The sticker price lies, in two directions. First, the model spread is enormous - the same 1024x1024 image is $0.003 on Flux Schnell and $0.167 on GPT Image 1 at high quality, a roughly 50x range. Second, nobody keeps the first generation. A 2-3x regeneration rate means you pay for two or three throwaways for every image you actually ship.
Diego generating 40K kept product images/month at 2.5 generations each = 100K generations. On GPT Image 1 (medium) that's ~$4,200/mo; on Imagen 4 Fast ~$2,000/mo; on Flux Schnell ~$300/mo. The quality gap is real but for clean product shots on a white background it's far smaller than the 14x price gap - so the model choice is the whole ballgame.
Two more levers move the bill. (1) Quality tier - on token-billed models like GPT Image, High can cost ~4x Low; use Standard for production and reserve High for hero shots. (2) Batch API - where the provider offers it (OpenAI, Google), asynchronous batch generation runs ~50% cheaper, which is free money for any non-real-time pipeline.
Here are the inputs that move the result the most. Play with the sliders and check it out. The number updates live.
Estimate the monthly cost of generating images — including the regenerations you throw away to get a keeper. Prices live from the LiteLLM catalog.
💡Two levers dominate: the model (≈30× spread from Flux Schnell to GPT Image HD) and revisions — every regeneration to land one good image multiplies the bill.
Same calculator, three team sizes. Click a tab to see how the numbers shift.
Clean white-background product shots. Quality gap vs premium is tiny for this task, so the cheapest fast model wins. 40K keepers at 2.5x = 100K generations, ~$300/mo on Flux Schnell.
Healthy range: <$0.02/kept image
Campaign images with embedded text and brand polish. GPT Image High earns its premium here. Low volume, high revisions (4x) because the bar is high. ~$1.3K/mo for 2K keepers.
Healthy range: $1K-1.5K/mo (premium justified for hero work)
Large catalog, non-real-time. Batch API halves the bill. 200K keepers at 2x = 400K generations; batch turns ~$16K/mo into ~$8K/mo.
Healthy range: Batch −50% pays for itself instantly
Same calculator, different applications. Sizes above, workloads here. Pick the one that looks like yours.
Pre-loaded scenarios for the most common applications. Click a tab to see realistic numbers, then hit "Try this scenario" to load it into the calculator above.
One-shot avatar per user, low revision rate. Flux Dev balances cost and quality. ~$3.7K/mo at 100K.
Healthy range: <$5K/mo at 100K/mo
High-churn social creative, lots of variants. Imagen 4 Fast keeps unit cost low while staying usable. 3x revisions reflect creative iteration.
Healthy range: $4K-5K/mo
At 500K+ keepers, hosted per-image fees dwarf GPU rental. Use this to find the volume where self-hosting SD pays off (model GPU separately in self-host-breakeven).
Healthy range: API cost flags when self-host wins
Open the full calculator. Pick a model, enter your tokens, see per-call, daily, monthly, and annual cost.
🚀 Open the full calculator →Image-gen pricing is a 30x spread and the sticker price lies. Real monthly math across GPT Image, DALL-E, Imagen, Flux and Stable Diffusion - including the regenerations you throw away.
Each input shapes the cost. Click an input on the calculator to set it. The explanations below match the live calculator field by field.
| Input | Default | Typical ballparks |
|---|---|---|
modelKey
|
gpt-image-1 | — |
quality
|
medium | — |
imagesPerMonth
moves the needle
|
50,000 | — |
revisionsPerImage
moves the needle
|
2.5 | — |
batchApi
|
false | — |
Ballparks are broad industry starting points (sourced ranges; * = rough estimate) — your result gets more accurate as you replace them with measured numbers.
Try them live in the calculator;
API & agent users get the same data from the MCP resource aicost://input-reference/image-generation-cost.
Cost per kept image is the unit that matters. Not the sticker $/image - the all-in number after revisions. Under $0.05/kept image is healthy for production; above $0.15 you're either over-tiered or regenerating too much.
Revision waste is the hidden line item. If 60% of your spend is on generations you discard, the fix is a better prompt or a reference-image workflow, not a cheaper model.
The model spread is bigger here than anywhere in text. Up to ~30-50x between the cheapest and priciest at usable quality. Shop aggressively and test on your actual images.
Honest limitations. Every model is wrong; some are useful. Where this one falls short:
For these, use: Vision Cost for images IN (analysis/OCR). Self-Host Breakeven for the GPU-vs-API crossover at high volume.
Cost isn't the only dimension. Click any constraint to see how recommendations change.
The model is the dominant cost lever - up to 50x spread. Pick the cheapest model that clears your quality bar on YOUR images, then optimize revisions and batch.
DALL-E, Imagen, Flux, and Stable Diffusion grant commercial use; some models and hosts differ. Confirm licensing before shipping generated images in a product.
If users upload reference images for img2img, treat them as sensitive data and use a no-train / enterprise tier.
Image generation is slow relative to text (2-15s). For non-interactive pipelines use batch and pocket the discount; for interactive UX, show progress and consider a faster model.
Generation APIs differ in shape but a routing layer (LiteLLM, or an aggregator) makes multi-model switching cheap - useful given how fast prices move.
Instrument the revision rate - it's the single most actionable cost metric. Caching identical prompts and seeds avoids paying twice for the same image.
Tradeoff analysis is where most AI projects go sideways. Talk to a CFO-grade AI cost analyst →
The gaps we just listed are real, and they are the expensive ones: your actual prompts, your switching cost, your MLOps overhead. An AICost expert spends the hour on your AI and cloud costs, not a generic playbook. You leave with a written report: the way forward, in 30 days of concrete steps.
Not sure yet? The $39 AICost Blueprint credits toward a Session, and the Session fee credits toward any plan. You never pay twice for the same ground. See all pricing →
You have seen the shape of it. Open the calculator with your model, your tokens, your volume.
🚀 Open the full calculator →Author: Subu Vdaygiri, Founder & CEO of CloudIntelligence.ai. 17 years Fortune 100 (Ingram Micro, Siemens). Wharton CTO program · Kellogg CPO program · 10× AWS+Azure certified.
Last-verified date is the most recent successful daily snapshot
(aicost_pricing_snapshots) or, when no snapshot exists yet,
the latest successful crawler run (aicost_crawler_runs).
10 of 10
vendors are currently verified. Aggregator services (TokenCost, AI Pricing Guru, etc.)
are not listed.
Derived from industry conventions, not directly published by the vendor. Typical conventions: cached input = 10% of base (90% off), Batch API = 50% of base (50% off).
| Vendor / Model | Field | Why it’s inferred |
|---|---|---|
| Anthropic — Claude Sonnet 4.6 | cachedInput |
Derived at 10% of input rate — Anthropic publishes 90% cache-hit discount on this tier. |
| Anthropic — Claude Sonnet 4.5 | cachedInput |
Derived at 10% of input rate; same 90% cache-hit convention as Sonnet 4.6. |
| Anthropic — Claude Sonnet 4.5 | batchInput |
Derived at 50% of standard input — Anthropic documents uniform 50% Batch discount. |
| Anthropic — Claude Sonnet 4.5 | batchOutput |
Derived at 50% of standard output — Anthropic documents uniform 50% Batch discount. |
| Anthropic — Claude Haiku 4.5 | cachedInput |
Derived at 10% of input rate — Anthropic 90% cache-hit discount convention. |
| OpenAI — GPT-5.4 Mini | cachedInput |
Derived at 10% of input — OpenAI documents automatic 90% discount on cache hits across GPT-5.x tier. |
| OpenAI — GPT-5.4 Nano | cachedInput |
Derived at 10% of input — OpenAI 90% cache-hit convention. |
| OpenAI — GPT-5.4 Nano | batchInput |
Derived at 50% of input — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Nano | batchOutput |
Derived at 50% of output — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Pro | cachedInput |
Derived at 10% of input — OpenAI 90% cache-hit convention. |
| OpenAI — GPT-5.4 Pro | batchInput |
Derived at 50% of input — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.4 Pro | batchOutput |
Derived at 50% of output — OpenAI Batch API uniform 50% discount. |
| OpenAI — GPT-5.2 | cachedInput |
Derived at 10% of input; no residency uplift. |
| OpenAI — GPT-5.2 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.2 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 | cachedInput |
Derived at 10% of input. |
| OpenAI — GPT-5 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.5 Pro | cachedInput |
Derived at 10% of input — OpenAI does not publish a cached rate for *-pro models; using the family convention. |
| OpenAI — GPT-5.5 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.5 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.2 Pro | cachedInput |
Derived at 10% of input — pro-tier convention. |
| OpenAI — GPT-5.2 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.2 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5.1 | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5.1 | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 Pro | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 Pro | batchOutput |
Derived at 50% of output. |
| OpenAI — GPT-5 Nano | cachedInput |
Derived at 10% of input. |
| OpenAI — GPT-5 Nano | batchInput |
Derived at 50% of input. |
| OpenAI — GPT-5 Nano | batchOutput |
Derived at 50% of output. |
| Google — Gemini 3 Flash | cachedInput |
Derived at 10% of input — Google caching discount convention ~90%. |
| Google — Gemini 3.1 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 3.1 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 3.1 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.5 Pro | cachedInput |
Derived at 10% of input. |
| Google — Gemini 2.5 Flash | cachedInput |
Derived at 10% of input. |
| Google — Gemini 2.5 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 2.5 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.5 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash | cachedInput |
Derived at 25% of input per Google 2.0 family caching rates. |
| Google — Gemini 2.0 Flash | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash-Lite | cachedInput |
Derived at 10% of input — Google caching convention. |
| Google — Gemini 2.0 Flash-Lite | batchInput |
Derived at 50% of input — Google Batch API uniform 50% discount. |
| Google — Gemini 2.0 Flash-Lite | batchOutput |
Derived at 50% of output — Google Batch API uniform 50% discount. |
| xAI — Grok 4 (legacy) | cachedInput |
Extrapolated at 25% of base. |
Pricing is cross-verified against the
LiteLLM community registry
when available. Daily snapshots are kept in aicost_pricing_snapshots;
every change is logged to aicost_price_changelog with old & new
values for full audit trail. Read the full methodology →