The same deterministic engines behind the web calculators run as a Model Context Protocol server. Ask cost questions inside Claude Code, Codex, ChatGPT, Cursor or Perplexity; wire them into CI, Streamlit apps and nightly jobs over REST. Six enterprise jobs below - each one something you can run today.
Each rail is a real workflow: the pain, the exact chain of engines and platform capabilities, the outcome you walk away with, and the one-liner to run it.
One invoice line per vendor; no idea which team, product, or agent spent it.
Spend attributed to workload and team, leakage and concentration flagged, and the cost of keeping that attribution running priced honestly.
In Claude Code
Routing 45% plus caching 30% is not 75%. Vendor decks sum levers that overlap.
A stacked-savings number that survives finance review, the eval budget that proves quality held, and an executable gateway config - not a slide.
Every engine returns
Teams that forecast cloud within 1-3% miss AI by 2-3x. Agentic spend compounds.
A range forecast with a breach month, a 20-40% reserve sized to your agentic mix, and alert thresholds instead of hard stops.
REST, from any script
A budget in a spreadsheet stops nothing. One runaway agent loop can burn a five-figure day.
Four enforcement tiers proven on live traffic, cost regressions caught pre-merge, and alerts before the bill becomes a surprise.
What a blocked call returns
Forecasts nobody reconciles are theater. Finance wants actuals, attribution, and an export their tooling reads.
Showback by tenant and workload, forecast accuracy you can quote, an open-standard export, and an audit trail where a silently altered row is a detected security event - not a shrug.
The buyer's uptime question
Every consultancy sells a point-in-time assessment. It is stale the week a model, a price, or a rule changes.
A governance posture that re-computes itself: same engines the assessment used, callable nightly from your own automation, with tamper-evident history. The PDF becomes a snapshot of a living system.
Deterministic by design
Every consultancy can hand you a governance assessment. Ours is different because the assessment IS the engine chain above: the day a model, a price, a route certification or a rule changes, you re-run it - from your own automation, on your schedule - and the hash-chained audit trail proves nothing was quietly edited in between. That is the continuous-governance claim behind AICostAdvisory.com, delivered as infrastructure instead of a deliverable.
Every audience runs the same engines in a different order. The full chain catalog powers this page, the web calculators, and the MCP server from one source of truth.
The engines model and decide for free. When the workload is live, two of the rails above become products: the wall that enforces, and the ledger that proves. Names and claims mirror the pricing page - one source for every price.
Blocks the waste, lets the work through. Budgets, daily caps, spike anomaly, kill switch - enforced in your gateway.
Runs on: Enforce it dynamically - caps that actually fire See plans →Prove it. Persistent ledger, modeled vs actual, realized savings.
Runs on: Audit it in the ledger - showback that holds up See plans →The cloud bill under the model bill, too. AWS CUR, Azure and BigQuery ingest, attribution by feature, customer and team, and a CI/CD cost gate that fails the build over budget.
Runs on: See where every AI dollar goes See plans →
We are the vendor-neutral decision layer that runs BEFORE the routing layer: decide which models are permitted,
what they cost, and what the caps are - then export that as configuration into whichever gateway your team already operates.
Policy engines emit an executable policy_patch;
anything a target cannot express natively is preserved under x_aicost_policy
so nothing is silently dropped. Actuals flow back through the LiteLLM connector, the billing file-inbox, or your FOCUS 1.x export.
Comparing them first? The router-compare engine holds the verified capability and pricing capture.
policy_patch compiles to model_list, fallbacks, budgets, cache and retry config. The Ledger pulls spend back from its database, so forecast-vs-actual closes automatically.
policy_patch compiles to gateway config: caching, rate limits, fallback order. Their edge guardrails and DLP run in front; our caps and routing policy ride in the same config.
Policy export to gateway config with guardrails at every tier on their side; our eligibility gates decide WHICH models the config may contain.
Policy export for teams already standardized on Kong: budgets and caps expressed as gateway plugins alongside existing API governance.
We run BEFORE it: eligibility and budget policy decide the permitted model set; OpenRouter moves the traffic. Billing exports drop into the file-inbox for Ledger reconcile.
Routing rules and budgets configured from our policy output; pass-through billing exports feed the Ledger the same way.
Their traces answer WHAT happened per request; our envelope answers WHAT IT SHOULD COST and why. Cost exports reconcile into the Ledger.
A routing brain that sits on top of an existing gateway. We qualify the model set it is allowed to route across before it ever picks.
Dual auth on one header: OAuth for humans in assistants, bearer keys for apps and pipelines. Nothing is degraded by using a key.
Get a key, connect an assistant, or book a session where we chain the right engines for your stack and hand you the gateway config to deploy.
Read the MCP docs → Book a Solution Session