Direct vendor APIs + AICost
Anthropic, OpenAI, Google, Mistral and more: the richest optimization surface, with real-time usage and daily reconciliation. Your model calls never route through us.
Your app→Vendor API→Vendor invoice
AICost: out-of-band
Richest optimization surface: caching, batch and routing engines apply at list prices
What you do (and what you never have to change)
- Nothing changes in your calls. Engines model caching ROI, batch eligibility and routing savings directly against verified list pricing.
- When you add a gateway (LiteLLM, Portkey, Helicone... 7 products verified), routing rules and caps export into it as policy_patch; the LiteLLM connector pulls actuals back two-way.
- No gateway? Vendor usage exports go to the actuals inbox monthly, same reconciliation.
- Product teams call the engines with scoped HMAC client tokens - they never touch your vendor keys.
DAILY costs - you never wait for the invoice: Two real-time options: emit the usage block your app already receives on every response into the modeled lane, or pull the vendor usage API daily into the inbox. Either way you never wait for the invoice.
And the same engines that price it recommend the fix: cheaper eligible models, caching and batch wins, plan-vs-API breakpoints - refreshed as your usage lands, so optimization is a daily habit, not a quarterly project.
What you'll see on the showback dashboard
$ / dayby modelby workloadby agentcache + batch savings realizedper-key attribution
Per-agent attributionPass metadata / user on each request, or issue one API key per agent - vendor usage APIs group by key. In the modeled lane, tags ride along on every ledger row.
Actuals pathGateway logs (two-way via LiteLLM connector) or vendor usage export -> file-inbox
What we never ask you to do
Never in your inference path. Latency, availability and data flow are untouched.
Never holding your keys. Actuals arrive as telemetry and exports you already generate.
Never a migration. The decision layer meets your estate where it is.