Every model call is measured as it happens and carries who asked, for which workspace, project and use case, under which key. Budgets are checked before the call, not discovered afterwards. And every saving is proved by comparing the same workload before and after.
Every call is written to a ledger as it happens, with its tokens, its cost and what it was for — including AI coding and agents built elsewhere.
Each line belongs to a unit because the call carried its workspace, team, project or key. Never because a ratio said so.
A position per scope: limit, burn rate over complete days, projection — and whether that budget can actually fire.
Your real traffic priced on other models, recommended only where measured evidence says quality holds.
Before and after on the same workload, split into how much work was done and what it cost per token.
Most reported AI savings are a quieter month. We split every change into the part explained by doing more or less work and the part explained by what that work costs per token — and only the second one is ever called a saving. The agent that recommends a change is not the agent that grades it.
Same workload within tolerance, and the unit price fell. Measured on both sides, from the ledger.
The same work now costs more per token. Reported as plainly as a win.
The volume moved: this is a different month, not a cheaper one. No saving is claimed.
Nothing metered before the change. The period cannot be audited, and we say so.
The team starts as a measurement function. It earns the right to act on the record of what its recommendations were worth.
Allocation, budget positions, burn rates and arbitrage shortlists. Read-only from day one, and that alone usually finds the first surprise.
Model and caching changes with the evidence attached, and the ceilings the team would set. A person applies them.
Once tightenings have been upheld without reversal, the operator may lower one ceiling on a defined trigger and report at once — within a floor it cannot cross.
Raising a budget, switching a model, touching a key. These change what an agent is, not what it may spend.
Nothing is autonomous by default, and the first tightening that stops legitimate work sends the control straight back to recommend-only.
Priced per band of governed AI spend, plus a share of the savings the before/after audit proves — never a share of an estimate. Carbon is not measured per workload anywhere on this platform, so no statement from this team carries a carbon figure; for what the runtime does to cut token consumption, see cost & carbon. Your own platform and finance team runs it.