Offer · Govern Early access

Cost tools read the invoice after the month.
We meter the call before it is made.

Every model call is measured as it happens and carries who asked, for which workspace, project and use case, under which key. Budgets are checked before the call, not discovered afterwards. And every saving is proved by comparing the same workload before and after.

Five steps, every month

1

Meter

Every call is written to a ledger as it happens, with its tokens, its cost and what it was for — including AI coding and agents built elsewhere.

2

Attribute

Each line belongs to a unit because the call carried its workspace, team, project or key. Never because a ratio said so.

3

Budget

A position per scope: limit, burn rate over complete days, projection — and whether that budget can actually fire.

4

Arbitrate

Your real traffic priced on other models, recommended only where measured evidence says quality holds.

5

Prove

Before and after on the same workload, split into how much work was done and what it cost per token.

A saving is unit price on the same workload

Most reported AI savings are a quieter month. We split every change into the part explained by doing more or less work and the part explained by what that work costs per token — and only the second one is ever called a saving. The agent that recommends a change is not the agent that grades it.

Audited saving

Same workload within tolerance, and the unit price fell. Measured on both sides, from the ledger.

Audited increase

The same work now costs more per token. Reported as plainly as a win.

Not comparable

The volume moved: this is a different month, not a cheaper one. No saving is claimed.

No baseline

Nothing metered before the change. The period cannot be audited, and we say so.

What the agents may not do when they start

Never, at entry

  • Raise a budget, widen one, or turn one off
  • Switch an agent's model to chase a saving
  • Rotate, scope or disable one of your provider keys
  • Pause or stop an agent — that belongs to fleet governance
  • Spread unallocated spend across the units that can be traced
  • Claim a saving nobody measured on both sides

Always

  • One agent, and only one, can lower a ceiling — never below that scope's own median day
  • Every tightening carries a written reason and reaches its owner the same hour
  • Unallocated spend is its own line, with what would make it traceable
  • Every figure names its period, its scope and its cost basis
  • Every report says what the ledger does not cover

Autonomy earned, one control at a time

The team starts as a measurement function. It earns the right to act on the record of what its recommendations were worth.

Measure

Allocation, budget positions, burn rates and arbitrage shortlists. Read-only from day one, and that alone usually finds the first surprise.

Recommend

Model and caching changes with the evidence attached, and the ceilings the team would set. A person applies them.

Tighten

Once tightenings have been upheld without reversal, the operator may lower one ceiling on a defined trigger and report at once — within a floor it cannot cross.

Always human

Raising a budget, switching a model, touching a key. These change what an agent is, not what it may spend.

Nothing is autonomous by default, and the first tightening that stops legitimate work sends the control straight back to recommend-only.

What is live, what early access adds

Live on the platform today

  • Per-call metering of every model call, tagged by workspace, team, agent, project, key and use case
  • Budgets checked before the call, with warn, throttle and block
  • A cancelled run stops billing: an aborted call records what it burned instead of vanishing
  • Per-scope provider keys, so a key can be a cost boundary
  • Spend from agents built elsewhere, once they run through the governance gateway
  • Installation on your premises

Built with early-access customers

  • The chargeback statement per business unit, project and use case
  • Budget positions that say which of your budgets can actually fire
  • Model arbitrage from your measured traffic and adequacy evidence
  • Audited before/after savings, and the gain-share billed on them
  • Metering of compute that is not inference, and of the vendor coding CLIs that report nothing today
  • Governed cloud rightsizing — the cloud connectors read inventory and health today, not bills

Priced per band of governed AI spend, plus a share of the savings the before/after audit proves — never a share of an estimate. Carbon is not measured per workload anywhere on this platform, so no statement from this team carries a carbon figure; for what the runtime does to cut token consumption, see cost & carbon. Your own platform and finance team runs it.

See every offer · Agent Fleet Guardian