Token salaries · gateway caps · cost telemetry
Hard-stop budgets for AI agents: never get a surprise bill
Autonomous agents can loop, retry, and burn tokens faster than you can notice. The fix isn't hoping they behave — it's a budget they cannot exceed. In PAI every Pie gets a salary: a warning at 80%, an automatic pause at 100%, enforced at the gateway and signed into the audit chain.
Why AI agent costs run away
The same autonomy that makes an AI agent useful makes its spend hard to predict. A stuck task retries. A research loop fans out. A misfired webhook fires a thousand times. With raw LLM tooling, the first sign of trouble is often the invoice. LLM cost control can't be a dashboard you check after the fact — it has to be a limit the agent runs into the moment it goes too far. That's what a hard-stop budget is.
Token salaries: every Pie gets paid, and capped
In PAI, each Pie has a token salary — a spend allowance you set, the same way you'd give an employee a budget. It's a dial you watch from the same chat thread you use for approvals. The homework-helper Pie gets pocket money; the research Pie gets a research budget. Nobody gets a blank cheque.
80% warn, 100% auto-pause
Two thresholds do the work. At 80% of its salary, a Pie triggers a soft warning — a heads-up so you can top up, raise the cap, or simply let it ride. At 100%, the Pie auto-pauses: it cannot spend another token until you act. The warn is a notification; the stop is a wall. You never wake up to a five-figure bill from an agent that quietly went sideways overnight.
Two layers of enforcement
The auto-pause isn't the only line of defence. Spend is metered at the model gateway — the single point your keys flow through — with hard caps enforced there, plus cost-spike anomaly rules that flag a sudden jump even before a cap is hit. So you get double enforcement: the salary pause at the workforce layer and the gateway cap underneath it. An AI spend limit that lives in two places is one a single misconfigured flag can't switch off.
Per-agent and per-company — one rollup for you
Budgets nest. Every Pie has its own cap; every company has its own cap; and you see one rollup across the lot. Run three businesses and each gets an isolated budget with its own burn-rate view — the agency client's spend never bleeds into yours, and the “family” company can run on the tightest leash of all. (More on that on multi-company AI.)
Cost telemetry on every heartbeat
You can't govern what you can't see. Every time a Pie wakes on its heartbeat, its cost is recorded — so the burn-rate dial is live, not a monthly guess. Because you run BYOK (bring your own keys), that spend is your real, at-cost provider spend with no markup. Privacy-routed work that runs on a local model costs nothing per token, so it never eats the cloud budget at all.
The MOAT: budgets you can prove
Here's the part raw orchestration can't match. Every budget event — a soft warn, an auto-pause, a cap raise — is signed into a tamper-evident audit chain as an integrity event you can verify offline. A budget the runtime merely reports is a budget you have to trust. A budget whose every breach is cryptographically recorded is one you can prove. That signed cost trail is part of the Governance Plane — the leash and the flight recorder — and it's why an AI workforce here is an accountable cost centre, not a liability with a credit line.
Keep reading
Budgets are one control. These cover the rest of the leash.
AI CEO Autopilot
Set a goal; an AI CEO staffs and runs it — and asks before it spends, every time.
Managed AI agents that don't get stuck
Stall detection and run contracts stop the loops that drive runaway spend in the first place.
Run multiple businesses with one AI workforce
Per-company budgets with a single rollup — isolation enforced by the database.
AI agent budgets — FAQ
Give it a budget it physically cannot exceed. In PAI every Pie has a token salary; spend is metered at the model gateway on every heartbeat. At 80% of the cap you get a warning; at 100% the Pie auto-pauses. Because the cap is enforced at the gateway — not by asking the agent to behave — a runaway loop stops itself instead of emptying your account.
The soft warn (80%) is a heads-up: a message so you can top up, raise the salary, or let it ride. The hard stop (100%) is structural: the Pie pauses and cannot spend another token until you act. One is a notification; the other is a wall. You get both, per agent.
Yes. Each Pie has its own salary, and each company has its own cap with a single rollup for you. If you run three businesses, every business has an isolated budget and the family company can have the strictest of all. A cost-spike anomaly rule watches across the board for sudden jumps.
They make it honest. You bring your own AI provider keys, so spend goes to your provider account at cost — no token markup, no middleman. The budget is denominated in your real spend, not a marked-up credit. Local models (on a Box Plus) cost nothing per token at all, so privacy-routed work doesn't touch the cloud budget.
Yes. Gateway caps and the cost-spike anomaly rules are live and tested in our stack. The token-salary layer with the 80% soft-warn and 100% auto-pause rides the workforce layer and is enableable — switched on during setup. Either way, every budget incident is signed into the tamper-evident audit chain as an integrity event.
Govern spend structurally, not by hoping
See hard-stop budgets alongside every other capability — and the honest status of each — then choose your plan.