Token salaries · gateway caps · cost telemetryEnableable

Hard-stop budgets for AI agents: never get a surprise bill

Autonomous agents can loop, retry, and burn tokens faster than you can notice. The fix isn't hoping they behave — it's a budget they cannot exceed. In PAI every Pie gets a salary: a warning at 80%, an automatic pause at 100%, enforced at the gateway and signed into the audit chain.

The spend controls

Six controls, one hard limit

  • Token salariesEach Pie gets a spend allowance you set — nobody gets a blank cheque.
  • Hard stop at 100%Warn at 80%, auto-pause at 100% — the Pie physically can't overspend.
  • Cost on the chainEvery budget event is signed into a tamper-evident audit chain you verify offline.
  • BYOK, at costBring your own keys; spend hits your provider account directly, with no markup.
  • Per-Pie, per-companyBudgets nest: every Pie and company gets its own cap, one rollup for you.
  • Approvals in chatTop up or raise a cap from the same thread you approve actions in.

Why AI agent costs run away

The same autonomy that makes an agent useful makes its spend unpredictable: a stuck task retries, a research loop fans out, a misfired webhook fires a thousand times. Real LLM cost control isn't a dashboard you check later — it's a limit the agent hits.

80% warn, 100% auto-pause

Two thresholds do the work. At 80% of its salary a Pie sends a soft warning — top up, raise the cap, or let it ride. At 100% it auto-pauses and cannot spend again until you act. One is a notification; the other is a wall.

The MOAT: budgets you can prove

Here's what raw orchestration can't match: every budget event is signed into a tamper-evident audit chain you can verify offline. A budget the runtime merely reports is one you trust; a budget whose every breach is signed is one you can prove.

AI agent budgets — FAQ

Give it a budget it physically cannot exceed. In PAI every Pie has a token salary; spend is metered at the model gateway on every heartbeat. At 80% of the cap you get a warning; at 100% the Pie auto-pauses. Because the cap is enforced at the gateway — not by asking the agent to behave — a runaway loop stops itself instead of emptying your account.

The soft warn (80%) is a heads-up: a message so you can top up, raise the salary, or let it ride. The hard stop (100%) is structural: the Pie pauses and cannot spend another token until you act. One is a notification; the other is a wall. You get both, per agent.

Yes. Each Pie has its own salary, and each company has its own cap with a single rollup for you. If you run three businesses, every business has an isolated budget and the family company can have the strictest of all. A cost-spike anomaly rule watches across the board for sudden jumps.

They make it honest. You bring your own AI provider keys, so spend goes to your provider account at cost — no token markup, no middleman. The budget is denominated in your real spend, not a marked-up credit. Local models (on a Box Plus) cost nothing per token at all, so privacy-routed work doesn't touch the cloud budget.

Yes. Gateway caps and the cost-spike anomaly rules are live and tested in our stack. The token-salary layer with the 80% soft-warn and 100% auto-pause rides the workforce layer and is enableable — switched on during setup. Either way, every budget incident is signed into the tamper-evident audit chain as an integrity event.

Govern spend structurally, not by hoping

See hard-stop budgets alongside every other capability — and the honest status of each — then choose your plan.

Explore the possibilities →