Agents with tools and no typed budget fail in a predictable way. The identity can call the tool. Nobody set how many times, for how much, to which class of recipient, or for how long that grant lasts. The model keeps proposing. The runtime keeps saying yes until something expensive or irreversible lands.
This post is about typed action budgets: tool or category allowlists, amount and count ceilings, and time-to-live for approvals and grant windows. It sits beside scoping permissions and human approval, which covers identity vs permission vs approval. It is not the finance-only payment and ledger gate matrix, and it is not MCP retry and remap binding. Those posts stay important. Here the focus is the numeric and temporal envelope that turns "allowed tool" into "allowed within this budget."
Permissions are not budgets
A permission answers whether the identity may use a tool at all. A budget answers how much of that permission may be consumed before the next human decision.
Without budgets you get familiar failure modes:
- a draft-and-send mailer that was meant for five follow-ups sends fifty;
- a CRM updater that was meant for three allowlisted fields walks an export;
- a refund tool with "ask first" still has no amount ceiling after the first yes;
- a sticky "always allow" preference that never expires covers a wider tool later;
- a weekend grant for one vendor change stays live on Monday for every vendor.
OWASP's excessive agency guidance names the pattern: too much functionality, too much permission, too much autonomy. Budgets are the autonomy layer expressed as numbers and clocks, enforced in deterministic code outside the model. A prompt that says "be careful with spend" is not a budget.
What a typed action budget is
A typed budget is a structured grant attached to a tool, a tool category, or a concrete approval. It should be machine-checkable on every execute call.
| Dimension | What it bounds | Example |
|---|---|---|
| Tool or category | Which execute surfaces may run | `send_email` only; or category `crm.write` excluding `crm.export` |
| Amount ceiling | Money, quantity, or size per call and per window | Max $250 per refund; max 20 MB per file write |
| Count ceiling | How many times in a window | Max 10 sends per hour; max 3 vendor updates per day |
| Recipient or target class | Who or what may be touched | Existing customers only; no new payees; tenant-scoped record IDs |
| Expiry / TTL | How long the grant or approval remains valid | Approval valid 15 minutes; daily auto budget resets at 00:00 UTC |
| Window | The period that amount and count accumulate in | Per run, per hour, per day, per approval |
Typed means the runtime knows the shape before the model proposes arguments. The check is not "does this look reasonable." It is "does this call fit the remaining grant."
Map autonomy tiers to budget shapes
Autonomy is not one switch. Match the tier to the budget you are willing to enforce.
| Autonomy tier | Typical work | Budget shape | Human role |
|---|---|---|---|
| Read / draft | Search, summarize, draft in a queue | Tool allowlist only; no external execute; optional count on heavy reads | Review drafts when useful |
| Bounded auto | Low-impact writes inside policy | Tool/category + amount + count + recipient class + daily window | Spot-check and exception review |
| Human-gated | Money, delete, new payee, external send above threshold | Zero auto amount for that class; approval carries immutable payload + short TTL | Named approver before execute |
| Unavailable | Open-ended export, free-form shell, admin schema change | No tool; no budget can unlock it through ordinary approval | Elevated break-glass path only if policy exists |
Read/draft should not share a credential that can also pay vendors. Bounded auto should fail closed when a ceiling is hit, not silently widen. Human-gated means the approval is itself a short-lived budget for one payload. Unavailable means no amount of polite prompting creates the tool.
For money-moving rows that need a finance-specific matrix, use the finance approval gates post. The budget idea here is the general control that those finance rows specialize.
Example matrix: same agent, different budgets
Build from real tools. This pattern is a starting board, not a universal policy.
| Action | Tier | Tool/category | Amount | Count | Expiry / window |
|---|---|---|---|---|---|
| Search CRM contacts | Read | `crm.read` | N/A | Soft cap if search is expensive | Session or day |
| Draft follow-up email | Draft | `mail.draft` | N/A | Optional draft cap | Draft lives until edited or discarded |
| Send email to existing contact | Bounded auto or gated | `mail.send` | Max recipients = 1 | 10 / hour | Daily auto budget; above threshold needs approval TTL 15 min |
| Update allowlisted CRM fields | Bounded auto | `crm.write.fields` | Max 5 fields / call | 50 / day | Day window; field allowlist immutable |
| Create refund | Human-gated | `payments.refund` | Max $250 / approval | 1 execute per approval | Approval TTL 15 min; no sticky always-allow |
| Change vendor bank details | Human-gated | `vendors.bank.update` | N/A (identity change) | 1 per approval | Approval TTL 30 min; out-of-band verify outside the agent |
| Export full tenant data | Unavailable | None | None | None | Break-glass only outside ordinary agent tools |
| Shell / arbitrary HTTP | Unavailable in production agents | None | None | None | Use narrow typed tools instead |
Notice prepare and execute stay different tools. A draft budget must not imply a send budget. That is the same prepare/execute split used in permissions and approval design.
What must stay immutable when a human approves or a budget grants
Approvals and auto budgets both grant execute rights. The grant must bind to facts the model cannot rewrite after the click or after the daily window opens.
Immutable on approval:
- tool name (or exact tool version / handler id);
- canonical argument payload or hash (payee, amount, currency, record ids, recipients);
- policy rule that required the gate;
- approver identity and timestamp;
- expiry time;
- remaining count = 1 unless dual-control policy says otherwise.
Immutable on an auto budget grant:
- tool or category id;
- amount and count ceilings;
- recipient/target class rules;
- window start and end;
- which identity and tenant the budget belongs to.
If the agent remaps to a different tool, widens amount, swaps payee, or retries after expiry, deny. That is where this post meets retry and remap binding: the first yes is not a blank check for a different call later.
Enforce outside the model
Put budget checks in the execution path the model cannot edit:
- MCP tool handlers and the narrow execution service in front of each side effect;
- deterministic validators for amount, count, recipient class, and TTL before the external API call;
- durable counters and audit rows so a restarted process does not reset spend;
- fail-closed behavior when the budget store is unavailable.
Do not keep the only copy of remaining balance in the prompt or in chat memory. The model is a proposer. The budget ledger is infrastructure.
Self-hosted layouts still need the same idea inside one trust boundary. Trust boundaries for self-hosted agents decide who shares a Gateway. Budgets decide how much that shared boundary may spend before the next human gate. The deployment checklist covers host and channel hardening; budgets cover the run-time envelope once tools are live.
Practical design rules
- 1Name budgets after tools or categories operators already understand. Avoid vague pools like "ops allowance" that every write can drain.
- 2Prefer small windows. A daily send cap beats a monthly cap you only notice at month end.
- 3Expire approvals in minutes for money and external sends. Long-lived approvals become ambient admin.
- 4Separate ceilings per rail or system when risk differs (email vs payout vs ledger post).
- 5Show remaining budget to the reviewer when you ask for a raise. "Approve +$500 once" is clearer than "approve higher limit" with no numbers.
- 6Log denials for ceiling hits. Silence teaches nobody why the agent stopped.
- 7Never let ordinary approval create an Unavailable tool. Break-glass is a different path with its own owners and audit.
How this fits OrchestriAI delivery
On delivery work, AI agent systems define the autonomy tier and the human gate. Custom MCP server development puts tool allowlists, amount/count checks, and approval TTL at the handler boundary. Systems integration keeps counters, queues, and side effects on identities that match the budget you intended.
The operating rule is short: narrow tools, typed budgets, approval bound to an exact payload with a clock. That is how you give agents useful autonomy without open-ended spend, open-ended rate, or open-ended time.
