Prerequisites
- You can create policy rules through
/v1/policies. - Read budgets and quotas, in particular which ceiling is admission-only.
API=http://localhost:8000 and a bearer token in $TOKEN.Steps
Choose the ceiling
Amounts are accepted as a string or a number; use a decimal string such as
"250.00" to avoid float rounding. They come back as strings.Create the cap
subject_type: "agent" with the
agent UUID bounds one agent; subject_type: "user" with a user id bounds
whatever that person launches. Lower scopes may only lower the number.Adjust an existing cap instead of stacking one
The workspace baseline already contains a monthly and a per-run cap. Find the row
and patch it rather than adding a second:
PATCH replaces params wholesale, so include period even when only the
amount changes. To switch a cap off without deleting it, send
{"enabled": false} — disabled rules are skipped when the layer compiles.Verify
Preview the merged ceiling before running anything:BudgetWarning at 80 percent
and BudgetExceeded when it stops:
Troubleshooting
422 when creating the rule or previewing
422 when creating the rule or previewing
A lower scope tried to raise a higher one’s ceiling. Budgets merge by taking
the minimum, and the resolver rejects rather than silently clamping so the
mistake is visible. Raise the parent first, or lower the child.
The cap was set and the task still overspent
The cap was set and the task still overspent
The monthly cap is admission control only. A task that starts under the cap
runs to completion no matter how far past the cap the workspace goes, and
the check is a read-then-decide with no lock, so concurrent task creation
can cross it. For a hard per-task bound use the run budget, which is
enforced inside the loop and before each call.
Month-to-date looks wrong
Month-to-date looks wrong
It is summed from
total_cost on the workspace’s task rows from the first
of the current UTC month. Spend from a task still in flight is not fully
counted until it finishes, so the figure trails reality while work is
running.A task is rejected because a ceiling is missing
A task is rejected because a ceiling is missing
Runtime execution requires an explicit run budget, total and per-call token
ceilings, and agent-loop limits in the resolved governance snapshot.
Deleting the persisted workspace defaults does not reveal a built-in numeric
fallback; it makes the runtime contract invalid. Restore the missing policy
rows and re-check the preview output.
max_tokens_per_call appears lower than expected
max_tokens_per_call appears lower than expected
It is enforced on every LLM call. The resolver takes the strictest positive
value from the effective policy, the request, and the model’s declared
output capability, so inspect all three sources before changing the
workspace rule.
The numbers do not match between the loop and the per-call gate
The numbers do not match between the loop and the per-call gate
They are two enforcement points reading one resolved ceiling — the tighter
of the per-request budget and the policy value. If they disagree, the
request carried its own
budget_usd ; the minimum wins, so check what the
caller sent.A budget denial appears as a failed activity, not a graceful stop
A budget denial appears as a failed activity, not a graceful stop
The per-call gate raises rather than returning a message the model can
answer. The graceful path — the loop noticing exhaustion and completing —
comes from the in-workflow tracker. Both are expected; which you see depends
on where the ceiling was crossed.
Related
Budgets and quotas
The enforcement points and their thresholds
The policy engine
How ceilings merge across scopes
Authorize a tool call
The other restriction you write as a policy rule