A warning is a message. A cap changes execution.
Simon Willison argues that pay-by-usage products need hard budget caps by default. His test is concrete: after the configured amount, the service stops accepting more paid work instead of sending an email while charges continue.[1]
That distinction matters more when an agent can create resources, retry a failed call, or keep working after the operator leaves. An alert depends on a person noticing it. A hard stop changes what the system can do.
Do not ask one limit to do two jobs. Stop the task locally. Stop the bill at the provider.
The provider's stop has a shape.
AWS now documents project-level spend limits for part of its new account experience. When a project reaches the limit, AWS pauses the project and stops its resources. The docs say the feature is aimed at experiments, learning, and sandbox work. They also warn that rollout is limited and that the minimum limit is the greater of $20 or AWS's conservative spend estimate.[2]
Google Cloud's public-preview Spend Caps take a narrower route. A cap applies to one service in one project for a fixed monthly period. Google says it restricts further cost-incurring usage after the threshold. Other services remain active. Fixed commitments can continue billing. For supported AI services, Google says enforcement happens within minutes, not at the exact instant the displayed total crosses the line.[3]
Both controls are useful. Neither supports the sentence "cloud spend cannot exceed this number" without qualification. Scope, billing delay, minimums, existing commitments, paused resources, and preview availability all belong in the run card.
The task needs its own fuse.
A provider cap is the last wall. Put a smaller stop in the agent job. Count the paid calls, elapsed time, retries, created resources, or estimated cost that the runtime can observe. Refuse another unit of work when the first local boundary opens.
The local fuse should fail closed when metering disappears. A missing price, unknown provider route, or unreadable usage response is not zero cost. Stop, retain the partial receipt, and require a person to choose the restart route.
Design the failure before turning the key.
Name the workload that may stop. Decide whether partial output remains readable. Keep the cleanup path outside the failed agent loop. Give one owner permission to raise or reset the limit. Record the old limit, new limit, reason, and time whenever that owner changes it.
A monthly cap is too coarse for a five-minute task. A per-task fuse is too weak for a leaked key used by another process. Use both. Test them with a small disposable workload before an unattended run depends on them.