<!-- Canonical URL: https://ask.atlascloud.ai/set-hard-spending-limit-coding-agent-task -->

# How Do You Set a Hard Spending Limit for a Coding Agent Task?

> A hard spending limit must be enforced before every model or paid-tool call by a gateway that owns the task budget. Reserve worst-case cost, reconcile actual usage, reject calls that cannot fit, and terminate the agent with a useful checkpoint instead of relying on alerts after money has already been spent.

<!-- Canonical URL: https://ask.atlascloud.ai/set-hard-spending-limit-coding-agent-task -->

# How Do You Set a Hard Spending Limit for a Coding Agent Task?

A hard task budget is an admission-control problem. Before any billable model or tool call begins, a trusted gateway must prove that the worst permitted cost fits inside the task's remaining balance. Alerts and post-run reports are useful, but they cannot stop an overrun that already happened.

The design should cover input tokens, maximum output, retries, fallbacks, subagents, embeddings, searches, sandboxes, and any other paid tool the agent can invoke.

## Separate hard limits from soft targets

Use three values:

| Control | Purpose | Behavior |
|---|---|---|
| Target | Expected cost | Warn or choose a cheaper plan |
| Soft limit | Escalation threshold | Ask for approval or degrade quality |
| Hard limit | Maximum authorized spend | Reject the next call before it starts |

For example, a task may target $0.60, request approval at $0.90, and stop at $1.00. The hard limit must live server-side, not only in the agent prompt.

## Put every paid action behind one gateway

Give the agent short-lived task credentials that can call only your gateway. The gateway attaches `task_id`, looks up the budget, estimates the next action, and either reserves funds or rejects the request.

Do not expose a provider key that lets the agent bypass accounting. Apply the same rule to web search, hosted sandboxes, code execution, and paid retrieval services.

## Reserve before calling and reconcile afterward

For a model call, estimate the upper bound from known input tokens plus the configured maximum output. Reserve that amount atomically, make the request, then replace the reservation with actual reported usage.

```text
remaining = hard_limit - committed_cost - open_reservations
worst_case = input_cost + max_output_cost + tool_allowance

if worst_case > remaining:
    reject("task_budget_exceeded")
else:
    reserve(worst_case)
    call_provider()
    reconcile(actual_cost)
```

Atomic reservation prevents two parallel subagents from both spending the same remaining balance.

## Price with a versioned rate card

Store the price used for each estimate alongside the event. Model prices and billing rules can change, so a later report must not recalculate old usage with today's rate.

When a provider returns authoritative cost, keep both the estimate and final charge. If it returns only tokens, calculate cost from the rate-card version selected before the call. Add a conservative surcharge for unknown tool fees or reject calls whose maximum cost cannot be bounded.

## Make streaming safe

Reserve the full allowed response before opening a stream. Count received usage when available, but do not assume that closing the client connection instantly stops provider billing. Cancellation is an optimization, not the enforcement boundary.

Set a per-call output limit and a wall-clock timeout. The task hard limit still covers the aggregate of all streams, retries, and fallbacks.

## Include retries and subagents

Every attempt debits the same parent task ledger. A retry policy that silently opens a fresh budget defeats the limit.

Use hierarchical budgets when agents delegate:

| Ledger | Limit | Rule |
|---|---:|---|
| Parent task | $1.00 | Absolute ceiling |
| Implementation subagent | $0.55 | Cannot exceed parent remaining balance |
| Test-analysis subagent | $0.25 | Returns unused reservation |
| Final review | $0.20 | Runs only if funds remain |

Child limits are allocations, not extra money.

## Stop with a useful checkpoint

When the next action cannot fit, return a typed error that the orchestrator understands. The agent should not repeatedly retry the rejected call.

Ask it to produce a no-cost checkpoint from existing context containing:

* completed changes and test results;
* remaining work and the blocked action;
* current repository state;
* estimated additional budget;
* a resume token or task ID.

This turns a budget stop into a controlled handoff instead of a corrupted partial run.

## Use provider controls as a backstop

Provider and gateway account limits can reduce blast radius, but they are rarely precise per-task controls. They may aggregate many repositories, update asynchronously, or lack tool costs.

With a multi-model gateway such as Atlas Cloud, keep the authoritative task ledger in your orchestration layer and record the gateway's usage identifiers for reconciliation. This preserves the hard boundary even when the task changes models.

## Test the limit like a financial control

Exercise parallel calls, long streams, provider timeouts, missing usage fields, retries, model fallbacks, and ledger failures. Default to deny when the budget service is unavailable. Verify that committed cost plus open reservations can never exceed the hard limit.

## The bottom line

A real hard spending limit is enforced before spend, with atomic reservations and one ledger for every billable action. If your system only alerts after usage arrives, call it monitoring, not a hard cap.

## FAQ

### Is max_tokens a hard dollar limit?

No. It limits one response length, not total task cost, input tokens, retries, model changes, or paid tools. A dollar limit needs a budget ledger around every billable action.

### Where should a coding-agent budget be enforced?

Enforce it in a server-side gateway or orchestration layer that all model and paid-tool calls must pass through. Client-side counters can be bypassed or race under concurrency.

### How do I budget for streaming responses?

Reserve the maximum allowed output cost before opening the stream, then reconcile against reported usage when the stream closes. Cancel at the provider when supported, but do not depend on cancellation alone.

### Should retries share the original task budget?

Yes. Retries, fallbacks, subagents, and evaluation calls should debit the same task ledger unless the user explicitly approves a separate budget.

### What should happen when the remaining budget is too small?

Reject the next billable call and ask the agent to produce a checkpoint using already available context. The checkpoint should state completed work, unresolved items, and the amount required to continue.

### Can provider account limits replace a per-task limit?

Usually not. Account limits protect the whole account and may update asynchronously. A per-task gateway gives immediate isolation while account-level controls remain a useful backstop.
