Shipping an agent without a governor means shipping a retry loop that discovers your provider’s limits by triggering them. This build log walks through the design of a client-side rate-limit governor — and the documented provider behaviours that forced each decision along the way. It is a design walkthrough grounded in provider documentation and engineering references; no governor was deployed for this article, so the failure modes below are documented behaviours rather than test results.

The governor is not a retry loop

A retry loop reacts after the 429 arrives; a governor decides before the request leaves your process. It has three possible outcomes for every call: send now, reject locally, or park the request for a delayed retry. That distinction matters because of two findings from Google’s SRE book. Handling Overload argues for client-side throttling: an overloaded backend should accept only what it can process and reject the rest gracefully — never turn away all traffic. Addressing Cascading Failures documents the complementary problem: retries amplify load on an already-degraded dependency. A naive retry loop is load amplification you installed yourself. The governor exists to make the shed-or-send decision early, locally, and with enough context to distinguish a transient limit from a terminal one.

The failure that changed the design

Desk research surfaced three documented provider behaviours that a plain counter-and-retry design gets wrong.

The first is the Anthropic monthly spend cap. Per Anthropic’s rate limit docs, hitting the cap returns HTTP 429 with error.details.error_code = enforced_spend_limit_reached and no Retry-After header. The documented failure mode: blind retries — including SDK auto-retries — keep failing until access resumes. A governor that treats every 429 as retryable would hammer a terminal condition all month. An adapter must detect that error code and classify it as non-retryable.

The second is burst semantics. The same docs describe enforcement as a token bucket with continuous refill, not a fixed-interval reset — so a limit like 60 RPM may effectively be about one request per second, and short bursts still return 429 even when your minute average looks compliant. The governor would fail here if it tracked a fixed one-minute window: it would green-light a twenty-request burst and watch half of it get rejected.

The third is OpenAI’s token accounting. OpenAI’s guide separates RPM from TPM and computes each request’s token cost as the maximum of max_tokens and an estimate derived from the request’s character count. An oversized max_tokens alone can trip TPM while RPM looks perfectly healthy. A governor that only counts requests would pass that request straight into a 429.

Provider limit landscape

The limits are separate axes with different reset semantics, and one scalar counter cannot represent them.

OpenAI enforces RPM and TPM independently, and its Batch API does not count against synchronous rate limits — a genuine escape hatch for async workloads. Anthropic exposes anthropic-ratelimit-* response headers reporting the most restrictive limit currently in effect — which makes them the right runtime input for re-tuning local buckets. Gemini spans RPM, TPM on input tokens, and RPD (requests per day) across usage tiers, with a separate Batch API. A non-LLM reference confirms the pattern is broader than chat models: Stripe returns 429 Too Many Requests and publishes concurrency limits alongside rate limits.

The design consequence: the governor needs one bucket per axis — request rate, token rate, daily quota, concurrency — and each axis resets differently. A per-second token bucket and a per-day counter share almost no semantics. Modelling them as a single number is how governors pass requests that fail on the axis nobody counted.

Rate-limiting algorithms compared

The core of the governor is a classic algorithm choice, and the four candidates differ exactly where agent traffic is hardest: bursts, memory, and multi-replica accuracy.

Algorithm Burst handling Memory Accuracy across replicas Notes
Fixed window Poor — two full windows’ worth can pass at a boundary Minimal: one counter plus window start per key Only if every replica shares the counter and clock; local windows multiply Simplest to reason about; boundary spikes are the classic failure
Sliding-window counter Better — weights the previous window to smooth edges Small: two counters per key, or a timestamp log for exact counts Needs a shared store plus agreed time slices An approximation; still requires coordination to mean anything across replicas
Token bucket Good — tokens accrue at a fixed rate; bucket depth sets burst tolerance Tiny: token count and last refill timestamp per key Local buckets across replicas sum to N× the intended rate; needs a central counter Matches provider enforcement; non-conforming requests can be dropped, queued, or marked
Leaky bucket Shapes output toward a constant rate via a queue; rejects when full Queue depth per key can grow under load Same coordination problem as token bucket Smooths traffic nicely, but adds queuing latency you rarely want in a request path

Token bucket wins the base slot for three reasons. It natively expresses both a sustained rate and a burst tolerance, which matches how providers actually enforce limits — Anthropic’s continuous-refill bucket is literally this algorithm. It has constant memory: a token count and a last-refill timestamp per key. And non-conforming requests can be dropped, queued, or marked, which maps exactly onto the governor’s three-outcome decision model. Fixed windows fail at boundaries precisely where agent loops burst; sliding-window counters smooth the edges but still need shared state; leaky buckets buy smoothness with queue latency.

Why token bucket plus provider cost estimators

A token bucket alone enforces request rate and burst; it knows nothing about tokens. So the base algorithm is paired with per-request cost estimators, one per provider. For OpenAI, the estimator replicates the documented cost rule — max(max_tokens, character-count estimate) — and the governor pre-charges the token bucket with that figure before sending. For Anthropic, the anthropic-ratelimit-* headers report the most restrictive limit currently in effect, so the adapter should re-read them on every response and re-tune buckets instead of trusting a startup config forever.

Go’s golang.org/x/time/rate is a direct implementation of the pattern: NewLimiter(r, b) builds the bucket, Allow() is the cheap non-blocking check, and Reserve() / WaitN(ctx) support delayed admission. A practical governor runs one limiter per axis and admits a request only when every limiter allows.

Code shape and the burst-size error

The decision model compresses to three returns: Allow (decrement buckets, send), Reject (shed locally, caller decides), Backoff (park with a deadline). Three pieces carry the documented failure modes:

type Decision int

const (
    Allow   Decision = iota
    Reject
    Backoff
)

func (g *Governor) Admit(req Request, body string) Decision {
    // 1. Cost estimator: OpenAI charges max(max_tokens, character estimate).
    cost := max(req.MaxTokens, estimateTokens(body))

    // 2. If the cost can never fit the bucket, refuse up front. x/time/rate's
    //    WaitN returns an error when n exceeds the burst size, so never hand
    //    it an oversized n.
    if cost > g.tokenBurst {
        return Reject
    }

    // 3. Reserve before spending: Reserve is refundable via Cancel, so a
    //    partial admission never leaks capacity from the request bucket.
    res := g.toks.ReserveN(time.Now(), cost)
    if !res.OK() {
        return Reject
    }
    if !g.reqs.Allow() {
        res.Cancel() // refund the token reservation
        return g.parkWithJitter(req) // Backoff
    }
    return Allow
}

// Terminal-error classifier: a spend-cap 429 is never retryable.
func (g *Governor) Classify(err error) Outcome {
    var ae *APIError
    if errors.As(err, &ae) && ae.Code == "enforced_spend_limit_reached" {
        return Terminal // no Retry-After; retries never succeed — alert, do not retry
    }
    return Retryable
}

// Full-jitter park helper: sleep = rand(0, min(cap, base*2^attempt)).
func (g *Governor) parkWithJitter(req Request) Decision {
    ceiling := min(g.cap, g.base<<uint(req.Attempt))
    wait := time.Duration(rand.Int63n(int64(ceiling) + 1)) // full jitter
    g.queue.Park(req, time.Now().Add(wait))
    return Backoff
}

The burst-size note is not cosmetic: the documented behaviour in x/time/rate is that WaitN returns an error when n exceeds the bucket’s burst size (and Reserve().OK() returns false), so the wrapper checks the cost against the burst and refuses an oversized request up front instead of handing it to the limiter. The retry helper uses AWS’s exponential backoff with full jitter, which breaks synchronized retry storms across a fleet — AWS open-sourced a simulator demonstrating the effect.

Local state vs distributed limits

A token bucket in one process enforces one process’s behaviour. Run N replicas against a shared provider account with local buckets and you have configured N× the intended rate — the provider will demonstrate this with 429s. The fix is shared state: either a central counter or partitioned leases where each replica owns a fraction of the budget.

LiteLLM Router is the working reference: it tracks RPM/TPM usage in Redis across horizontally scaled replicas and layers cooldowns, fallbacks, timeouts, and fixed-plus-exponential-backoff retries on top. The more ambitious ancestor is Doorman, Google/YouTube’s distributed client-side rate limiter — Apache-2.0 licensed, and archived in 2024, which says something about the maintenance cost of doing this properly. For this build, the honest move is to document the limitation inside the component: v1 enforces per-process limits only, and running it multi-replica without a shared counter is an explicit, detectable misconfiguration — not a silent one.

Backoff and Retry-After reality

When a 429 does arrive, Retry-After is either delta-seconds or an HTTP-date — and MDN notes client and server support is still inconsistent. The governor honours the header when present, with three guardrails: parse both forms, cap the honoured delay at a configured ceiling so a bad value cannot park a worker indefinitely, and fall back to full-jitter backoff when the header is absent — which, per the Anthropic spend-cap behaviour, is exactly the case where blind retrying is worst.

The retry budget matters as much as the delay. Full-jitter backoff exists because synchronized retries arrive as a storm; a bounded per-window retry budget caps that amplification. And no retry at all should happen for non-idempotent work without a dedupe key — a retried agent action that already executed is a correctness bug, not a rate-limit problem.

SRE tradeoffs: reject gracefully, protect the dependency

Two tradeoffs shaped the edges. First, rate is not concurrency. NGINX separates limit_req for rate from limit_conn for concurrency, and its own editor note observes that burst and delay are computed at a sliding millisecond rate — a reminder that time granularity matters. The governor mirrors the separation: one token bucket per rate axis, plus a semaphore-style cap on in-flight requests.

Second, shedding policy. Handling Overload’s rule — accept what you can process, reject the rest gracefully, never turn away all traffic — becomes a priority ladder: shed low-priority speculative calls first, keep headroom for interactive traffic, and fail fast locally rather than consuming a provider slot on a request you expected to drop. Every retry draws from a bounded budget so the governor cannot become the cascading-failure amplifier it was built to prevent, and parked requests live in a queue with deadlines — past the deadline, the request is rejected, not silently aged.

FAQ

Should I retry a 429 from the Anthropic spend cap? No. The documented behaviour is enforced_spend_limit_reached with no Retry-After, and retries — including SDK auto-retries — fail until access resumes. Classify the error as terminal, alert, and stop the retry path.

Can I route everything through Batch APIs to avoid rate limits? Only async workloads. OpenAI’s Batch API does not count against synchronous rate limits, and Gemini offers a separate Batch API — but batch responses are not for latency-sensitive agent loops. Keep the governor on the synchronous path.

Does a local token bucket enforce provider-wide limits? No. It enforces one process’s behaviour. N replicas with local buckets send N× the intended rate. Use a shared counter — the LiteLLM Redis pattern — or partition the budget into per-replica leases.

What burst size should the bucket have? Keep it small and let the sustained rate carry the load. Provider headers report remaining quota, not the burst a provider tolerates, so size the depth from the enforcement granularity: Anthropic may enforce 60 RPM as roughly one request per second, so a large burst fails even when the minute average looks fine.

Should I always trust Retry-After? Honour it when present — it is delta-seconds or an HTTP-date — but cap the honoured delay and fall back to full-jitter backoff when it is missing, since support remains inconsistent.

The Bottom Line

The governor is a token bucket per limit axis, a per-provider cost estimator, a terminal-error classifier, a capped honouring of Retry-After with full-jitter fallback, and an honest single-node scope statement. The algorithm was never the hard part — the documented edge cases were: a 429 that must never be retried, a burst that fails at one request per second, and a token budget charged at max(max_tokens, char estimate). Build those three behaviours first. Everything else is a counter.

How This Guide Was Built

This is a desk-research build log. The design and every failure mode above are drawn from provider documentation (OpenAI, Anthropic, Google Gemini, Stripe), engineering references (Google SRE, AWS, NGINX, MDN), and open-source implementations (golang.org/x/time/rate, LiteLLM Router, Doorman). To be plain about scope: no governor was deployed to production for this article, and no benchmarks were run. Every “would fail here” is a documented provider behaviour, not a test result — validate against your own traffic before shipping.

  • ToolBrain — tool reviews, LLM comparisons, and AI workflow guides
  • CodeIntel Log — code quality, debugging, and software engineering benchmarks

Cross-links automatically generated from NiteAgent.

← Back to all posts