AI Agent Error Recovery: Retry, Timeout, Backoff, and Circuit-Breaker Patterns

Retry, timeout, backoff, and circuit-breaker patterns for AI agent error recovery. How to force 429 and 500 failures against a mock and prove your agent backs off and never double-sends.

Ashley Innocent

Ashley Innocent

1 September 2026

AI Agent Error Recovery: Retry, Timeout, Backoff, and Circuit-Breaker Patterns

Apidog for Enterprise

On-Premises Deploy

SSO & RBAC

SOC 2 Compliant

Explore Apidog Enterprise

Your agent calls an API. The API returns a 429. Your agent retries right away, gets another 429, retries again, and now you have a loop that hammers a throttled service until the run dies or the bill spikes. Nobody wrote that loop on purpose. It falls out of the naive version of “handle the error,” and it is the single most common thing developers ask about on the Anthropic SDK discussion board.

Error recovery is the part of agent building that separates a clean demo from something you can page someone about. The model is not the problem. The problem is what your code does when a tool call comes back slow, throttled, or broken. Get recovery right and a flaky dependency becomes a short pause the user never notices. Get it wrong and one 500 turns into an incident. This guide covers the four patterns that carry most of the load: retries with backoff, timeouts, circuit breakers, and idempotency keys. Then it shows how to test them against a mock, before a user finds the gaps for you. For the wider picture of how agents fail, start with why AI agents break in production.

button

You can’t test recovery against a healthy API

Here is the trap. Your dependency works fine in development. You write your agent, the calls succeed, the demo is clean, and you ship. The recovery code never ran once, because a healthy API never returns the errors it is supposed to handle. The first time your backoff logic runs is in production, against a real outage, with real users watching. That is the worst place to discover a typo in a retry loop.

So the rule is simple. To test recovery, you produce the failures on purpose. Stand up a mock of the API the agent calls, program it to return a 429, a 500, a timeout, or a malformed body, point the agent at it, and watch what it does. The failure becomes something you trigger in a test instead of something that triggers you at 3am. Apidog stands up that mock and scripts the responses, and it runs through the testing section at the end.

Retry with exponential backoff and jitter

A retry is the first line of defense, and the naive version is the trap. Catch the error, call again immediately. Against a transient blip that works. Against a service under load it makes things worse, because every failed client retries at the same instant and the stampede keeps the service down.

Two fixes stack together. Exponential backoff spaces the attempts out: wait 1 second, then 2, then 4, then 8, doubling up to a ceiling. The service gets room to recover instead of a wall of immediate retries. Jitter adds a random offset to each wait so a thousand clients that all failed at the same moment don’t all retry at the same moment either. Without it, backoff still produces synchronized waves.

Cap two things: the delay, so you don’t wait minutes between attempts, and the attempt count, so a permanent failure gives up instead of retrying forever. Three to five attempts covers almost every transient error. Past that you are usually retrying something that is not going to succeed. The Anthropic SDK does a chunk of this for its own calls: it retries connection errors and specific status codes with exponential backoff, and you set the ceiling with a max-retries option. It does not cover the other APIs your agent’s tools hit, so you wrap those yourself. Teams that run money through their retries learn this early, and our breakdown of retry logic for high-stakes APIs shows where a careless retry does real damage.

Set a timeout on every call

A retry only helps if the request fails. The nastier case is a request that never comes back: a dependency accepts your connection, then hangs. With no timeout, the tool call blocks and the whole run stalls behind one dead socket. No error, no recovery, only a stuck agent burning wall-clock time and token budget on nothing.

Every outbound call needs a timeout. Set a connect timeout for establishing the connection and a read timeout for waiting on the response, then a total budget for the whole agent run so a chain of slow-but-legal calls can’t outlive the user’s patience. When a timeout fires, treat it like any other retryable error: back off and try again, up to your cap.

Pick the numbers from real latency, not a guess. Set each timeout above the dependency’s p99 with headroom. Too tight and you abort calls that would have succeeded. Too loose and a hung dependency ties up the agent long past the point of usefulness. Give streaming responses their own budget, since a long completion is legitimately slow and a short fixed timeout kills it mid-stream.

Trip a circuit breaker when a dependency is down

Backoff handles a service that is briefly busy. It is the wrong tool for a service that is flat-out down. If a dependency has been failing for a minute, the next request will almost certainly fail too, and retrying it piles more load onto something already broken while the user waits for a failure you could have predicted.

A circuit breaker fixes this with three states. Closed is normal: requests flow and the breaker counts failures. When failures cross a threshold, it trips to open: it stops sending requests and fails fast for a cool-down window, so you are not paying the timeout on every call to a dead service. After the window, it goes half-open and lets a single probe through. If the probe succeeds, the breaker closes and traffic resumes; if it fails, it opens again and waits.

For an agent, the breaker turns “the payment API is down” into one fast, clean failure the agent can reason about, instead of forty slow timeouts that drain the token budget and the clock. Wire it per dependency, not globally, so a dead search API doesn’t stop the agent from using a healthy billing one.

Make retries safe with idempotency keys

Every pattern so far assumes retrying is safe. Often it is not. Your agent sends POST /charge, the server processes it, and the response times out on the way back. The agent never saw the success, so it retries, and now the customer is charged twice. The retry did exactly what you asked. The design was the bug.

An idempotency key closes the gap. The client generates a unique key per logical action and sends it with the request, usually as an Idempotency-Key header. The server records the key on first receipt and, if it sees the same key again, returns the original result instead of doing the work twice. Now a retry is safe by construction: the second POST /charge with the same key is a no-op that hands back the first charge.

The key has to stay stable across retries of the same action and change between different actions. Generate it once when you build the request, not inside the retry loop, or every attempt gets a fresh key and the dedupe never fires. Any tool call that creates or changes state (charges, orders, emails, records) needs one. Our guide on idempotency keys covers generation and server-side handling in full.

Survive rate limits and the RateLimitError loop

Rate limits deserve their own handling because they come with instructions. A rate-limit-exceeded response usually arrives as a 429 carrying a Retry-After header that tells you exactly how long to wait, in seconds or as a date. Respect it. If the server says wait 30 seconds and you retry in 2, you get another 429, and you have built the RateLimitError loop that fills the SDK discussion board: catch the limit, retry too soon, get limited harder, repeat until the run dies. A separate SDK thread covers the same wall developers hit here.

The fix is to let the server set the pace. When you get a 429, read Retry-After and wait at least that long before retrying. If the header is missing, fall back to exponential backoff with jitter. Cap the attempts so a sustained limit ends in a clean failure instead of an infinite wait. The Anthropic SDK already honors Retry-After for its own calls; the work is applying the same rule to the other rate-limited APIs your agent touches.

There is a proactive side too. If a provider allows a set number of requests per minute, meter your own calls with a token bucket so you stay under the ceiling instead of finding it by getting throttled. Recovery handles the limits you hit; pacing keeps you from hitting them.

How to test the recovery path

Now put it together. The patterns above are only as good as your proof they run, and the proof is a test that forces the failures a healthy API won’t give you. The shape reuses across every scenario:

  1. Mock the dependency. Stand up a mock of the API your agent’s tool calls, so you control every status code, header, body, and delay, and no real charge or email fires during the test.
  2. Program a sequence. Script the mock to answer a series of calls in order: first a 429 with Retry-After: 2, then a 500, then a 200 with a valid body. One endpoint, three scripted responses, a full recovery arc in a single run.
  3. Drive the agent at the mock. Point the agent’s tool at the mock URL instead of the real service and run the scenario end to end.
  4. Assert on behavior. Check what matters: the agent waited at least 2 seconds after the 429 before retrying, retried after the 500, succeeded on the third call, and never exceeded your attempt cap.

That one scenario proves backoff and Retry-After in a single pass. Add a second scenario for the give-up path: script the mock to fail every time and assert the agent stops at the cap and returns a clean error instead of looping. Add a third for the circuit breaker: fail enough calls in a row and assert the agent trips and fails fast instead of paying a timeout on every attempt.

The idempotency check is the one people skip, and it is the one that saves money. Script the mock to accept a mutating call, drop the response so the agent thinks it failed, then accept the retry. Now assert on request shape: both requests carried the same Idempotency-Key, and the mock saw one logical action, not two. A fresh key on the retry, or a duplicate call, means you found a double-send before a customer did. The broader method for testing agents that call your APIs sets up the harness end to end.

Recovery also needs a surface where a stuck run becomes somebody's problem. In Sharkly, a run that ends in a failed state is visible on the task itself, and the Inbox separates items needing your reply from routine updates. A follow-up comment on that task can start another run carrying the full history, so the retry is not a fresh prompt written from memory.

The error-recovery checklist

Before an agent goes to production, walk this list:

Tick all seven and your agent recovers on purpose instead of by luck.

Where Apidog fits (and where it doesn’t)

Keep the tool’s job honest. Apidog is not an agent framework, a model host, or a runtime. It does not build, run, or orchestrate your agent, and it does not grade the model’s output. What it owns is the API layer your agent calls, which is exactly where recovery is won or lost.

That gives it three jobs. It mocks the dependencies your agent hits, so you get a controllable stand-in instead of the live service. It programs the failure responses (429 with Retry-After, 500, timeout, malformed body) that a real API will not produce on command, so you can rehearse recovery. And it validates the requests the mock receives (idempotency key present and stable, correct shape, expected call count) so a double-send or a dropped header fails a test instead of a customer. That is the honest fit: Apidog mocks the failures your agent has to survive and checks what it sends back.

Frequently asked questions

Doesn’t the Anthropic SDK handle retries for me? For its own calls, yes. The SDK retries certain errors with exponential backoff and respects Retry-After, and you set the ceiling with a max-retries option. It does not cover the other APIs your agent’s tools call. Those need the same patterns applied by you.

When do I need an idempotency key? On any call that creates or changes state: charges, orders, sent messages, new records. Read-only calls are safe to retry without one. Generate the key once per action so it stays stable across retries.

Rehearse one failure this week

You don’t have to build all four patterns at once. Pick the one that would hurt most, usually the rate-limit loop or a non-idempotent retry, and rehearse it against a mock. Program the 429, drop a response, and watch what the agent sends. The first time you see clean backoff and a single idempotency key where you feared a double-charge, you will trust the agent for a better reason than a green demo.

Download Apidog to mock the failures, script the sequence, and assert on what your agent does when the API pushes back.

button

Explore more

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, and no shared benchmark

GPT-6.1 Sol vs Claude Sonnet 5.5: same $2/$10 price, launched a day apart, no shared benchmark. Specs, vendor claims, and how to test both yourself.

30 September 2026

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol Free?

Is GPT-6.1 Sol free? No: it's for Plus and up in ChatGPT Work and Codex, and the API has no free tier. Here are the cheapest routes, with real cost math.

30 September 2026

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

How to Use Claude Sonnet 5.5 for Free: Every Route That Works (and the Ones That Don't)

Is Claude Sonnet 5.5 free? Yes on Claude.ai (web, iOS, Android). Every free route checked, plus what isn't: Claude Code, the API, and Copilot Free.

29 September 2026

Practice API Design-first in Apidog

Discover an easier way to build and use APIs

AI Agent Error Recovery: Retry, Timeout, Backoff, and Circuit-Breaker Patterns