The API team renamed a field from customer_name to customer_full_name. They announced it, they updated the docs, and every human-maintained client got a pull request. Your agent got nothing, because nobody thought of it as a client. It kept sending the old field, the API kept accepting the request and ignoring the unknown key, and for two weeks every record it created had an empty name.
Agents are the API consumers least able to notice a change and most likely to paper over it. A human client throws an exception. An agent reads a 200, decides the call worked, and moves on. Sometimes it improvises around the problem in a way that looks like success.
This guide covers why agents are unusually fragile to API drift, which changes break them that would not break ordinary clients, how to pin and detect versions, and how to catch drift in CI before a run does. Our post on why AI agents break in production covers the failure modes; this is the one that arrives from outside your codebase.
Apidog matters here because detection is a spec problem. If you have the previous version of an API definition and the current one, the diff is mechanical.
Why agents notice less than clients do
Four properties combine badly.
Silent tolerance. Most APIs ignore unknown fields in a request body. A renamed field means the new one is absent and the old one is discarded, with a 200 on the way out. Nothing raises.
Improvisation. When a response is missing a value, a model will often continue with a plausible substitute rather than stopping. That is helpful behavior in conversation and dangerous behavior against an API.
Descriptions in the prompt. Agent tool descriptions encode assumptions about the API in text. When the API changes, the descriptions become subtly wrong, and wrong descriptions produce wrong calls without any code being involved. Our post on tool schema design covers how much behavior rides on that text.
No compiler. A typed client breaks at build time when a field disappears. An agent’s contract lives in JSON schemas and prose, and nothing checks it until a call fails, or worse, until one quietly does not.
The upshot: changes that are safe for typical clients are not always safe for agents, and you should classify them separately.
Which changes actually break agents
The usual additive-versus-breaking split still applies, and agents add a middle category.
Genuinely breaking, for everyone. Removing an endpoint, removing a field, renaming a field, changing a type, making an optional parameter required, changing the URL. Agents break here too, only more quietly.
Safe for typed clients, risky for agents:
- A new required field. Every existing caller breaks, but an agent breaks with a validation error it may try to fix by inventing a value. That is worse than a hard failure.
- A new enum value. Ordinary clients ignore what they do not handle. An agent may reason about the unfamiliar value and draw a conclusion your product never intended.
- A tightened validation rule. A field that used to accept any string now requires a pattern. The agent has no way to learn the pattern except by failing, which is why the rule belongs in the error message, as covered in our post on API error design for agents.
- A changed default. Pagination default drops from 100 to 20 and the agent, which never sent a limit, now sees a fifth of the data and reports on it as if complete.
- Reworded documentation. No behavior change at all, but if your tools are generated from the spec, as in our guide to turning an OpenAPI spec into agent tools, the description text changed and tool selection can shift with it.
Safe for agents too. Adding an optional field, adding an endpoint, adding an optional parameter with a preserved default, loosening validation.
That middle list is the one to watch, because nothing in a standard change review flags it.
Pin the version, always
The first defense is refusing to move implicitly.
Send an explicit version on every request, whatever mechanism the API offers: a path segment, a header, or an account-level pin. GitHub’s API versioning documentation uses a date header, and Stripe pins a version per account with an explicit upgrade step. Both give you the same property: nothing changes underneath you until you decide.
DEFAULT_HEADERS = {
"X-API-Version": "2026-06-01",
"User-Agent": "billing-agent/1.4 (+https://example.com/agents)",
}
The User-Agent is worth as much as the version pin. When an API provider needs to warn callers about a deprecation, they look at traffic. An agent that identifies itself gets the email; one that sends a default library string does not.
If you own the API, publish a version and hold it. Our guide to the best API versioning strategy covers the options, and managing API versioning in Apidog covers keeping several live at once.
For third-party APIs with no versioning at all, pin what you can: record the response shape you built against and check it, which is the next section.
Detect drift before a run does
Pinning buys time. It does not stop the eventual upgrade, and it does nothing for APIs that change without versioning. So detect.
Diff the spec on a schedule. If the provider publishes an OpenAPI document, fetch it daily and compare against the copy you generated tools from. Fields removed, types changed, requirements added, enums extended, descriptions edited. In Apidog you can keep the imported definition in the project and see what moved between versions, which turns “did anything change” into a report rather than an investigation.
Contract test the endpoints you call. For each tool the agent has, send a known-good request and assert on the response shape: required fields present, types correct, enum values within the set you expect. This catches drift in APIs that publish no spec at all, which is most of them. Our API contract testing guide covers the pattern, and bidirectional contract testing covers running it from both sides.
Assert on shape at runtime. Validate responses in the tool wrapper against the schema you expect, and log a warning when something unexpected appears. This is the last line, and it is the one that catches the change nobody announced.
def check_shape(tool_name, payload, expected):
missing = [f for f in expected["required"] if f not in payload]
extra = [f for f in payload if f not in expected["properties"]]
if missing:
log.error("api_drift", tool=tool_name, missing=missing)
raise ApiDriftError(f"{tool_name}: missing fields {missing}")
if extra:
log.warning("api_new_fields", tool=tool_name, fields=extra)
return payload
Fail on missing, warn on extra. A missing required field means the agent is about to work with incomplete data, which is the failure worth stopping for. New fields are usually additive and worth knowing about without interrupting a run. Route both into the trace record described in our post on tracing agent tool calls.
Watch behavior, not only schemas. Some drift is invisible to a shape check: a default that changed, a rate limit that tightened, a response that got slower. Track calls per completed task, retry rate per endpoint, and average response size per tool. A step change in any of them usually means something moved upstream.
Upgrading without breaking the agent
When you do move to a new version, treat it as a change to the agent, because it is.
Regenerate the tools rather than hand-editing them, so descriptions and schemas move together. Then read the diff of the generated tool definitions. That diff is the real blast radius, and it is often smaller or larger than the API changelog implies.
Run the agent against a mock of the new version before pointing it at anything live. This is the highest-value step and the one most often skipped: a mock built from the new spec lets you run your whole task suite against the new shapes with no risk, following our post on running agents against mocks instead of production.
Re-run the selection suite. Description changes shift which tool the model picks, and that regression is invisible to a schema diff. Assert on tool choice for a fixed set of prompts, as in our guide to testing non-deterministic agents.
Roll out behind a flag, on a slice of traffic, with the old version still pinned and ready. Watch the same four numbers for a day. Agent regressions show up as more calls per task and more retries long before anyone files a complaint.
Three drifts that made it to production
The renamed field. The opening story. A 200 on every call, empty names on every record, discovered two weeks later by a human reading a report. A runtime shape check on the response would have caught it on the first call, because the field the agent expected to read back was gone.
The tightened pagination default. A provider dropped the default page size from 100 to 20. The agent never sent a limit, so it started seeing 20 records and summarizing them as the complete set. Nothing errored. The summaries were simply wrong, in a way that read as confident. The fix was one line, sending an explicit limit, and the lesson is broader: rely on defaults and you have an undeclared dependency on someone else’s decision.
The new enum value. A payment API added status: "disputed". Typed clients ignored it. The agent reasoned about it, decided a disputed charge counted as a refund, and reported reconciled books that were not. Explicit enum validation would have raised on the unfamiliar value instead of letting the model interpret it.
The pattern: each change was announced, each was additive or minor by the provider’s own classification, and each was breaking for an agent. That gap is the thing to design around.
Treat deprecations as a work item
Providers usually warn you. The warning arrives in a changelog, an email, or a Deprecation header on the response, and it is easy for none of those to reach the person who maintains the agent.
Wire them into your normal queue. The Deprecation header and the Sunset header are both standardized, so a generic check works across providers. Log them when they appear, and alert on the first sighting rather than the thousandth. A header that shows up on 3 percent of calls today is a full outage on the sunset date.
Keep an inventory too: which agent, which provider, which version, which endpoints, and who owns it. Ten lines in a file is enough. When a deprecation notice arrives, the question “does this affect us” should take a minute, not an afternoon of grepping.
Drift is work, so give it an owner
Detection produces a queue: a spec diff, a failing contract test, a deprecation header seen for the first time. Each one is a small piece of work with a deadline attached, and the failure mode is that it sits in a channel nobody owns until the sunset date arrives.
Put them where your team already tracks work. If your agents run as coding runtimes rather than as a service you deployed, the platform managing them can close the loop: Sharkly assigns a Task to an Agent or a Crew and keeps the goal, the execution trace, and the review in one place, so “the payments API deprecated this endpoint” becomes an assigned task with a result rather than a message in a thread. Whatever you use, the rule is the same. A drift alert with no owner is a deprecation you will meet again on the day it breaks.

A checklist
- Every request sends an explicit API version and an identifying
User-Agent. - Third-party spec documents are fetched and diffed on a schedule.
- Every tool the agent can call has a contract test asserting response shape.
- Tool wrappers validate responses at runtime: fail on missing, warn on new.
- Behavioral metrics are tracked per endpoint so silent drift surfaces.
- Version upgrades regenerate tools rather than editing them by hand.
- The task suite and the selection suite both run against a mock of the new version first.
- Rollout is flagged and reversible, with the previous version still pinned.
The API team will keep shipping changes, and that is fine. What you need is for your agent to be a client that notices, which takes a version pin, a contract test, and a runtime shape check. Download Apidog to diff the spec and mock the next version before it reaches a live run.
Frequently asked questions
How often should I check a third-party spec for changes? Daily is enough for most, and cheap to automate. For APIs with no published spec, lean on contract tests running in CI instead, since they detect the same drift from the outside.
Should I always pin to the oldest working version? No. Pin so upgrades are deliberate, then upgrade on a schedule. Sitting on an old version until it is removed converts a planned change into an emergency.
What if the agent works fine after a change? Verify rather than assume. The dangerous outcomes are the ones that still return 200, such as a renamed field silently dropped. A shape assertion tells you what a green run cannot.
Do I need to version my own API differently for agents? Not differently, but more strictly. Treat new required fields, new enum values, and changed defaults as breaking for agent consumers even when they are additive for typed clients, and announce them the same way.
How do I know which agents call which endpoints? From your traces. Tool name plus endpoint per run gives you the dependency map, and it tells you exactly who is affected by a deprecation. Our post on tracing agent tool calls covers the record shape.
Can the agent adapt to a changed API on its own? Sometimes, and you should not rely on it. A model that improvises around a missing field produces plausible output with no signal that anything went wrong. Fail loudly and fix the tools instead.



