The API contract changed underneath the frontend and nothing failed
One agent edits the backend, another edits the frontend, a field quietly changes type, and every build stays green. The screen renders blank where a number belongs, and only something watching the deployed endpoint will tell you.
uptime_percent used to be a number. It is now a string, because a Rust
serialiser gained a format!("{:.1}", pct) to round it to one decimal place,
and format! produces text.
The backend tests pass: they assert on the value, and "99.9" is the value. The
frontend builds: TypeScript is describing a response it was told about, not one
it received. The page renders. In the cell where a percentage belongs there
is now 99.9%%, or NaN%, or nothing at all, depending on which arithmetic ran
first.
Nothing in either build raised a hand.
Why the build cannot save you here
The failure has a precise location: the boundary where a typed program stops being able to check its own beliefs.
Inside the backend, types are enforced. Inside the frontend, types are enforced.
Between them is JSON, which enforces nothing. The frontend's
interface UptimeSummary is not a check on anything — it is a comment the type
checker happens to trust. When the two sides disagree, the disagreement lives in
the one place neither language is looking.
Every failure in this family is the same shape:
- A field changed type: number to string, string to enum object, scalar to array of one.
- A field was renamed:
openIncidentsbecameopen_incidents, so the old name is nowundefinedandundefinedrenders as an empty cell. - A field became optional and is now absent on some rows — usually the interesting rows, the ones where a check has not run.
- A field got nested:
{ latency: 42 }became{ latency: { p50: 42, p95: 81 } }, and the template now stringifies an object. - An array became paginated: what was
[...]is now{ items: [...], next: null }, and.mapis not a function.
None of these throw on the server. Most do not throw on the client either.
JavaScript's defining characteristic in this situation is that it keeps going.
An absent value propagates through template interpolation and arrives on screen
as blank, as the literal word undefined, or as NaN once it has been through
a calculation. The page looks like a page. It is simply wrong.
What agents changed about the odds
This has always happened, in every codebase with a client and a server written by different people. What changes when the writing is done by agents is that the two sides of a contract can now be edited hours apart by processes that share no memory, and the diff each produces is small enough to approve without opening the other side.
When the API and the client are edited from two separate sessions, neither session holds both sides in context. Each makes a locally correct change. One was asked to round a percentage and did so cleanly, with a test. The other was asked to add a sparkline and did so cleanly, with a test. Nothing anywhere contained both facts at once.
It happens with one agent across two sittings for the same reason: the context that knew the response shape ended when the session did.
And the pull request looks good. A diff adding format!("{:.1}", pct) is small,
tidy, well-named and obviously improving something. It gets approved quickly,
because the question a reviewer asks is "is this change correct" and the answer
is yes. The question that would have caught it — "who else believes this field
is a number" — needs the whole set of consumers in view.
There is machinery for that. A checked-in OpenAPI or GraphQL schema with generated clients moves the disagreement to build time, and if you have it, use it — it is strictly better than what follows. The gap is that it only covers the endpoints described by the schema, and it is a statement about the code in the repository rather than the pair of services currently deployed.
Integration tests test the version you had
The usual answer is contract testing, and it is a good answer that has a specific hole.
An integration test asserts that the API and the client agree at the moment CI ran, against the code in that branch. That is genuinely valuable. It is also a claim about a repository, not about a deployment.
The things it does not cover are exactly the things that bite:
- The API is deployed and the frontend is not, or the reverse. The two versions agreed in the repository and do not agree in production.
- The response shape depends on data. The test fixture has a monitor with results; production has one that has never run, and that row is the one with the null.
- Something upstream changed. A third-party API you proxy altered its own payload, and your serialiser passed the change straight through.
- A cache, a proxy, or a CDN is serving a body from before the deploy.
- A rollback restored an older API against a newer client.
In every case CI was green and the deployed pair disagree. The assertion has to run against the running system, on a schedule, or it is answering a question about the past.
Assert the shape, continuously
The correction is unglamorous, and it is detection rather than prevention: check that the deployed endpoint still returns what the deployed client expects, repeatedly, from outside. It will not stop the deploy. It will tell you within one check interval instead of within one customer email.
HarpyWatch does this with two checks that do different amounts of work.
json-api requests an endpoint and asserts on the response body itself —
that a path exists, that a value matches, that a field is present. It is the
right tool when there are two or three facts about a response you actually care
about and you do not want to describe the rest of it.
json-schema validates the whole body against a JSON Schema you supply.
This is the one that catches the type change, because a schema says
"type": "number" and "99.9" is not one. It also catches a field going
missing when the schema marks it required, and a field going nested when the
schema says it is a scalar. The check fails at the endpoint that changed rather
than three services downstream, which is what makes the failure legible instead
of a hunt.
You do not need a schema for every endpoint. You need one for the handful whose breakage would render a screen wrong without erroring — the summary payloads, the counts, the anything-a-dashboard-draws. For most products that is a handful of endpoints, not the whole API surface.
Two honest limits
First, a schema is a thing you wrote, and it can rot. If the API legitimately gains a field and nobody updates the schema, you get either a false failure or, if the schema is permissive, a silent gap. Schemas need to live next to the code that produces the response and change with it.
Second, this catches shape, not meaning. uptime_percent: 0 is a perfectly
valid number and may be catastrophically wrong. No schema will tell you that a
value is stale or fabricated — only that it is the right type. Shape
validation is the floor, not the ceiling.
Both limits are worth accepting, because the failure being prevented is the one that produces a confident, populated, entirely wrong screen — and nothing in either codebase's type system is going to notice it.
Cite this
The HarpyWatch team. “The API contract changed underneath the frontend and nothing failed”. The Nest, HarpyWatch, 24 September 2026 UTC. https://harpywatch.com/blog/the-contract-changed-underneath
More from The Nest
Staging configuration travels, and agents help it
A canonical tag pointing at staging, a noindex that shipped, a test API key in production. Each is invisible on a developer's machine and expensive live, and copying config between environments is the fastest way to make something work.
Monitor the site nobody is looking at yet
The instinct is to add monitoring once a site has traffic worth protecting. That gets it backwards: an empty site is the one failure mode nobody reports, because your first hundred visitors leave instead.
Four hops to nowhere: what refactors do to your redirects
Every restructure moves URLs, and every move adds a redirect. Nobody ever removes one. The result is a chain a crawler stops following and a browser spends a second on before it sees a single byte of content.