The outages that live beneath your application

Between a visitor typing your name and a request reaching your server sits a system you did not write and cannot review: a registration, a delegation, a zone, and a chain of caches. Here is how each of them fails.

article By 6 min read

Your application can be perfect and your site can be unreachable. Between a visitor typing a name and a request arriving at your server sits a system with no code in it. It has four parts: a registration with an expiry date, a delegation to a set of nameservers, a zone containing records, and a chain of resolver caches holding old answers until their TTL runs out.

When it breaks, it usually breaks completely, and it does so with no error for you to read — the request never reached anything of yours that logs.

Four failures, in order of how quietly they happen

A lapsed registration. A domain is a subscription. It renews on a clock measured in years, its reminders go to whichever address was on the registrant record at the time — often a person who has left, or a role alias that was never forwarded — and the card on file expires long before the domain does. When it lapses, the name stops resolving, or resolves to a registrar parking page, which is the more dangerous outcome, because it returns 200. Everything you host is gone from the internet and the certificate you renewed on schedule has nothing to be presented for.

A nameserver change nobody propagated. You move DNS providers. The new zone is built, the records look right, and the delegation at the registrar is updated — or half of it is. For the duration of the old delegation's TTL, resolvers are split: some ask the new provider, some ask the old, and the two disagree. The failure is partial and correlates with which resolver the visitor's network uses — so the bug reports arrive as "it works for me", truthfully, from everyone who says it, and you cannot reproduce it from any machine whose resolver happens to be on the right side.

A record pointing at a recycled IP. This is the one that deserves more attention than it gets. You tear down a server. The A record stays. That address goes back into the provider's pool and is allocated to somebody else, possibly within hours. Now your hostname resolves to a machine you have no relationship with. staging.example.com serves a stranger's site — and because your name still resolves and something answers with a 200, any check that asserts only "connected, status code below 400" is green. Worse, whoever now holds that address can satisfy an HTTP-based domain validation challenge for your hostname, because the challenge is served from whatever the record points at. They do not need your registrar account. They need only the machine your record still names.

A record that was right and stopped being. A CNAME to a platform that renamed its endpoints. A TXT record for domain validation that was cleaned up by someone tidying a zone. An MX record removed alongside a service migration, so the domain is fine and the mail is not — including, if you are unlucky, the mail your registrar sends about expiry.

Why this class is different

Three properties, shared by all four.

Not in the code, so review does not reach it. There is no diff and no test. The change happened in a registrar's console or a DNS provider's UI, possibly months ago, possibly by someone who no longer works with you.

Not local, so a development environment cannot reproduce it. Your machine has a resolver cache, an /etc/hosts file and a browser that has been open for days, and all three lie in the same direction.

Cached, so breakage and visibility are separated by a TTL you set long ago and have forgotten. A 24-hour TTL means a bad record published today starts failing tomorrow, for some people, in an order set by their resolvers.

The agent-adjacent part

None of this is caused by an agent. An agent cannot let your domain lapse and cannot change your delegation. That is why it is worth writing about: the registrar account and the zone were never in a file anyone was reading, and they do not become better understood because the code around them got easier to produce.

Concretely: an agent asked to move a site to a new host will change what needs changing in the repository and tell you what to point at it. It will not ask which of your other twelve A records still names a machine you own, because that question has no artefact. Nobody is being careless. There is nothing to review.

What to observe, and from where

Two checks cover most of it, and both have to run from outside your network, against public resolvers, on a schedule.

dns-resolution resolves the name the way a visitor's resolver does, and reports what came back. It catches the delegation that half-propagated, the record that disappeared, and the resolution that started failing while your own laptop kept serving a cached answer. It is deliberately upstream of everything else in a monitoring grid: if this is failing, no result below it means anything, and a set of checks that all fail together with dns-resolution at the top is a different sentence from a set that fail independently.

domain-expiry watches the registration itself, against a date known years in advance — which makes it the most preventable outage on this list. The alert is useful weeks out, not days. Auto-renewal and a calendar reminder do most of this job; the check exists because both of those depend on an account and a payment method nobody has looked at recently, and it is the only one of the three that fails loudly rather than quietly.

Two things this pair does not do, stated plainly. Neither tells you whether the address a record points at is still yours — resolution succeeding and the answer being correct are different questions, and the recycled-IP case satisfies the first while failing the second. Checking that is an inventory problem: what you own, what you resolve to, and the fact that those two lists should reconcile. And neither watches your registrar account, your renewal payment method, or who receives the reminder. Those live in an account nothing external can see.

For the recycled-IP case specifically, the closest external signal is content: dom-content against a hostname you expect to serve your page will notice when it starts serving somebody else's. That is a coarse instrument, and it only works on hostnames you thought to check.

The unglamorous conclusion

There is no clever mechanism here. This layer fails differently from your application because nothing about it is in version control: the changes are made in a web console, by whoever had the login, on a clock measured in years rather than deploys, and there is no artefact recording that they happened.

The defence is unglamorous too: something outside your infrastructure resolves your names regularly and tells you when the answer changes, and somebody can state the expiry date of every domain you own without logging in to find out. That is the difference between a five-minute fix and a weekend on the phone to a registrar's support queue, arguing about a name that has entered its redemption period.

Cite this

The HarpyWatch team. “The outages that live beneath your application”. The Nest, HarpyWatch, 10 September 2026 UTC. https://harpywatch.com/blog/beneath-the-application

More from The Nest

Staging configuration travels, and agents help it

A canonical tag pointing at staging, a noindex that shipped, a test API key in production. Each is invisible on a developer's machine and expensive live, and copying config between environments is the fastest way to make something work.

article 6 min read

Monitor the site nobody is looking at yet

The instinct is to add monitoring once a site has traffic worth protecting. That gets it backwards: an empty site is the one failure mode nobody reports, because your first hundred visitors leave instead.

article 6 min read

The API contract changed underneath the frontend and nothing failed

One agent edits the backend, another edits the frontend, a field quietly changes type, and every build stays green. The screen renders blank where a number belongs, and only something watching the deployed endpoint will tell you.

article 6 min read