Four hops to nowhere: what refactors do to your redirects
Every restructure moves URLs, and every move adds a redirect. Nobody ever removes one. The result is a chain a crawler stops following and a browser spends a second on before it sees a single byte of content.
A link in a two-year-old newsletter is clicked, and here is what happens before the reader sees anything:
http://example.com/blog/post-name
→ https://example.com/blog/post-name (301, http to https)
→ https://www.example.com/blog/post-name (301, canonical host)
→ https://www.example.com/articles/post-name (301, section rename)
→ https://www.example.com/articles/post-name/ (301, trailing slash)
→ 404
Five requests. Five DNS-and-connection round trips in the worst case. And the destination does not exist, because the article was retired in a content cull and the redirect rule was written against a pattern rather than a list.
Four redirects, added eighteen months apart, each correct when it was written, and none ever deleted — because deleting a redirect is how you break the thing it was protecting. Chains are an emergent property: whoever added the fourth was not looking at the first three, and no file anywhere shows all four together.
A coding agent changes the rate. Asked to rename a route, it will add the redirect, which is more than most humans manage — forgetting is the older failure. But it adds it at whichever layer it happens to be editing: the framework's route table, or nginx, or the CDN config. Three layers, no single file that shows the whole picture, and nothing in the task that prompts it to check whether the target of its new redirect is itself a redirect.
The four shapes, and what each one does
The chain. Two or more hops to reach content. Every hop is a full request and, when it crosses an origin, a fresh DNS lookup, TCP connection and TLS handshake. That is several round trips before any bytes of content, so the cost scales with latency rather than bandwidth: on a mobile connection with 100ms round-trip time, a cross-origin hop costs several such round trips on its own, and a faster pipe does not help. Google's documentation says Googlebot follows up to ten redirect hops in a single crawl attempt and advises keeping chains short; treating three or more as a defect is the conventional reading and it is sound. How much link signal survives a long chain is disputed, and hard for you to measure either way — which is itself the argument. Point the link at the final destination and the question never arises.
The loop. /a redirects to /b, /b redirects to /a. Browsers stop
after a fixed limit — Chrome at twenty hops — and show
ERR_TOO_MANY_REDIRECTS. A crawler gives up and drops the URL. This is a total outage for that page and it is
invisible to anyone whose browser has cached an earlier answer, which
frequently includes the person who deployed it. The classic cause is two rules
in two layers disagreeing: nginx forcing a trailing slash while the application
strips it.
The mixed http/https hop. A chain that dips into http, even for one hop,
sends that request in the clear. Any header on it — a cookie without Secure,
a session token in a query parameter — is on the wire. It also gives anyone on the path — a
compromised router, a hostile access point — the chance to rewrite that
response and send the browser somewhere else entirely, before HSTS has had a
chance to apply.
The usual shape is http://old.example → http://new.example →
https://new.example: someone fixed the host and left the scheme to the layer
below, which is one hop too late.
The redirect to a 404. The most common one, and the most damaging. A
pattern-based rule (/blog/(.*) → /articles/$1) is written when the sections
matched. Then articles get retired, and the rule cheerfully maps a live inbound
link to a page that does not exist. A 301 into a 404 is worse for you than a plain 404 would have been: the
original URL has been declared moved, and the destination it was moved to
returns nothing. You have retired the old page without producing a new one.
Here is how the treatment differs, roughly:
| Shape | What a browser does | What a search engine does |
|---|---|---|
| One hop, 301 | Follows, caches the redirect | Passes signals, indexes the target, eventually replaces the old URL |
| Chain of 3+ | Follows, pays the latency | Follows, but crawl budget is spent and Google's own advice is to shorten it |
| Loop | ERR_TOO_MANY_REDIRECTS |
Drops the URL |
| Any http hop | Follows, request sent in the clear | Follows; the insecure hop is a security finding, not an SEO one |
| 301 into 404 | Shows the 404 | Deindexes the old URL, finds nothing at the new one |
One more distinction, routinely got wrong: a 302 says "keep the old URL,
this is temporary". Six months of a 302 that was always meant to be permanent
means the old URL is still the indexed one. Generated code uses whatever the
framework's redirect() helper defaults to, which is often 302. Check yours.
Why you cannot see this from inside
You can read your nginx config, your route table and your CDN rules and still not know what happens to a request, because the answer is the composition of all three plus whatever the CDN does with trailing slashes on its own initiative. You can reconstruct it from edge logs afterwards, if you have them and the patience. Making the request from outside and following it answers the question directly.
That is what HarpyWatch's redirect-chain check does — it is our product, so
weigh the recommendation accordingly. It starts at a URL, follows every hop,
records each status code and location, and reports the full path, the number of
hops, whether any hop was insecure, and whether the terminus is a loop or an
error. Its companion broken-links catches the last shape in the table: it
crawls what your pages link to and reports what does not resolve, including the
links that resolve through three hops into a 404 that clicking around the site
would never reach. Neither is a clever check. Both answer a question that is easy to assume you
already know the answer to.
The limit is real: an external check follows the links it can discover from a public page, plus whatever URLs you give it explicitly. It does not know about the inbound link in somebody else's PDF or the QR code on a printed flyer — exactly the class of link redirects exist to protect, and the class most likely to point at a URL no one has considered since 2021. For those you have only the redirect configuration itself, which is an argument for keeping it in one readable place rather than in three.
Do it before DNS moves
A restructure is a scheduled event. The URLs that will move are known before the deploy, and so are the chains they will create: both are properties of configuration that already exists in a branch.
So put the new site on a staging hostname and run the checks against that before you point anything at it. Take your most-linked URLs, run them through, and look at the hop counts. A chain found at that moment costs one line in a config file. The same chain found a month later costs whatever the traffic to those URLs was worth — and you will not be able to measure that, because traffic that did not arrive leaves no record anywhere.
Cite this
The HarpyWatch team. “Four hops to nowhere: what refactors do to your redirects”. The Nest, HarpyWatch, 22 September 2026 UTC. https://harpywatch.com/blog/four-hops-to-nowhere
More from The Nest
Staging configuration travels, and agents help it
A canonical tag pointing at staging, a noindex that shipped, a test API key in production. Each is invisible on a developer's machine and expensive live, and copying config between environments is the fastest way to make something work.
Monitor the site nobody is looking at yet
The instinct is to add monitoring once a site has traffic worth protecting. That gets it backwards: an empty site is the one failure mode nobody reports, because your first hundred visitors leave instead.
The API contract changed underneath the frontend and nothing failed
One agent edits the backend, another edits the frontend, a field quietly changes type, and every build stays green. The screen renders blank where a number belongs, and only something watching the deployed endpoint will tell you.