---
title: "Four hops to nowhere: what refactors do to your redirects"
description: "Every restructure moves URLs, and every move adds a redirect. Nobody ever removes one. The result is a chain a crawler stops following and a browser spends a second on before it sees a single byte of content."
url: https://harpywatch.com/blog/four-hops-to-nowhere
kind: article
author: "The HarpyWatch team"
published: 2026-09-22T09:00:00+00:00
tags: ["Deployment", "Monitoring", "SEO"]
publisher: "HarpyWatch"
---

# Four hops to nowhere: what refactors do to your redirects

A link in a two-year-old newsletter is clicked, and here is what happens before
the reader sees anything:

```
http://example.com/blog/post-name
  → https://example.com/blog/post-name      (301, http to https)
  → https://www.example.com/blog/post-name  (301, canonical host)
  → https://www.example.com/articles/post-name    (301, section rename)
  → https://www.example.com/articles/post-name/   (301, trailing slash)
  → 404
```

Five requests. Five DNS-and-connection round trips in the worst case. And the
destination does not exist, because the article was retired in a content cull
and the redirect rule was written against a pattern rather than a list.

Four redirects, added eighteen months apart, each correct when it was written,
and none ever deleted — because deleting a redirect is how you break the thing
it was protecting. Chains are an emergent property: whoever added the fourth
was not looking at the first three, and no file anywhere shows all four
together.

A coding agent changes the rate. Asked to rename a route, it will add the
redirect, which is more than most humans manage — forgetting is the older
failure. But it adds it at whichever layer it happens to be editing: the
framework's route table, or nginx, or the CDN config. Three layers, no single
file that shows the whole picture, and nothing in the task that prompts it to
check whether the target of its new redirect is itself a redirect.

## The four shapes, and what each one does

**The chain.** Two or more hops to reach content. Every hop is a full request
and, when it crosses an origin, a fresh DNS lookup, TCP connection and TLS
handshake. That is several round trips before any bytes of content, so the cost
scales with latency rather than bandwidth: on a mobile connection with 100ms
round-trip time, a cross-origin hop costs several such
round trips on its own, and a faster pipe does not help.
Google's documentation says Googlebot follows up to ten redirect hops in a
single crawl attempt and advises keeping chains short; treating three or more
as a defect is the conventional reading and it is sound. How much link signal survives a
long chain is disputed, and hard for you to measure either way — which is
itself the argument. Point the link at the final destination and the question
never arises.

**The loop.** `/a` redirects to `/b`, `/b` redirects to `/a`. Browsers stop
after a fixed limit — Chrome at twenty hops — and show
`ERR_TOO_MANY_REDIRECTS`. A crawler gives up and drops the URL. This is a total outage for that page and it is
invisible to anyone whose browser has cached an earlier answer, which
frequently includes the person who deployed it. The classic cause is two rules
in two layers disagreeing: nginx forcing a trailing slash while the application
strips it.

**The mixed http/https hop.** A chain that dips into http, even for one hop,
sends that request in the clear. Any header on it — a cookie without `Secure`,
a session token in a query parameter — is on the wire. It also gives anyone on the path — a
compromised router, a hostile access point — the chance to rewrite that
response and send the browser somewhere else entirely, before HSTS has had a
chance to apply.
The usual shape is `http://old.example` → `http://new.example` →
`https://new.example`: someone fixed the host and left the scheme to the layer
below, which is one hop too late.

**The redirect to a 404.** The most common one, and the most damaging. A
pattern-based rule (`/blog/(.*)` → `/articles/$1`) is written when the sections
matched. Then articles get retired, and the rule cheerfully maps a live inbound
link to a page that does not exist. A 301 into a 404 is worse for you than a plain 404 would have been: the
original URL has been declared moved, and the destination it was moved to
returns nothing. You have retired the old page without producing a new one.

Here is how the treatment differs, roughly:

| Shape | What a browser does | What a search engine does |
| --- | --- | --- |
| One hop, 301 | Follows, caches the redirect | Passes signals, indexes the target, eventually replaces the old URL |
| Chain of 3+ | Follows, pays the latency | Follows, but crawl budget is spent and Google's own advice is to shorten it |
| Loop | `ERR_TOO_MANY_REDIRECTS` | Drops the URL |
| Any http hop | Follows, request sent in the clear | Follows; the insecure hop is a security finding, not an SEO one |
| 301 into 404 | Shows the 404 | Deindexes the old URL, finds nothing at the new one |

One more distinction, routinely got wrong: a **302** says "keep the old URL,
this is temporary". Six months of a 302 that was always meant to be permanent
means the old URL is still the indexed one. Generated code uses whatever the
framework's `redirect()` helper defaults to, which is often 302. Check yours.

## Why you cannot see this from inside

You can read your nginx config, your route table and your CDN rules and still
not know what happens to a request, because the answer is the composition of
all three plus whatever the CDN does with trailing slashes on its own
initiative. You can reconstruct it from edge logs afterwards, if you have them
and the patience. Making the request from outside and following it answers the
question directly.

That is what HarpyWatch's `redirect-chain` check does — it is our product, so
weigh the recommendation accordingly. It starts at a URL, follows every hop,
records each status code and location, and reports the full path, the number of
hops, whether any hop was insecure, and whether the terminus is a loop or an
error. Its companion `broken-links` catches the last shape in the table: it
crawls what your pages link to and reports what does not resolve, including the
links that resolve through three hops into a 404 that clicking around the site
would never reach. Neither is a clever check. Both answer a question that is easy to assume you
already know the answer to.

The limit is real: an external check follows the links it can discover from a
public page, plus whatever URLs you give it explicitly. It does not know about
the inbound link in somebody else's PDF or the QR code on a printed flyer —
exactly the class of link redirects exist to protect, and the class most likely
to point at a URL no one has considered since 2021. For those you have only the
redirect configuration itself, which is an argument for keeping it in one
readable place rather than in three.

## Do it before DNS moves

A restructure is a scheduled event. The URLs that will move are known before
the deploy, and so are the chains they will create: both are properties of
configuration that already exists in a branch.

So put the new site on a staging hostname and run the checks against that
before you point anything at it. Take your most-linked URLs, run them through,
and look at the hop counts. A chain found at that moment costs one line in a
config file. The same chain found a month later costs whatever the traffic to
those URLs was worth — and you will not be able to measure that, because
traffic that did not arrive leaves no record anywhere.
