---
title: "Reading a site the way a machine does"
description: "Your browser smooths over a great deal of what your site actually sends. Here are the terminal commands that show you what a crawler, a link checker or a monitoring probe gets instead."
url: https://harpywatch.com/blog/reading-a-site-the-way-a-machine-does
kind: article
author: "The HarpyWatch team"
published: 2026-09-15T09:00:00+00:00
tags: ["Deployment", "Monitoring", "SEO"]
publisher: "HarpyWatch"
---

# Reading a site the way a machine does

Open your site in a browser and it looks fine. That is the browser's job. It
follows redirects without telling you, fills in a missing certificate
intermediate from a cache of ones it has already seen, runs the JavaScript that
assembles the page, and recovers from malformed HTML without saying so anywhere
you will look. Some of that shows up in developer tools if you go looking.
Almost none of it shows up in the window.

Much of what visits your site does less. Some crawlers render JavaScript, on a
second pass and not always; plenty take the first response and stop. A link
checker reads a status code. A server-to-server callback often runs against a
trust store that has never seen your site, so nothing is cached for it.

Here is how to look at your own site the way they do. Every command below runs
in a normal terminal and needs no account. Replace `example.com` with yours.

## The first response, with nothing helping it

Start with the raw exchange. `-s` silences the progress meter, `-D -` dumps the
headers, `-o /dev/null` throws the body away.

```
curl -s -D - -o /dev/null https://example.com/
```

Note what you did *not* pass: no `-L`. curl does not follow redirects unless
you ask, which is the point — you are seeing the first thing your server says,
which is what a lot of software acts on.

Read the status line first. A `200` on the front page is unremarkable. A `301`
where you expected a `200` means every client that does not follow redirects is
one hop short of your content, and every link pointing here spends a round trip
it did not need to.

## The redirect chain, one hop at a time

Every hop costs a round trip and gives something else a chance to go wrong:
headers dropped, query strings eaten, a cookie set on one host and not sent to
the next. `http://` to `https://` to `www.` to a trailing slash is four requests
to serve one page, and nobody built it on purpose — each hop was added
separately and correctly.

```
curl -sIL -o /dev/null -w '%{num_redirects} hops -> %{url_effective} (%{http_code})\n' http://example.com
```

To see the hops themselves rather than the count:

```
curl -sIL http://example.com | grep -Ei '^(HTTP/|location:)'
```

Start from `http://`, not `https://`. That is where visitors, old bookmarks and
printed URLs start, and it is the version of the chain that gets tested least.

Two failures to look for. A chain that never terminates — A redirects to B
redirects to A — is invisible to any check that follows redirects and reports
only the final status, because there is no final status. And a redirect that
takes its destination from a query parameter will send a visitor anywhere:

```
curl -sI 'https://example.com/login?next=https://example.org/' | grep -i '^location:'
```

If the `Location` header comes back pointing at `example.org`, you have an open
redirect. It is a phishing vector wearing your domain, and it is the default
implementation of "return to the page you came from".

## What the headers admit

The header block is a statement about how your site expects to be treated. Read
it as one.

```
curl -sI https://example.com/ | sort
```

Worth noticing: whether `Cache-Control` on an HTML document outlasts your
deploy cycle, whether cookies carry `Secure`, `HttpOnly` and `SameSite`, whether
`Content-Security-Policy` exists at all, and whether
`Access-Control-Allow-Origin` says `*` on anything returning private data.

Compression deserves its own command, because you have to ask for it:

```
curl -sI -H 'Accept-Encoding: gzip, br' https://example.com/ | grep -i content-encoding
```

No `Content-Encoding` line means the response came back uncompressed. That is
usually not a config change anyone made — it is a proxy inserted in front, or a
`Content-Type` the compressor does not recognise, and it shows up as a site
that got slower for no visible reason.

## The page without JavaScript

This is the one that surprises people most.

```
curl -s https://example.com/ | wc -c
curl -s https://example.com/ | grep -o '<title>[^<]*' 
```

What you get here is what a client that does not execute JavaScript receives.
For a server-rendered site that is the page. For a client-rendered one it is a
`<div id="root">` and a script tag, and every crawler that does not run a full
browser sees an empty document with your title on it.

That is not automatically a problem. It is a problem if you assumed otherwise.

## The files you wrote once and never looked at again

```
curl -s https://example.com/robots.txt
curl -s https://example.com/sitemap.xml | head -40
```

`Disallow: /` shipping to production is expensive out of all proportion to its
size, because it works exactly as written: crawlers stop fetching, pages drop
out of the index over the following weeks, the site stays up the whole time, and
the cause is a two-line file nobody has opened since launch.

For the sitemap, check two things beyond it parsing: that the URLs in it are
the canonical ones, and that they still exist. Pull a handful and ask:

```
curl -s https://example.com/sitemap.xml \
  | grep -o '<loc>[^<]*' | sed 's/<loc>//' | head -20 \
  | xargs -I{} curl -s -o /dev/null -w '%{http_code} {}\n' {}
```

Anything that is not a `200` is a page you are actively telling search engines
to fetch and which is not there.

## What the TLS handshake actually presents

Your browser will not tell you your chain is incomplete, because it probably
has the missing intermediate cached from another site.

```
echo | openssl s_client -connect example.com:443 -servername example.com 2>/dev/null \
  | openssl x509 -noout -subject -issuer -dates
```

That gives you the leaf: who it is for, who signed it, and the validity window.
For the chain as sent, drop the second `openssl` and read the `Certificate
chain` block at the top of the output: each entry's `i:` line should be the next
entry's `s:` line, up to a root the client trusts. If it stops early, your site
works for you and fails for a freshly installed client.

The `-servername` flag is not optional. Without it you are not sending SNI, and
on shared infrastructure you will be handed a completely different
certificate than a real visitor gets.

To check what the server will actually negotiate rather than what it prefers:

```
openssl s_client -connect example.com:443 -servername example.com -tls1_2 </dev/null 2>&1 | grep -E 'Protocol|Cipher'
```

Swap `-tls1_2` for `-tls1_3`, or `-tls1_1` if you want to confirm the old ones
are refused.

## And resolution, underneath all of it

```
dig +short example.com A
dig +short example.com CNAME
```

DNS fails independently of your servers, and its failures are total: an
NXDOMAIN or a SERVFAIL means no client reaches anything, however healthy the
servers are. It is also the layer most likely to be holding a stale record from
a migration — pointing at an address that still answers, with something older
than you think behind it.

## Doing this once tells you about one moment

Everything above is a snapshot, from one machine on one network, of one second.

Run it against a staging hostname before you point DNS at it and you find the
broken redirect, the missing header and the sitemap entry that 404s while they
still cost nothing to fix.

None of these failures announce themselves when they start, though. A chain
shortens on a renewal nobody attended. A header disappears when someone puts a
proxy in front. A sitemap goes stale when a route is renamed. The gap between
the last time you ran these commands and the next is the window in which all of
that is true and unobserved.

Running them on a schedule, from outside your infrastructure, and being told
when an answer changes — that is what monitoring is, mechanically.

A note on what we build, since it is the same thing: HarpyWatch runs this set
and a couple of dozen more against your hostnames and puts the answers in one
grid. There is nothing in it you have not just run by hand. The difference is
the schedule and the memory of the previous answer.

Run the commands first. If the answers are boring, you have learned something
worth knowing.
