Reading a site the way a machine does
Your browser smooths over a great deal of what your site actually sends. Here are the terminal commands that show you what a crawler, a link checker or a monitoring probe gets instead.
Open your site in a browser and it looks fine. That is the browser's job. It follows redirects without telling you, fills in a missing certificate intermediate from a cache of ones it has already seen, runs the JavaScript that assembles the page, and recovers from malformed HTML without saying so anywhere you will look. Some of that shows up in developer tools if you go looking. Almost none of it shows up in the window.
Much of what visits your site does less. Some crawlers render JavaScript, on a second pass and not always; plenty take the first response and stop. A link checker reads a status code. A server-to-server callback often runs against a trust store that has never seen your site, so nothing is cached for it.
Here is how to look at your own site the way they do. Every command below runs
in a normal terminal and needs no account. Replace example.com with yours.
The first response, with nothing helping it
Start with the raw exchange. -s silences the progress meter, -D - dumps the
headers, -o /dev/null throws the body away.
curl -s -D - -o /dev/null https://example.com/
Note what you did not pass: no -L. curl does not follow redirects unless
you ask, which is the point — you are seeing the first thing your server says,
which is what a lot of software acts on.
Read the status line first. A 200 on the front page is unremarkable. A 301
where you expected a 200 means every client that does not follow redirects is
one hop short of your content, and every link pointing here spends a round trip
it did not need to.
The redirect chain, one hop at a time
Every hop costs a round trip and gives something else a chance to go wrong:
headers dropped, query strings eaten, a cookie set on one host and not sent to
the next. http:// to https:// to www. to a trailing slash is four requests
to serve one page, and nobody built it on purpose — each hop was added
separately and correctly.
curl -sIL -o /dev/null -w '%{num_redirects} hops -> %{url_effective} (%{http_code})\n' http://example.com
To see the hops themselves rather than the count:
curl -sIL http://example.com | grep -Ei '^(HTTP/|location:)'
Start from http://, not https://. That is where visitors, old bookmarks and
printed URLs start, and it is the version of the chain that gets tested least.
Two failures to look for. A chain that never terminates — A redirects to B redirects to A — is invisible to any check that follows redirects and reports only the final status, because there is no final status. And a redirect that takes its destination from a query parameter will send a visitor anywhere:
curl -sI 'https://example.com/login?next=https://example.org/' | grep -i '^location:'
If the Location header comes back pointing at example.org, you have an open
redirect. It is a phishing vector wearing your domain, and it is the default
implementation of "return to the page you came from".
What the headers admit
The header block is a statement about how your site expects to be treated. Read it as one.
curl -sI https://example.com/ | sort
Worth noticing: whether Cache-Control on an HTML document outlasts your
deploy cycle, whether cookies carry Secure, HttpOnly and SameSite, whether
Content-Security-Policy exists at all, and whether
Access-Control-Allow-Origin says * on anything returning private data.
Compression deserves its own command, because you have to ask for it:
curl -sI -H 'Accept-Encoding: gzip, br' https://example.com/ | grep -i content-encoding
No Content-Encoding line means the response came back uncompressed. That is
usually not a config change anyone made — it is a proxy inserted in front, or a
Content-Type the compressor does not recognise, and it shows up as a site
that got slower for no visible reason.
The page without JavaScript
This is the one that surprises people most.
curl -s https://example.com/ | wc -c
curl -s https://example.com/ | grep -o '<title>[^<]*'
What you get here is what a client that does not execute JavaScript receives.
For a server-rendered site that is the page. For a client-rendered one it is a
<div id="root"> and a script tag, and every crawler that does not run a full
browser sees an empty document with your title on it.
That is not automatically a problem. It is a problem if you assumed otherwise.
The files you wrote once and never looked at again
curl -s https://example.com/robots.txt
curl -s https://example.com/sitemap.xml | head -40
Disallow: / shipping to production is expensive out of all proportion to its
size, because it works exactly as written: crawlers stop fetching, pages drop
out of the index over the following weeks, the site stays up the whole time, and
the cause is a two-line file nobody has opened since launch.
For the sitemap, check two things beyond it parsing: that the URLs in it are the canonical ones, and that they still exist. Pull a handful and ask:
curl -s https://example.com/sitemap.xml \
| grep -o '<loc>[^<]*' | sed 's/<loc>//' | head -20 \
| xargs -I{} curl -s -o /dev/null -w '%{http_code} {}\n' {}
Anything that is not a 200 is a page you are actively telling search engines
to fetch and which is not there.
What the TLS handshake actually presents
Your browser will not tell you your chain is incomplete, because it probably has the missing intermediate cached from another site.
echo | openssl s_client -connect example.com:443 -servername example.com 2>/dev/null \
| openssl x509 -noout -subject -issuer -dates
That gives you the leaf: who it is for, who signed it, and the validity window.
For the chain as sent, drop the second openssl and read the Certificate chain block at the top of the output: each entry's i: line should be the next
entry's s: line, up to a root the client trusts. If it stops early, your site
works for you and fails for a freshly installed client.
The -servername flag is not optional. Without it you are not sending SNI, and
on shared infrastructure you will be handed a completely different
certificate than a real visitor gets.
To check what the server will actually negotiate rather than what it prefers:
openssl s_client -connect example.com:443 -servername example.com -tls1_2 </dev/null 2>&1 | grep -E 'Protocol|Cipher'
Swap -tls1_2 for -tls1_3, or -tls1_1 if you want to confirm the old ones
are refused.
And resolution, underneath all of it
dig +short example.com A
dig +short example.com CNAME
DNS fails independently of your servers, and its failures are total: an NXDOMAIN or a SERVFAIL means no client reaches anything, however healthy the servers are. It is also the layer most likely to be holding a stale record from a migration — pointing at an address that still answers, with something older than you think behind it.
Doing this once tells you about one moment
Everything above is a snapshot, from one machine on one network, of one second.
Run it against a staging hostname before you point DNS at it and you find the broken redirect, the missing header and the sitemap entry that 404s while they still cost nothing to fix.
None of these failures announce themselves when they start, though. A chain shortens on a renewal nobody attended. A header disappears when someone puts a proxy in front. A sitemap goes stale when a route is renamed. The gap between the last time you ran these commands and the next is the window in which all of that is true and unobserved.
Running them on a schedule, from outside your infrastructure, and being told when an answer changes — that is what monitoring is, mechanically.
A note on what we build, since it is the same thing: HarpyWatch runs this set and a couple of dozen more against your hostnames and puts the answers in one grid. There is nothing in it you have not just run by hand. The difference is the schedule and the memory of the previous answer.
Run the commands first. If the answers are boring, you have learned something worth knowing.
Cite this
The HarpyWatch team. “Reading a site the way a machine does”. The Nest, HarpyWatch, 15 September 2026 UTC. https://harpywatch.com/blog/reading-a-site-the-way-a-machine-does
More from The Nest
Staging configuration travels, and agents help it
A canonical tag pointing at staging, a noindex that shipped, a test API key in production. Each is invisible on a developer's machine and expensive live, and copying config between environments is the fastest way to make something work.
Monitor the site nobody is looking at yet
The instinct is to add monitoring once a site has traffic worth protecting. That gets it backwards: an empty site is the one failure mode nobody reports, because your first hundred visitors leave instead.
The API contract changed underneath the frontend and nothing failed
One agent edits the backend, another edits the frontend, a field quietly changes type, and every build stays green. The screen renders blank where a number belongs, and only something watching the deployed endpoint will tell you.