Set it up in an afternoon and then stop

Adopting monitoring for a small site is five decisions, not a project. Here is what each one actually involves, and why finishing a rough version beats designing a perfect one.

article By 6 min read

The most common state for a small site is not "badly monitored". It is "monitoring was on the list for eight months".

The reason is rarely cost or difficulty. It is that the task has no obvious edge. Once you start thinking about what to watch, the list grows — every route, every third party, every environment — and a job that was going to take an hour becomes a design exercise, and design exercises get postponed indefinitely in favour of things with deadlines.

So here is the version with an edge. Five decisions, in order, each one answerable in a few minutes. When you have made all five, stop. A rough setup that exists beats an excellent one that is still being scoped.

One: list the addresses that matter

Not the pages that matter. The addresses whose failure you would want to be woken for.

For most products it fits on one screen, and it is shorter than the first draft:

  • the marketing site's front page, because it is what a stranger sees
  • the application's front door, the URL people bookmark
  • the login or session endpoint, which fails independently of everything else
  • the two or three API routes a customer's integration actually calls
  • whatever a customer would check first if they suspected you were down

Write them out. If the list is over fifteen, you are listing pages rather than failures — cut back to the ones where the answer to "would I want to know at 2am?" is yes.

In HarpyWatch's model, the thing you create is an asset, and an asset is a hostname. Checks hang off it. If you run app.example.com, example.com and api.example.com, that is three assets, and they will fail separately, which is the whole reason they are separate rows.

Two: choose the checks, then choose two more

A check is one kind of observation against an asset on a schedule. Start with the ones you would have picked anyway:

  • http-uptime on each asset — status code and response time
  • ssl-cert-expiry and cert-chain-trust on anything served over HTTPS
  • dns-resolution, which is upstream of all of it and fails on its own
  • domain-expiry, once, on each registered name

That is the set anybody would have picked. Now add two you would not have thought to ask for, chosen by what your site actually is:

  • publishing anything on a feed? feed-freshness
  • care about search traffic? robots-policy and sitemap-health
  • an API other people call? json-schema and cors-policy
  • a page assembled by JavaScript? browser-render, which uses a real browser
  • anything realtime? websocket-handshake
  • anything you have ever hand-tuned for speed? compression and cache-policy

Two is enough. The point is not coverage. It is that the failures which sit unnoticed longest are the ones that leave a healthy 200 in place — a feed that stopped updating in March, a certificate chain missing its intermediate, a Disallow: / nobody has opened the file to see — and none of those will ever occur to you as something to ask for.

Three: pick an interval you will not resent

This is the decision people get wrong, and they get it wrong in the direction of ambitious.

The interval is not a quality setting. It is the maximum time a failure can exist before anyone is told, and it is also the rate at which you generate noise. A one-minute check on a flaky endpoint is a pager that goes off during dinner about a condition that resolved itself before you opened the laptop.

A workable default:

What it is A sensible starting interval
The front door, the app, the login The shortest your plan allows
Secondary pages and API routes Five minutes
Certificates, domains, robots, sitemaps Hourly or daily — they change slowly

Be honest with yourself about the second column of the first row: if you would not act on a two-minute outage at 3am, do not configure a check that will wake you for one.

Two things worth stating plainly about intervals in HarpyWatch. The floor depends on the plan — each plan sets a minimum interval, and the API refuses anything below it rather than silently rounding. And that floor is not a feature-list line item, it is a delay: a plan whose minimum is five minutes means a failure can be up to five minutes old before anything happens. That is the number to price against, not the badge on the pricing page.

Also worth knowing before you tune everything: browser-render runs an actual browser, so it costs more to run than a header request and carries its own, longer floor. Use it where JavaScript assembles the page, not everywhere.

Four: decide where the alert goes, and who reads it

Alerts go to email, Slack, or a webhook. Pick one. Do not pick all three on the first afternoon.

The question that matters is not the channel, it is whether the destination has a person attached at the hour the alert will arrive. The classic failure is an alias — alerts@ — that three people are technically on and nobody reads, which is the same architecture as no monitoring plus a false sense of having some. A named individual's inbox is worse-designed and works better, for the dull reason that responsibility does not divide: one person cannot assume somebody else has already seen it.

Slack is a reasonable default for a small team, with one caveat: put it in a channel people are actually in, not a dedicated #alerts that becomes wallpaper in a fortnight. Webhooks are for when you have somewhere to route to already.

If your site is customer-facing and you want to answer "is it just me?" before anybody emails you, a public status page is derived from these same results rather than being a second thing somebody updates by hand. That is worth setting up — but do it next week, not this afternoon.

Five: watch it for a week, then change exactly one thing

Leave it alone for a few days and see what comes in.

If nothing fires at all, do not read that as a clean bill of health until you have watched the alert path work once. Point a check at a URL you know is wrong, and confirm the message arrives where you think it does. Silence from a check that cannot deliver looks exactly like silence from a healthy site, and that is the failure this whole exercise was meant to avoid.

If something fires repeatedly and you have started ignoring it, that is the important finding. You have an alert you do not believe, which is worse than no alert, because it trains you to dismiss the channel the real one will arrive on.

Fix it by changing one thing: lengthen that check's interval, raise its threshold, or delete it. Not all three, and not everything else at the same time.

Why finishing beats perfecting

The setup described here is not comprehensive, and the gaps are whole categories rather than edge cases. Everything above observes the outside of the site, so nothing in it reaches your queue depths, your background jobs, or the one account whose data has been quietly wrong since Tuesday. For those you want instrumentation inside the application, which is a different tool and a different afternoon. Any list of five decisions has that shape.

But the alternative on offer is not a better setup — it is the eight-month version, where the perfect configuration is still a document and the site is observed by nobody. Five running checks tonight beats thirty planned ones.

The last thing worth saying is the one people skip: point the same checks at a staging hostname before you move DNS to it. Everything above works against any address that resolves. A broken redirect, a missing header, a certificate chain one link short — all of them are visible from outside before a single visitor arrives, and that is the cheapest moment they will ever be found.

Cite this

The HarpyWatch team. “Set it up in an afternoon and then stop”. The Nest, HarpyWatch, 6 October 2026 UTC. https://harpywatch.com/blog/set-it-up-in-an-afternoon

More from The Nest

Staging configuration travels, and agents help it

A canonical tag pointing at staging, a noindex that shipped, a test API key in production. Each is invisible on a developer's machine and expensive live, and copying config between environments is the fastest way to make something work.

article 6 min read

Monitor the site nobody is looking at yet

The instinct is to add monitoring once a site has traffic worth protecting. That gets it backwards: an empty site is the one failure mode nobody reports, because your first hundred visitors leave instead.

article 6 min read

The API contract changed underneath the frontend and nothing failed

One agent edits the backend, another edits the frontend, a field quietly changes type, and every build stays green. The screen renders blank where a number belongs, and only something watching the deployed endpoint will tell you.

article 6 min read