Set it up in an afternoon and then stop
Adopting monitoring for a small site is five decisions, not a project. Here is what each one actually involves, and why finishing a rough version beats designing a perfect one.
The most common state for a small site is not "badly monitored". It is "monitoring was on the list for eight months".
The reason is rarely cost or difficulty. It is that the task has no obvious edge. Once you start thinking about what to watch, the list grows — every route, every third party, every environment — and a job that was going to take an hour becomes a design exercise, and design exercises get postponed indefinitely in favour of things with deadlines.
So here is the version with an edge. Five decisions, in order, each one answerable in a few minutes. When you have made all five, stop. A rough setup that exists beats an excellent one that is still being scoped.
One: list the addresses that matter
Not the pages that matter. The addresses whose failure you would want to be woken for.
For most products it fits on one screen, and it is shorter than the first draft:
- the marketing site's front page, because it is what a stranger sees
- the application's front door, the URL people bookmark
- the login or session endpoint, which fails independently of everything else
- the two or three API routes a customer's integration actually calls
- whatever a customer would check first if they suspected you were down
Write them out. If the list is over fifteen, you are listing pages rather than failures — cut back to the ones where the answer to "would I want to know at 2am?" is yes.
In HarpyWatch's model, the thing you create is an asset, and an asset is a
hostname. Checks hang off it. If you run app.example.com, example.com and
api.example.com, that is three assets, and they will fail separately, which
is the whole reason they are separate rows.
Two: choose the checks, then choose two more
A check is one kind of observation against an asset on a schedule. Start with the ones you would have picked anyway:
http-uptimeon each asset — status code and response timessl-cert-expiryandcert-chain-truston anything served over HTTPSdns-resolution, which is upstream of all of it and fails on its owndomain-expiry, once, on each registered name
That is the set anybody would have picked. Now add two you would not have thought to ask for, chosen by what your site actually is:
- publishing anything on a feed?
feed-freshness - care about search traffic?
robots-policyandsitemap-health - an API other people call?
json-schemaandcors-policy - a page assembled by JavaScript?
browser-render, which uses a real browser - anything realtime?
websocket-handshake - anything you have ever hand-tuned for speed?
compressionandcache-policy
Two is enough. The point is not coverage. It is that the failures which sit
unnoticed longest are the ones that leave a healthy 200 in place — a feed that
stopped updating in March, a certificate chain missing its intermediate, a
Disallow: / nobody has opened the file to see — and none of those will ever
occur to you as something to ask for.
Three: pick an interval you will not resent
This is the decision people get wrong, and they get it wrong in the direction of ambitious.
The interval is not a quality setting. It is the maximum time a failure can exist before anyone is told, and it is also the rate at which you generate noise. A one-minute check on a flaky endpoint is a pager that goes off during dinner about a condition that resolved itself before you opened the laptop.
A workable default:
| What it is | A sensible starting interval |
|---|---|
| The front door, the app, the login | The shortest your plan allows |
| Secondary pages and API routes | Five minutes |
| Certificates, domains, robots, sitemaps | Hourly or daily — they change slowly |
Be honest with yourself about the second column of the first row: if you would not act on a two-minute outage at 3am, do not configure a check that will wake you for one.
Two things worth stating plainly about intervals in HarpyWatch. The floor depends on the plan — each plan sets a minimum interval, and the API refuses anything below it rather than silently rounding. And that floor is not a feature-list line item, it is a delay: a plan whose minimum is five minutes means a failure can be up to five minutes old before anything happens. That is the number to price against, not the badge on the pricing page.
Also worth knowing before you tune everything: browser-render runs an actual
browser, so it costs more to run than a header request and carries its own,
longer floor. Use it where JavaScript assembles the page, not everywhere.
Four: decide where the alert goes, and who reads it
Alerts go to email, Slack, or a webhook. Pick one. Do not pick all three on the first afternoon.
The question that matters is not the channel, it is whether the destination has
a person attached at the hour the alert will arrive. The classic failure is an
alias — alerts@ — that three people are technically on and nobody reads,
which is the same architecture as no monitoring plus a false sense of having
some. A named individual's inbox is worse-designed and works better, for the
dull reason that responsibility does not divide: one person cannot assume
somebody else has already seen it.
Slack is a reasonable default for a small team, with one caveat: put it in a
channel people are actually in, not a dedicated #alerts that becomes wallpaper
in a fortnight. Webhooks are for when you have somewhere to route to already.
If your site is customer-facing and you want to answer "is it just me?" before anybody emails you, a public status page is derived from these same results rather than being a second thing somebody updates by hand. That is worth setting up — but do it next week, not this afternoon.
Five: watch it for a week, then change exactly one thing
Leave it alone for a few days and see what comes in.
If nothing fires at all, do not read that as a clean bill of health until you have watched the alert path work once. Point a check at a URL you know is wrong, and confirm the message arrives where you think it does. Silence from a check that cannot deliver looks exactly like silence from a healthy site, and that is the failure this whole exercise was meant to avoid.
If something fires repeatedly and you have started ignoring it, that is the important finding. You have an alert you do not believe, which is worse than no alert, because it trains you to dismiss the channel the real one will arrive on.
Fix it by changing one thing: lengthen that check's interval, raise its threshold, or delete it. Not all three, and not everything else at the same time.
Why finishing beats perfecting
The setup described here is not comprehensive, and the gaps are whole categories rather than edge cases. Everything above observes the outside of the site, so nothing in it reaches your queue depths, your background jobs, or the one account whose data has been quietly wrong since Tuesday. For those you want instrumentation inside the application, which is a different tool and a different afternoon. Any list of five decisions has that shape.
But the alternative on offer is not a better setup — it is the eight-month version, where the perfect configuration is still a document and the site is observed by nobody. Five running checks tonight beats thirty planned ones.
The last thing worth saying is the one people skip: point the same checks at a staging hostname before you move DNS to it. Everything above works against any address that resolves. A broken redirect, a missing header, a certificate chain one link short — all of them are visible from outside before a single visitor arrives, and that is the cheapest moment they will ever be found.
Cite this
The HarpyWatch team. “Set it up in an afternoon and then stop”. The Nest, HarpyWatch, 6 October 2026 UTC. https://harpywatch.com/blog/set-it-up-in-an-afternoon
More from The Nest
Staging configuration travels, and agents help it
A canonical tag pointing at staging, a noindex that shipped, a test API key in production. Each is invisible on a developer's machine and expensive live, and copying config between environments is the fastest way to make something work.
Monitor the site nobody is looking at yet
The instinct is to add monitoring once a site has traffic worth protecting. That gets it backwards: an empty site is the one failure mode nobody reports, because your first hundred visitors leave instead.
The API contract changed underneath the frontend and nothing failed
One agent edits the backend, another edits the frontend, a field quietly changes type, and every build stays green. The screen renders blank where a number belongs, and only something watching the deployed endpoint will tell you.