sitescope a small health monitor and status page for a few hosts

Use

Checks

Every check sitescope runs, what it looks at, the id it reports under, and what makes it warn or go critical.

Ids below use ZONE, HOST, NAME, FAM (v4 or v6) and PORT as placeholders. Thresholds are {warn, crit} and can be changed in Configuration.

DNS

Section dns, every 5 minutes (delegation every hour).

id checks warn crit
dns.primary the hidden primary answers SOA for every zone any zone missing
dns.soa.ZONE.HOST.FAM each secondary answers SOA with the aa flag, over IPv4 and IPv6 serial differs from the primary’s no answer, error rcode, not authoritative
dns.delegation.ZONE the TLD’s servers delegate to the expected name servers extra name servers missing name servers, no delegation
dns.resolve.NAME a public resolver (1.1.1.1) returns exactly the expected A and AAAA anything else
host.HOST.knot from the agent: every zone loaded, secondaries not near expiry expiry under 14 days under 3 days, zone missing

The delegation check asks the parent zone’s own servers, not a resolver, so it sees a registrar change before caches do.

Domains

Section domains, every 12 hours. domain.NAME reads the registration expiry over RDAP: warn under 45 days, crit under 14, crit when expired or not registered. Servers come from the IANA bootstrap file; .us and .co, which aren’t in it, have built-in fallbacks, and domains.servers can add more.

TLS

Section tls, every 6 hours. tls.NAME.PROTO.PORT.FAM connects over each IP family, verifies the chain against the system roots and the hostname, and reports days until the leaf expires: warn under 20, crit under 7. starttls: "smtp" checks a mail server’s certificate after STARTTLS on port 25. With alpn: "h2" the handshake offers h2 and http/1.1, and the check warns unless the server chooses h2.

Certificate Transparency

Section ct, once a day per domain. Every public certificate is logged in Certificate Transparency logs, so a certificate you didn’t ask for (a mis-issuing CA, or someone who briefly controlled your DNS or a web server) shows up there. ct.DOMAIN asks Cert Spotter for the domain’s currently valid certificates, subdomains included:

Acknowledge a certificate by adding its SHA-256 (or a prefix) to ignore. Domains default to domains.names; one request per domain covers all its subdomains. Without a key Cert Spotter allows 10 requests an hour, so the domains are spread evenly over the day (9 domains run 2h40m apart), a restart keeps that spacing, and a refused request is retried when Cert Spotter’s Retry-After says. A key in the vault as certspotter_token raises the limit; it is used once the vault is unlocked, and the checks run without it until then.

ct = {
  # Cloudflare's Universal SSL uses these CAs for proxied names
  domainIssuers."bllue.org" = [ "SSL.com" "Google Trust Services" ];
  # optional; "\\*" is a literal wildcard certificate name, "*" a glob
  names = [ "da.bllue.org" "ne.bllue.org" "status.bllue.org" "zircon.chooser.us" "bllue.org" "\\*.bllue.org" ];
};

HTTP

Section http, every minute. http.NAME fetches the URL without following redirects. Crit on a connection error or a status other than expectStatus (200); warn or crit when latency passes {2, 5} seconds.

For devices and private services, a target can also set:

Devices

Area devices, every minute, for things on the network that run no agent: printers, cameras, a picture frame.

A device that sleeps, like a picture frame, sets seenWithin, e.g. "6h". A failure is then ok while the device answered within that window, and the message says when it was last seen. After a hub restart the window starts again from the restart.

"ping": {"targets": [{"name": "printer.lan"}, {"name": "frame.lan", "seenWithin": "6h"}]},
"tcp": {"targets": [{"name": "camera.lan", "port": 554}]},
"ipp": {"targets": [{"name": "printer", "url": "ipp://printer.lan/ipp/print"}]}

Mail

id interval checks
mail.banner.HOST.FAM 5m SMTP greeting arrives, mentions expect, within {3, 10} s
mail.openrelay.HOST.FAM 24h from outside, RCPT TO an outside address is refused with expect (554). Accepted is crit. The conversation ends at RCPT; no message is ever sent
mail.blocklist.IP 1h each IP against each DNS blocklist, through the local resolver. Listed is warn, or crit for lists marked crit
host.HOST.postfix agent queue size {20, 200} and oldest message {1h, 4h}

Blocklist answers in 127.255.255.0/24 mean the list refused the query, usually because it came through a public resolver; they are not counted as listed.

Hosts

Section hosts, every minute. host.HOST.agent fetches the agent’s report: crit if it can’t. For a local host the hub collects the report itself, with the same checks. Every other host check reads that report, so it depends on the agent check and doesn’t retry by itself. A section the agent couldn’t collect is warn; a report older than three intervals is unknown.

id checks default
host.HOST.disk used space and inodes, worst real filesystem {80, 90} %
host.HOST.memory memory in use: total minus MemAvailable, so reclaimable page cache doesn’t count. The message shows the cache and the three largest services {90, 97} %
host.HOST.pressure what a memory shortage costs: share of time tasks stalled waiting for memory (PSI), swap-in pages/s, and OOM kills (any is crit) stall {10, 30} %, swap-in {100, 1000}/s
host.HOST.cpu CPU busy (not idle or iowait), with iowait, steal and CPU pressure in the message {85, 95} %
host.HOST.diskio each disk’s busy time, read and write throughput, and the share of time tasks stalled on I/O busy {80, 95} %, stall {25, 50} %
host.HOST.network per interface throughput; errors and drops per second errors {1, 10}/s, netMbps off
host.HOST.cgroups each service with a MemoryMax: memory in use as a share of it {85, 95} %
host.HOST.swap swap in use {60, 90} %
host.HOST.load 5-minute load per CPU {2, 4}
host.HOST.units failed systemd units: any is warn
host.HOST.unit.NAME one per entry in the host’s units: see below crit, or the unit’s severity
host.HOST.wireguard age of each peer’s latest handshake, named through wgPeers; not on a host with wireguard: false {600, 3600} s

CPU, memory pressure, disk I/O and network are rates over rateWindow (5 minutes) from the agent’s counters, so a brief spike doesn’t alert but a sustained one does. For the first two minutes after the hub or the host starts they report “collecting a baseline”.

Per host: skip and disks

A host’s skip drops checks that don’t fit it: any of disk, memory, swap, load, units, cgroups, cpu, pressure, diskio, network, wireguard, reboot, nixpkgs, updates, postfix or knot. Swap in use is a poor signal on a host with zram or plenty of free memory. There, pressure (time stalled on memory, and pages swapped back in) is what matters, so "skip": ["swap"] is reasonable.

A host’s disks sets thresholds per mount point, for volumes where one percentage doesn’t fit:

"disks": {
  "/archive": {"used": {"warn": 97, "crit": 99}},
  "/data": {"free": {"warn": 1000, "crit": 250}},
  "/scratch": {"ignore": true}
}

Units and timers

A host’s units names systemd units that must be up, beyond “no failed units”. Each entry is {name, user, severity, maxAge}. name includes the suffix. user: true asks the agent user’s own systemd manager, for an agent that runs as a user unit. severity is crit (the default) or warn.

"units": [
  {"name": "sshd.service"},
  {"name": "backup.timer", "user": true, "maxAge": "26h"},
  {"name": "fwupd-refresh.timer", "severity": "warn"}
]

The hub sends the list with each poll, so units are configured only in the hub’s config. An agent older than 1.4.0 ignores the list, and these checks report unknown until it is upgraded.

A WireGuard peer that only carries occasional traffic may go minutes without a handshake. Set PersistentKeepalive on it, raise the threshold, or list it in the host’s wgIgnore.

Hygiene

From the agent’s report, area hygiene.

Both are for NixOS hosts only. A host with another os in its entry gets these instead:

Monitoring

hub.vault, every tick: locked while the vault has been locked under hub.lockedAfter (15 minutes), then warn until someone runs sitescope unlock, so a restart nobody followed up on sends an email. Negative lockedAfter removes the check.

Cloud

Hourly (token expiry daily), and only while the vault is unlocked; until then these checks report locked.

id checks
linode.account balance and accrued charges (uninvoiced threshold in USD). An unpaid balance is warn; a payment_due or abuse-ticket notification is crit
linode.transfer network transfer used, {80, 95} % of the pool; any billable overage is warn
linode.maintenance scheduled maintenance and account notices: warn
linode.events in the last eventWindow (24h): failed events, and reboots, migrations, rebuilds, resizes, shutdowns, deletions, user and password changes: warn
linode.instance.NAME instance status is running, else crit
cloudflare.tokens daily: sitescope’s own token, plus the tokens named in tokens (default every active one), by days until they expire (tokenDays {30, 7}). Disabled, expired or a named token not found is crit; a name it can’t see because listing was refused is warn
cloudflare.records.ZONE every expected record exists: missing is crit. Any other record with an expected name and type, or a watched one, is warn