sitescope a small health monitor and status page for a few hosts

Start

Concepts

Checks, statuses, retries, notifications and history. How sitescope decides that something changed, and when to tell you.

Checks

Configuration is written in sections (dns, tls, hosts, …). The hub expands each section into flat checks, each with a stable id such as dns.soa.example.org.b.v6 or host.a.disk. Ids name the same thing across restarts, so history follows them.

Each check belongs to an area: dns, domains, tls, http, mail, hosts, hygiene, cloud or hub. Each check also has a visibility: public (listed on the public page under a label), grouped (folded into a service light) or private (not on the public page at all). See Status page and API.

Every check has its own timing: interval, timeout, retries and retryInterval, set per section or inherited from defaults.

Scheduling

The hub plans its work instead of polling a clock. Every run lands on a grid of hub.tick steps (1 minute) counted from the hub’s start, and between steps the hub sleeps.

Statuses

status means counts toward a light
ok healthy yes
warn a threshold’s warn bound crossed, or a soft problem yes
crit a crit bound crossed, or the thing is down yes
unknown not run yet, or the check itself couldn’t decide yes, as unknown
locked needs a vault secret and the vault is locked no

Thresholds are written { warn, crit }. Whether “above” or “below” is bad depends on the check: disk use is bad above, days to expiry is bad below. A bound of 0 turns it off.

Retries

A move into warn or crit must be seen retries more times, retryInterval apart, before it is committed. While retrying, the check keeps its previous status and shows “retrying 1/2”. A move to ok, unknown or locked is committed at once.

Checks derived from an agent’s report (host.a.disk, host.a.postfix, …) depend on the agent check host.a.agent and don’t retry themselves: the report already passed the agent’s retries.

Notifications

When a batch of checks finishes, the hub gathers the committed changes that are due and sends them in one email.

A failed send stays due and is retried on the hub’s next pass, usually within a minute. Only the notifiers that failed get it again.

ntfy

Email can’t reach you when the mail server is the thing that’s down. With alerts.ntfy.enabled, the hub also pushes to an ntfy topic. It pushes only changes to alerts.ntfy.min (crit, or warn) or worse, and their recoveries. Everything else, including the digest and the startup message, stays email only. Pushes follow the same batching and renotify rules as email.

The topic URL works like a password, so it comes from SITESCOPE_NTFY_URL in the environment file, never from the config. Add SITESCOPE_NTFY_TOKEN for a server that needs an access token. Host the topic somewhere other than the hub’s own network. A push carries only counts, like “2 crit, 1 recovered”, plus a link to the detail page. alerts.ntfy.details: true adds the check names. A failed push is retried twice, five seconds apart, before it waits for the next pass.

Heartbeat

Every hub.heartbeatInterval (1 minute by default) in which the store had no errors, the hub GETs SITESCOPE_HEARTBEAT_URL. Set the heartbeat service’s period to match, with some grace. If the hub dies, hangs, or loses its disk, the pings stop and the heartbeat service tells you. That covers the one thing sitescope can’t report on: itself.

History

The hub keeps history in /var/lib/sitescope/history.db (bbolt): every status change, plus one sample per hub.sampleEvery (15 minutes) for each check. History older than hub.retentionDays (35) is pruned every six hours, and so is history of checks that are no longer configured. Current states are saved every five minutes and on shutdown, so a restart picks up where it left off.