Start
Concepts
Checks, statuses, retries, notifications and history. How sitescope decides that something changed, and when to tell you.
Checks
Configuration is written in sections (dns,
tls, hosts, …). The hub expands each section
into flat checks, each with a stable id such as
dns.soa.example.org.b.v6 or host.a.disk. Ids
name the same thing across restarts, so history follows them.
Each check belongs to an area: dns,
domains, tls, http,
mail, hosts, hygiene,
cloud or hub. Each check also has a
visibility: public (listed on the public
page under a label), grouped (folded into a service light)
or private (not on the public page at all). See Status page and API.
Every check has its own timing: interval,
timeout, retries and
retryInterval, set per section or inherited from
defaults.
Scheduling
The hub plans its work instead of polling a clock. Every run lands on
a grid of hub.tick steps (1 minute) counted from the hub’s
start, and between steps the hub sleeps.
- Phases. Each check gets a fixed offset within its interval. Checks that run every tick (the agent polls) start together; slower checks get a stable spot in the first 15 minutes, so the 5-minute DNS and SMTP probes and the hourly ones don’t all hit their targets at once.
- Spread groups. Checks that share a rate-limited API are spaced evenly over their interval: the CT checks, one per domain, every 24 hours.
- Host checks run as soon as their agent’s report arrives.
- Restarts. A check that ran before a restart keeps its pace; one that never ran shows “first run at …” until its slot.
- Retries wait for the next step after
retryInterval; an API’sRetry-Afterbrings the next run forward. - Housekeeping follows the work: when a batch of results is in, the hub sends one email for what changed, then pings the heartbeat, saves state and prunes history when those are due. It wakes at least every 5 minutes.
Statuses
| status | means | counts toward a light |
|---|---|---|
ok |
healthy | yes |
warn |
a threshold’s warn bound crossed, or a soft problem | yes |
crit |
a crit bound crossed, or the thing is down | yes |
unknown |
not run yet, or the check itself couldn’t decide | yes, as unknown |
locked |
needs a vault secret and the vault is locked | no |
Thresholds are written { warn, crit }. Whether “above”
or “below” is bad depends on the check: disk use is bad above, days to
expiry is bad below. A bound of 0 turns it off.
Retries
A move into warn or crit
must be seen retries more times, retryInterval
apart, before it is committed. While retrying, the check keeps its
previous status and shows “retrying 1/2”. A move to ok,
unknown or locked is committed at once.
Checks derived from an agent’s report (host.a.disk,
host.a.postfix, …) depend on the agent check
host.a.agent and don’t retry themselves: the report already
passed the agent’s retries.
Notifications
When a batch of checks finishes, the hub gathers the committed changes that are due and sends them in one email.
- The first ok is silent. A check seen healthy for the first time sends nothing.
- Renotify interval. After a check has been notified,
its next change waits at least
alerts.renotifyInterval(1 hour by default). - Flapping is absorbed. If a check changes and changes back before its next notification is due, nothing is sent.
- Daily digest. After
alerts.digestTime, one email lists every check that isn’t ok. A hub started after that time skips the day’s digest, and none goes out in the first 10 minutes after a start, so a restart never mails a list of checks that simply haven’t run yet. - Unknown waits. A move to
unknowndoesn’t send email at once. A check still unknown afteralerts.unknownAfter(1 hour, counted from startup for a check that has never produced a result) is reported once as UNKNOWN, and its recovery is reported. Checks whose agent is down stay quiet; the agent’s own alert covers them. - Startup. Every start sends “sitescope started on
HOST - vault LOCKED, run sitescope unlock”, unless no check needs a
vault secret. If the vault is still locked
hub.lockedAfter(15 minutes) later,hub.vaultsends one WARN email, and a recovery email once you unlock.
A failed send stays due and is retried on the hub’s next pass, usually within a minute. Only the notifiers that failed get it again.
ntfy
Email can’t reach you when the mail server is the thing that’s down.
With alerts.ntfy.enabled, the hub also pushes to an ntfy topic. It pushes only changes to
alerts.ntfy.min (crit, or warn)
or worse, and their recoveries. Everything else, including the digest
and the startup message, stays email only. Pushes follow the same
batching and renotify rules as email.
The topic URL works like a password, so it comes from
SITESCOPE_NTFY_URL in the environment file, never from the
config. Add SITESCOPE_NTFY_TOKEN for a server that needs an
access token. Host the topic somewhere other than the hub’s own network.
A push carries only counts, like “2 crit, 1 recovered”, plus a link to
the detail page. alerts.ntfy.details: true adds the check
names. A failed push is retried twice, five seconds apart, before it
waits for the next pass.
Heartbeat
Every hub.heartbeatInterval (1 minute by default) in
which the store had no errors, the hub GETs
SITESCOPE_HEARTBEAT_URL. Set the heartbeat service’s period
to match, with some grace. If the hub dies, hangs, or loses its disk,
the pings stop and the heartbeat service tells you. That covers the one
thing sitescope can’t report on: itself.
History
The hub keeps history in /var/lib/sitescope/history.db
(bbolt): every status change, plus one sample per
hub.sampleEvery (15 minutes) for each check. History older
than hub.retentionDays (35) is pruned every six hours, and
so is history of checks that are no longer configured. Current states
are saved every five minutes and on shutdown, so a restart picks up
where it left off.