Skip to content

Monitoring

What to watch when failures are silent

Edit this page
On this page

Loomscope's failure modes are quiet. Almost nothing here errors loudly — a scanner that lost its ICMP privilege still reports success, a stale CVE feed still shows findings, a stopped scheduler still serves every page. So the things worth watching are mostly did the thing that should have happened actually happen, not did anything throw.

Prometheus scrapes /api/v1/metrics. Everything below is also checkable by hand, which is what to do while the scrape target is being wired up.

Prometheus

yaml
- job_name: loomscope
  metrics_path: /api/v1/metrics
  static_configs:
    - targets: ["loomscope.example.com"]
  authorization:
    type: Bearer
    credentials_file: /etc/prometheus/loomscope.token

The endpoint needs an API key carrying the metrics:read scope, and refuses anything else with a 401 naming the scope it wanted. It is not open: the figures are per organisation — host counts, open critical findings, which sites exist — and a multi-tenant deployment publishing those to anything that can reach the port would undo the isolation the rest of the product spends its effort on.

The scrape is scoped to the organisation the key belongs to, so one key per tenant.

MetricLabelsWatch for
loomscope_job_last_run_age_secondsjobThe most useful alert here. Above three times a job's interval means the scheduler stopped, and nothing else in the product will say so
loomscope_job_last_run_okjob0 — the last run of that job failed
loomscope_daemon_quiet_secondsstatusAbove 300 while status="online" — a daemon that stopped reporting without saying so
loomscope_daemonsstatusA drop in online
loomscope_snapshot_age_seconds—Above ~90000 (25h) — a day's snapshot was missed
loomscope_cve_feed_age_seconds—A stale mirror shortens the vulnerability list rather than erroring
loomscope_scan_sessionsstateA rising stalled count — daemons dying mid-scan
loomscope_findingsseverity, stateNew critical in state open
loomscope_certificates_expiringwindowexpired above zero
loomscope_hosts, loomscope_services—A sudden fall — discovery lost its reach

Ages are emitted only where the underlying thing exists: no snapshot has ever been taken means no loomscope_snapshot_age_seconds, rather than a zero that would read as "taken this second".

Minting a scrape key. There is no screen for this yet. Insert a row into api_keys with the metrics:read scope and the SHA-256 of the token, the same shape as a daemon or SCIM key.

Health

http
GET /api/health

Returns non-200 when the database round trip fails. The server image uses it as its HEALTHCHECK; the daemon image has none.

json
{ "ok": true, "version": "0.13.1", "ts": "2026-08-21T19:31:03.530Z" }

version is the release the running container was built from. It is the only thing that will tell you an upgrade actually replaced the process — a rollout that silently kept the old image reports ok exactly like one that worked. A build outside a container reports 0.1.0-dev.

That is the whole liveness story: it tells you the process is up, which release it is, and that the database is reachable. Nothing else.

What to watch, and why

SignalWhereWhy it matters
Control plane healthGET /api/healthProcess and database
Daemon last seendaemons.last_seen_at, or Settings → DaemonsA quiet daemon scans nothing, and nothing else says so
Job last runSettings → JobsA stopped scheduler is completely silent
Snapshot cadenceThe 60-day strip on HistoryThe fastest visual check that scheduled work is alive
CVE feed ageThe coverage banner on VulnerabilitiesA stale feed shortens the list rather than erroring
ICMP fallbackDaemon logsDiscovery silently runs at a fraction of its coverage
PostgreSQL connectionspg_stat_activityExhaustion presents as slowness, not as an error
DiskThe database volumeSnapshots and audit entries accumulate with no automatic pruning

Daemon liveness

sql
SELECT name, status, last_seen_at, now() - last_seen_at AS quiet_for
FROM daemons
ORDER BY last_seen_at;

A daemon quiet for more than five minutes is a problem. The daemon.offline event fires and can route to any notification destination, which is the closest thing to an alert that exists today.

ICMP fallback

bash
docker logs loomscope-daemon 2>&1 | grep -c "falling back to TCP"

Anything other than zero means that daemon is discovering web servers and missing printers, cameras, switches and everything else that is up but serving nothing — with no number anywhere looking wrong. Worth a check in whatever runs your log queries.

Connection budget

sql
SELECT count(*) AS in_use,
       current_setting('max_connections')::int AS max
FROM pg_stat_activity
WHERE datname = current_database();

Each control-plane process holds up to LOOMSCOPE_DB_POOL_MAX (default 25) across its pools, and the worker holds one per running job on top. Budget (servers + workers + 1) × pool_max and leave room for a rolling deploy and a psql session.

Logs

Structured JSON on stdout — pino for the control plane, zerolog for Go. Ship them with whatever you already use.

Credentials, API keys and SNMP community strings are redacted before they reach a log line.

Useful fields:

FieldUse
daemon_idCorrelate daemon and control plane
session_idFollow one scan end to end
jobWhich scheduled tick a line is about
org_idTenant, in a multi-tenant deployment

Set LOG_LEVEL=debug to see scan progress in detail. Do not leave it there — a busy daemon at debug is a lot of lines.

A minimal alerting set

If you only wire up four things:

  1. GET /api/health non-200 for more than a minute.
  2. A daemon quiet for more than five minutes — route daemon.offline to the channel people read.
  3. A job with no successful run in three times its interval — the snapshot job is the one that matters most.
  4. Database disk above 80%.

The first is the only one a conventional uptime check would have caught. The other three are the ones that actually bite.

Verifying after a change

Every restart, upgrade or outage:

bash
curl -fsS http://localhost:3000/api/health

then open Settings → Daemons and Settings → Jobs, and glance at the snapshot density strip on History.

The principle underneath: when a mechanism fails silently, absence of an error means nothing. The only honest check is to ask the thing that should have happened whether it did.