Monitoring
What to watch when failures are silent
Edit this pageOn this page
Loomscope's failure modes are quiet. Almost nothing here errors loudly — a scanner that lost its ICMP privilege still reports success, a stale CVE feed still shows findings, a stopped scheduler still serves every page. So the things worth watching are mostly did the thing that should have happened actually happen, not did anything throw.
Prometheus scrapes /api/v1/metrics. Everything below is also checkable
by hand, which is what to do while the scrape target is being wired up.
Prometheus
- job_name: loomscope
metrics_path: /api/v1/metrics
static_configs:
- targets: ["loomscope.example.com"]
authorization:
type: Bearer
credentials_file: /etc/prometheus/loomscope.tokenThe endpoint needs an API key carrying the metrics:read scope, and
refuses anything else with a 401 naming the scope it wanted. It is not
open: the figures are per organisation — host counts, open critical
findings, which sites exist — and a multi-tenant deployment publishing
those to anything that can reach the port would undo the isolation the
rest of the product spends its effort on.
The scrape is scoped to the organisation the key belongs to, so one key per tenant.
| Metric | Labels | Watch for |
|---|---|---|
loomscope_job_last_run_age_seconds | job | The most useful alert here. Above three times a job's interval means the scheduler stopped, and nothing else in the product will say so |
loomscope_job_last_run_ok | job | 0 — the last run of that job failed |
loomscope_daemon_quiet_seconds | status | Above 300 while status="online" — a daemon that stopped reporting without saying so |
loomscope_daemons | status | A drop in online |
loomscope_snapshot_age_seconds | — | Above ~90000 (25h) — a day's snapshot was missed |
loomscope_cve_feed_age_seconds | — | A stale mirror shortens the vulnerability list rather than erroring |
loomscope_scan_sessions | state | A rising stalled count — daemons dying mid-scan |
loomscope_findings | severity, state | New critical in state open |
loomscope_certificates_expiring | window | expired above zero |
loomscope_hosts, loomscope_services | — | A sudden fall — discovery lost its reach |
Ages are emitted only where the underlying thing exists: no snapshot has
ever been taken means no loomscope_snapshot_age_seconds, rather than a
zero that would read as "taken this second".
Minting a scrape key. There is no screen for this yet. Insert a row
into api_keys with the metrics:read scope and the SHA-256 of the
token, the same shape as a daemon or SCIM key.
Health
GET /api/healthReturns non-200 when the database round trip fails. The server image uses it
as its HEALTHCHECK; the daemon image has none.
{ "ok": true, "version": "0.13.1", "ts": "2026-08-21T19:31:03.530Z" }version is the release the running container was built from. It is the
only thing that will tell you an upgrade actually replaced the process —
a rollout that silently kept the old image reports ok exactly like one
that worked. A build outside a container reports 0.1.0-dev.
That is the whole liveness story: it tells you the process is up, which release it is, and that the database is reachable. Nothing else.
What to watch, and why
| Signal | Where | Why it matters |
|---|---|---|
| Control plane health | GET /api/health | Process and database |
| Daemon last seen | daemons.last_seen_at, or Settings → Daemons | A quiet daemon scans nothing, and nothing else says so |
| Job last run | Settings → Jobs | A stopped scheduler is completely silent |
| Snapshot cadence | The 60-day strip on History | The fastest visual check that scheduled work is alive |
| CVE feed age | The coverage banner on Vulnerabilities | A stale feed shortens the list rather than erroring |
| ICMP fallback | Daemon logs | Discovery silently runs at a fraction of its coverage |
| PostgreSQL connections | pg_stat_activity | Exhaustion presents as slowness, not as an error |
| Disk | The database volume | Snapshots and audit entries accumulate with no automatic pruning |
Daemon liveness
SELECT name, status, last_seen_at, now() - last_seen_at AS quiet_for
FROM daemons
ORDER BY last_seen_at;A daemon quiet for more than five minutes is a problem. The daemon.offline
event fires and can route to any notification
destination, which is the closest thing to an alert
that exists today.
ICMP fallback
docker logs loomscope-daemon 2>&1 | grep -c "falling back to TCP"Anything other than zero means that daemon is discovering web servers and missing printers, cameras, switches and everything else that is up but serving nothing — with no number anywhere looking wrong. Worth a check in whatever runs your log queries.
Connection budget
SELECT count(*) AS in_use,
current_setting('max_connections')::int AS max
FROM pg_stat_activity
WHERE datname = current_database();Each control-plane process holds up to LOOMSCOPE_DB_POOL_MAX (default 25)
across its pools, and the worker holds one per running job on top. Budget
(servers + workers + 1) × pool_max and leave room for a rolling deploy and
a psql session.
Logs
Structured JSON on stdout — pino for the control plane, zerolog for Go.
Ship them with whatever you already use.
Credentials, API keys and SNMP community strings are redacted before they reach a log line.
Useful fields:
| Field | Use |
|---|---|
daemon_id | Correlate daemon and control plane |
session_id | Follow one scan end to end |
job | Which scheduled tick a line is about |
org_id | Tenant, in a multi-tenant deployment |
Set LOG_LEVEL=debug to see scan progress in detail. Do not leave it there —
a busy daemon at debug is a lot of lines.
A minimal alerting set
If you only wire up four things:
GET /api/healthnon-200 for more than a minute.- A daemon quiet for more than five minutes — route
daemon.offlineto the channel people read. - A job with no successful run in three times its interval — the snapshot job is the one that matters most.
- Database disk above 80%.
The first is the only one a conventional uptime check would have caught. The other three are the ones that actually bite.
Verifying after a change
Every restart, upgrade or outage:
curl -fsS http://localhost:3000/api/healththen open Settings → Daemons and Settings → Jobs, and glance at the snapshot density strip on History.
The principle underneath: when a mechanism fails silently, absence of an error means nothing. The only honest check is to ask the thing that should have happened whether it did.