Skip to content

Scheduled jobs

Nine jobs, and how to confirm they run

Edit this page
On this page

Nine jobs run on intervals in the worker process. Every tick takes a PostgreSQL advisory lock, so running more than one worker gives you redundancy rather than duplicate work.

The schedule

JobEveryWhat it does
scan.schedule30 secondsEnqueues scans for networks that are due
scan.stall1 minuteMarks sessions whose daemon stopped reporting as stalled
notify.deliver1 minuteDelivers queued alerts and retries failures
cloud.k8s5 minutesSyncs Kubernetes accounts
cve.match5 minutesMatches newly observed services against advisories
cloud.sync15 minutesSyncs AWS accounts
topology.dependencies1 hourRebuilds application-topology edges from recent flows
notify.certificates1 hourEmits tls.expiring as each threshold is crossed
inventory.snapshot1 hourTakes today's snapshot if it does not exist yet

Two of those intervals are chosen rather than obvious:

notify.deliver every minute, because this is the queue carrying alerts. A minute of latency on "a critical CVE was found" is about the most anyone should accept, and the tick costs nothing when nothing is due.

inventory.snapshot hourly rather than daily at 02:00. The old scheduler computed the next 02:00 and slept until it, which a restart at 01:59 silently skipped. Taking a snapshot is idempotent for a day, so asking hourly whether today's exists is both simpler and more robust.

Where they run

In the worker container. LOOMSCOPE_JOBS_IN_WEB controls whether the web process also runs them, and it defaults to on.

That default is deliberate. Making it opt-in would mean any deployment that had not yet added the worker container silently ran no scans, no CVE matching and no snapshots — trading a known weakness for a silent one.

Running them in both places is safe rather than harmful: the advisory lock means only one process executes each tick. But it puts the work in the process serving requests, so set LOOMSCOPE_JOBS_IN_WEB=false once the worker is running. The shipped compose file does.

Retry

A failed tick is retried within the same run, with backoff, where the job's interval is long enough for the wait to matter.

The policy is derived from the interval rather than configured per job, so it cannot drift out of step with it:

IntervalAttemptsReasoning
≤ 2 minutes1The next scheduled run lands at roughly the same time as a retry would
> 2 minutes35s then 25s backoff, capped so retries finish well inside the interval

Transient failures — a database blip, a 5xx from a provider, a timeout — are retried. A 4xx is not: a rejected request will be rejected again.

The known gap

Scheduled work does not retry across a restart. A worker that dies mid-retry resumes the normal schedule rather than resuming the attempts. For a job running every 30 seconds that costs nothing; for the hourly snapshot it can cost a day's snapshot.

Closing it properly means a durable queue, which is a dependency decision still open. It is recorded rather than glossed over.

The connection budget

The worker holds a dedicated pool, separate from the query pool, sized so that scheduled work cannot be starved by request traffic — an earlier version shared one pool and deadlocked, with the job that could not start being the job that would have recorded that it had not started.

Each running job holds one connection on top of LOOMSCOPE_DB_POOL_MAX. Budget accordingly:

(servers + workers + 1) × LOOMSCOPE_DB_POOL_MAX  <  max_connections

Confirming they are running

This is worth doing after every restart, upgrade or outage, because the failure mode is silence rather than an error.

CheckWhere
Every job has a recent runSettings → Jobs — name, last run, duration, outcome
Snapshots are still being takenThe 60-day density strip on History
Scans are being enqueuedDiscovery — sessions appearing on schedule
Alerts are being deliveredSend a test from Settings → Integrations

The general principle, learned the hard way: when a mechanism fails silently, the absence of an error means nothing. The only honest check is to ask the thing that should have happened whether it did.

Running one by hand

Some jobs have an administrative endpoint for an immediate run:

bash
curl -X POST https://loomscope.example.com/api/v1/admin/snapshot
curl -X POST https://loomscope.example.com/api/v1/admin/dependency-tick
curl -X POST https://loomscope.example.com/api/v1/admin/cloud/{id}/sync

Session-authenticated, admin or owner.