Scheduled jobs
Nine jobs, and how to confirm they run
Edit this pageOn this page
Nine jobs run on intervals in the worker process. Every tick takes a PostgreSQL advisory lock, so running more than one worker gives you redundancy rather than duplicate work.
The schedule
| Job | Every | What it does |
|---|---|---|
scan.schedule | 30 seconds | Enqueues scans for networks that are due |
scan.stall | 1 minute | Marks sessions whose daemon stopped reporting as stalled |
notify.deliver | 1 minute | Delivers queued alerts and retries failures |
cloud.k8s | 5 minutes | Syncs Kubernetes accounts |
cve.match | 5 minutes | Matches newly observed services against advisories |
cloud.sync | 15 minutes | Syncs AWS accounts |
topology.dependencies | 1 hour | Rebuilds application-topology edges from recent flows |
notify.certificates | 1 hour | Emits tls.expiring as each threshold is crossed |
inventory.snapshot | 1 hour | Takes today's snapshot if it does not exist yet |
Two of those intervals are chosen rather than obvious:
notify.deliver every minute, because this is the queue carrying alerts.
A minute of latency on "a critical CVE was found" is about the most anyone
should accept, and the tick costs nothing when nothing is due.
inventory.snapshot hourly rather than daily at 02:00. The old scheduler
computed the next 02:00 and slept until it, which a restart at 01:59 silently
skipped. Taking a snapshot is idempotent for a day, so asking hourly whether
today's exists is both simpler and more robust.
Where they run
In the worker container. LOOMSCOPE_JOBS_IN_WEB controls whether the web
process also runs them, and it defaults to on.
That default is deliberate. Making it opt-in would mean any deployment that had not yet added the worker container silently ran no scans, no CVE matching and no snapshots — trading a known weakness for a silent one.
Running them in both places is safe rather than harmful: the advisory lock
means only one process executes each tick. But it puts the work in the
process serving requests, so set LOOMSCOPE_JOBS_IN_WEB=false once the
worker is running. The shipped compose file does.
Retry
A failed tick is retried within the same run, with backoff, where the job's interval is long enough for the wait to matter.
The policy is derived from the interval rather than configured per job, so it cannot drift out of step with it:
| Interval | Attempts | Reasoning |
|---|---|---|
| ≤ 2 minutes | 1 | The next scheduled run lands at roughly the same time as a retry would |
| > 2 minutes | 3 | 5s then 25s backoff, capped so retries finish well inside the interval |
Transient failures — a database blip, a 5xx from a provider, a timeout — are retried. A 4xx is not: a rejected request will be rejected again.
The known gap
Scheduled work does not retry across a restart. A worker that dies mid-retry resumes the normal schedule rather than resuming the attempts. For a job running every 30 seconds that costs nothing; for the hourly snapshot it can cost a day's snapshot.
Closing it properly means a durable queue, which is a dependency decision still open. It is recorded rather than glossed over.
The connection budget
The worker holds a dedicated pool, separate from the query pool, sized so that scheduled work cannot be starved by request traffic — an earlier version shared one pool and deadlocked, with the job that could not start being the job that would have recorded that it had not started.
Each running job holds one connection on top of LOOMSCOPE_DB_POOL_MAX.
Budget accordingly:
(servers + workers + 1) × LOOMSCOPE_DB_POOL_MAX < max_connectionsConfirming they are running
This is worth doing after every restart, upgrade or outage, because the failure mode is silence rather than an error.
| Check | Where |
|---|---|
| Every job has a recent run | Settings → Jobs — name, last run, duration, outcome |
| Snapshots are still being taken | The 60-day density strip on History |
| Scans are being enqueued | Discovery — sessions appearing on schedule |
| Alerts are being delivered | Send a test from Settings → Integrations |
The general principle, learned the hard way: when a mechanism fails silently, the absence of an error means nothing. The only honest check is to ask the thing that should have happened whether it did.
Running one by hand
Some jobs have an administrative endpoint for an immediate run:
curl -X POST https://loomscope.example.com/api/v1/admin/snapshot
curl -X POST https://loomscope.example.com/api/v1/admin/dependency-tick
curl -X POST https://loomscope.example.com/api/v1/admin/cloud/{id}/syncSession-authenticated, admin or owner.