Skip to content

Architecture

The components, and why there are so few

Edit this page
On this page

Loomscope is deliberately small. There is no Redis, no message broker, no websocket layer and no gRPC — every one of those would be another thing to operate in an environment where the customer, not us, is the one on call.

The shape of it

        ┌──────────────────┐
        │   Go daemon(s)   │   scans the network it sits on
        └────────┬─────────┘
                 │  outbound only: REST + long-poll
                 ▼
        ┌──────────────────┐
        │  Next.js control │   UI, REST API, cloud collectors
        │  plane           │
        └────────┬─────────┘
                 │  SQL, LISTEN/NOTIFY
                 ▼
        ┌──────────────────┐
        │   PostgreSQL 17  │   the single source of truth
        │   row-level      │
        │   security       │
        └────────▲─────────┘
                 │
        ┌────────┴─────────┐
        │   Job worker     │   scans, CVE matching, snapshots, alerts
        └──────────────────┘

Components

ComponentRequiredWhat it does
PostgreSQL 17YesEvery piece of state: inventory, jobs, sessions, realtime notifications
Control planeYesWeb UI, REST API, tRPC, cloud collectors. Stateless — all state is in PostgreSQL
Job workerYesScheduled work: scan dispatch, CVE matching, snapshots, alert delivery
DaemonAt least oneScans networks. Runs where the networks are. Outbound connections only
CVE ingest workerNoMirrors OSV and NVD. Only needed to isolate that load or feed an air-gapped mirror

The control plane is stateless, so it scales horizontally behind a load balancer. PostgreSQL is the ceiling.

How the daemon talks to the control plane

Everything is REST over HTTP, initiated by the daemon:

  1. Register once at boot, with its name, version, platform and capabilities. It gets back a daemon id and its polling intervals.
  2. Heartbeat every 30 seconds by default, reporting status and how many signatures it loaded.
  3. Long-poll for a job, holding the request open up to 60 seconds. The control plane answers immediately if a scan is already queued, and otherwise waits on a PostgreSQL LISTEN channel for one — so a scan enqueued in the UI reaches the daemon in milliseconds without anyone polling in a tight loop.
  4. Post observations in batches while a scan runs, and once more with final: true at the end.

No connection is ever made towards a daemon. That is the property that makes a daemon in a customer's DMZ, behind NAT, or on an isolated VLAN a normal deployment. The full contract is in the REST API reference.

Why REST and SSE rather than gRPC and websockets

Realtime updates in the browser use Server-Sent Events fed by PostgreSQL LISTEN/NOTIFY. The database is already there, already durable, already the thing that knows when something changed.

A websocket layer would need its own fan-out story across control-plane replicas; a broker would need operating; gRPC would need a proxy to survive the corporate middleboxes this software is specifically meant to run behind. REST over long-poll and SSE crosses every one of them, and degrades to ordinary HTTP when something in the path buffers.

Where the trust boundaries are

The daemon holds one credential: its own API key, hashed at rest in the control plane. It has no database access, no cloud credentials and no ability to read another daemon's work.

Cloud discovery runs in the control plane. AWS and Kubernetes credentials never reach a daemon host. This is a hard rule in the codebase, not a convention.

Credentials are encrypted at rest with AEAD, keyed by LOOMSCOPE_KMS_KEY which lives only in the environment. The database alone does not decrypt them — which is also why a backup restored under a different key comes back looking healthy and silently fails to authenticate anything.

Tenant isolation is in the database, not the application. Every org-scoped table has row-level security enabled and FORCEd, so a forgotten WHERE clause returns nothing instead of returning another organisation's estate.

The stack, and why

LayerChoiceReason
Control planeTypeScript, Next.js 15 (App Router, RSC)One language for UI and API; server components keep list pages off the wire
ScannerGo 1.23+Static binary, netip.Addr for honest v4/v6 handling, cheap goroutines
DatabasePostgreSQL 17inet/cidr types, RLS, LISTEN/NOTIFY, advisory locks — no add-ons
ORMDrizzle, SQL-first migrationsMigrations are readable SQL files, lintable by squawk
AuthBetter-AuthOIDC, 2FA, magic links, API keys and SCIM from one library
RealtimeSSE over LISTEN/NOTIFYNo broker, no second delivery path to operate
UITailwind, shadcn/ui, React Flow, RechartsTopologies at 5 000 nodes, tables at 100 000 rows
AIVercel AI SDKAnthropic, OpenAI or a local Ollama — swappable, and off by default

Repository layout

PathWhat lives there
apps/serverControl plane: UI, REST API, jobs, AI, cloud collectors
apps/docsThis documentation site
services/daemonGo scanner: ICMP/ARP/TCP/UDP, SNMP, TLS, Docker, NetFlow
services/cve-ingestOptional standalone OSV/NVD mirror worker
packages/contractsZod schemas — source of truth for the OpenAPI spec and the Go client
packages/tokensDesign tokens, shared by the control plane and this site
infraCompose files, Dockerfiles, systemd units, Helm chart

Performance targets

Targets, and this section used to call them "the numbers the design is held to, not aspirations" — which was a claim about evidence that did not exist. Only the second one has been measured.

  • Measured. 100 000 hosts per organisation with no UI degradation: every list is sorted, filtered and paged on the server with keyset cursors, and at 100 001 rows every sort is still an index scan. Method, hardware and dataset in ADR-0023.
  • Not measured. A /24 discovered in under 30 seconds cold; a 5 000-node topology rendered in under two seconds; p99 control-plane latency under 300 ms on 8 vCPU / 16 GB. These are what the design aims at. Nobody has timed them on stated hardware against a stated dataset.

See Scaling for the distinction and what it means for sizing.