Scaling
Replicas, connections and disk growth
Edit this pageOn this page
The control plane is stateless and scales horizontally. PostgreSQL is the ceiling. Everything below follows from those two sentences.
Measured, and not measured
This page distinguishes the two, because most of what a product says about its own performance is the second kind wearing the first kind's clothes.
Measured is one thing: the cost of the host list at 100 001 rows, below. The method, the hardware and the seeded dataset are in ADR-0023, and the seeding script is in the repository, so you can reproduce it or contradict it.
Not measured, and therefore stated as intent rather than fact: a /24 discovered in under 30 seconds cold, a 5 000-node topology rendered in under two seconds, and p99 control-plane latency under 300 ms. These are the targets the design aims at. Nobody has timed them on stated hardware against a stated dataset, so do not plan a purchase around them. When they are measured they will move into the section below with their method attached.
Sizing
A starting point, not a benchmark result. It comes from what the components are — PostgreSQL, a stateless Node process, a Go scanner — rather than from a sizing study, so treat it as the shape of the answer and watch your own numbers.
| CPU | RAM | Disk | |
|---|---|---|---|
| Up to ~5 000 hosts | 4 vCPU | 8 GB | 40 GB |
| Tens of thousands | 8 vCPU | 16 GB | 200 GB SSD, PostgreSQL on its own volume |
| Each daemon | 1 vCPU | 512 MB | negligible |
What 100 000 hosts actually costs
Measured rather than asserted, on a database seeded to 100 001 hosts. See ADR-0023.
The list query, which is the part that scales with row count:
| Sort | Plan | Execution |
|---|---|---|
| Last seen | Index Scan idx_hosts_org_seen_id | 0.15 ms |
| Hostname | Index Scan idx_hosts_org_hostname_id | 0.14 ms |
| Device type | Index Scan idx_hosts_org_devicetype_id | 0.23 ms |
| Search | Bitmap Index Scan idx_hosts_search (GIN trigram) | 7.68 ms |
Every sort is an index scan with no sort node and no sequential scan.
The keyset cursors and the covering (organization_id, <sort key>, id)
indexes mean the planner reads twenty-five index entries and stops, so the
cost of a page does not depend on how many rows are behind it. That is the
property the 100 000-host target rests on, and it holds.
Search is the one real cost, and it is the honest price of a substring
match: 7.7 ms through a trigram index instead of the full scan an
unindexed ILIKE would be.
The p99 latency target has not been verified on production hardware. Doing so needs a production build on 8 vCPU / 16 GB, which is a deployment rather than a laptop. What can be said from measurement is that the database is not the constraint at this size.
No index was added by this pass. The measurement said the existing ones already produce the optimal plan for every sort the interface offers, and an index no query uses is a write cost with no reader.
Scaling the control plane
Add replicas behind a load balancer. There is no session affinity requirement — sessions are cookies validated against the database.
Two things to keep straight as you add replicas:
Migrations. Never automatic. Under Compose you run them; under Helm a hook Job runs once before the Deployments are touched.
The connection budget. Each process holds up to LOOMSCOPE_DB_POOL_MAX
(default 25) across all its internal pools:
(server replicas + worker replicas + 1) × LOOMSCOPE_DB_POOL_MAX < max_connectionsLeave headroom for a rolling update, which runs old and new processes at
once, and for a psql session. Raise LOOMSCOPE_DB_POOL_MAX only alongside
max_connections — the default of 25 exists precisely because seven pools
sized independently once summed to more connections than the database
allowed.
Scaling the worker
More worker replicas buy redundancy, not throughput. Every tick takes a PostgreSQL advisory lock, so a second replica skips rather than duplicates — and each one costs connection budget.
One worker is the right answer for almost everyone. Two if a worker outage would matter more than the connections it costs.
Scaling daemons
This is where throughput actually comes from. Scanning is network-bound, and a daemon can only scan what it can reach.
Add a daemon per network segment, not per unit of load. A second daemon on the same segment does not halve the scan time; a daemon on a segment that had none doubles your coverage.
One daemon per site is the usual shape. Bind each with
LOOMSCOPE_DAEMON_SITE_CODE or from the site page.
In Kubernetes, the DaemonSet gives you one per node — restrict with
nodeSelector rather than assuming every node is a useful vantage point.
PostgreSQL
The ceiling, so it is worth treating well.
| Setting | Guidance |
|---|---|
max_connections | Budget as above. 200 is comfortable for a few replicas |
shared_buffers | 25% of RAM |
work_mem | Modest; the heavy queries are aggregations over indexed columns |
| Storage | SSD. The write pattern is observation reconciliation, which is small transactions at volume |
| Autovacuum | Leave it on. Observation tables churn |
Every WHERE pattern actually used has an index, and anything suspected of
scanning more than 10 000 rows is checked with EXPLAIN ANALYZE before it
ships.
Disk growth
Three things accumulate, and none of them prunes itself yet:
| Data | Grows with | Mitigation |
|---|---|---|
| Snapshots | inventory size × days retained | Prune old snapshots by hand |
| Audit entries | rate of privileged actions | Archive by query |
| Observations | scan frequency × range size | Reconciled, but history is kept |
For a 50 000-host estate with daily snapshots, plan on tens of gigabytes a year. There is no retention policy in the product; that is a stated gap, not an oversight you have found.
Large scans
Two levers on a scan job:
topPortsTcp(default 1 000) andtopPortsUdp(default 50). Most of the time in a scan is ports, not hosts.- Exclusions on the network, so a range with something fragile in it can be scanned without touching that address.
A /16 is not a /24 with more addresses — it is 256 times the work, and most of it never answers. Split large ranges into the subnets that actually contain things, and let the routing tables collected over SNMP tell you which those are.
Multi-tenancy at scale
Organisations are isolated in the database by row-level security, so a deployment holding many customers is a normal deployment. The practical limits are the same ones above — connections and disk — rather than anything about the tenancy model.
Per-tenant daemons are the norm: each tenant's daemon holds a key scoped to that tenant, and can post observations for nothing else.