Skip to content

Upgrading

Expand-and-contract, and rolling back

Edit this page
On this page

Loomscope is versioned semantically and every image is pinned. An upgrade is a deliberate act: change the tag, run migrations, restart. Nothing upgrades itself.

Before you start

  1. Read the changelog for every version between yours and the target. Breaking changes and known limitations are recorded there, including the ones that are awkward to admit.
  2. Take a backup and confirm you hold LOOMSCOPE_KMS_KEY. See Backup and restore.
  3. Upgrade one minor version at a time. Skipping several is not tested.

Docker Compose

bash
# 1. Update the tag in infra/docker-compose.yml, or override it.
VERSION=0.13.1

# 2. Pull the new images.
docker compose -f infra/docker-compose.yml pull

# 3. Stop the daemons. They tolerate a control plane that is briefly away.
docker compose -f infra/docker-compose.yml stop daemon

# 4. Start the new control plane, then migrate.
docker compose -f infra/docker-compose.yml up -d postgres server
docker compose -f infra/docker-compose.yml run --rm server node --experimental-strip-types apps/server/server/db/migrate.ts

# 5. Bring the rest back.
docker compose -f infra/docker-compose.yml up -d worker daemon

Kubernetes

bash
helm upgrade loomscope infra/helm/loomscope -n loomscope --reuse-values

Migrations run as a Helm hook Job before the Deployments are touched, and Helm waits for it. A failed migration stops the rollout rather than leaving new code running against an old schema.

Why migrations do not run automatically

With more than one control-plane replica, automatic migration on boot means every replica racing to alter the same schema. The Compose path therefore asks you to run them; the Kubernetes path uses a hook Job that runs exactly once.

The runner records each file in a __migrations table and skips what is already applied, so re-running it is safe. Each file executes as a single multi-statement query, which PostgreSQL wraps in an implicit transaction — a failing migration rolls back whole.

Expand and contract

Loomscope's deploy model is rolling: old and new control-plane processes run side by side during a rollout. A migration that breaks the old code crashes the running container.

So every schema change that would break a running replica is split across two releases:

  • Expand (release N): add the new column nullable, write to both, never drop the old one.
  • Contract (release N+1 or later): drop the old column or add NOT NULL, once nothing reads the old shape.

The practical consequence for you: a new schema stays readable by the previous release, so a rollback within one minor version is a matter of putting the old image back. Rolling back across a contract migration is not, which is what the backup is for.

Every migration is also linted by squawk in CI, indexes on large tables are created CONCURRENTLY, and any ALTER sets lock_timeout so it fails fast rather than locking out the application.

Rolling back

Within a minor version, and with no contract migration between:

bash
# Put the previous tag back and restart. The schema is still compatible.
docker compose -f infra/docker-compose.yml up -d

Across a contract migration, restore the backup. There is no down-migration path — down migrations that are never run are down migrations that do not work, and pretending otherwise is worse than saying so.

Upgrading daemons

Daemons and the control plane are versioned together and should be upgraded together. The daemon reports its version at registration, and the daemon list shows it — a daemon left behind is visible rather than silent.

A daemon one patch version behind will keep working; a daemon several minor versions behind may be posting observations against a schema the control plane has moved past.

After an upgrade

CheckHow
Control plane is healthycurl -fsS http://localhost:3000/api/health
The new version is the one runningThe same call — version in the response must be the release you deployed
Migrations appliedThe migration command printed each file it applied
Daemons reconnectedSettings → Daemons — all online, versions current
Scheduled work resumedSettings → Jobs — every job has a recent run
Credentials still decryptOpen a stored credential in the UI. If it renders, LOOMSCOPE_KMS_KEY survived the upgrade

That last one is worth doing every time. A container that lost its environment comes back looking completely healthy and silently fails to authenticate anything.