Update the DataMind Installer

An update replaces the DataMind Installer itself. The DataMind OS stack it manages is a separate compose project and is not pulled, recreated or restarted by an Installer update.

Understand what an update changes

A release can change two things, and the Installer checks both:

What changedHow it is detectedWhy it matters
The Installer imageThe running image digest differs from the digest published on the install's channelThe code that runs
The Installer's own docker-compose.ymlThe local compose file differs from the published oneNew environment variables or volume mounts cannot arrive without it

The image is built by .azure-pipelines/build-image.yml. The Installer's own compose file is published to blob storage by .azure-pipelines/publish-compose.yml on every push to main, so the published file is the canonical one.

Important

An update touches only the Installer's compose project. The DataMind OS containers carry the label com.docker.compose.project=unistream and are a different project; the self-update helper runs docker compose against the Installer's project only.

Install channels and image tags

The Installer resolves its target from the channel, not from a hand-typed version.

ItemProductionDevelopment
Channelproddev
Image tag the check targetsprod-latestdev-latest
Registryunistream.azurecr.iounistreamdev.azurecr.io
Compose filedocker-compose.ymldocker-compose.dev.yml

Check whether an update is available

In the interface, a pill in the header shows "Manager update available" when one exists; its tooltip shows the current digest and the remote digest. The check runs on load and every 30 minutes. Clicking the pill opens a confirmation dialog that states the Installer restarts, the DataMind OS stack is not touched, and to expect roughly one to three minutes without a connection.

Over the API, GET /api/system/self-update-check returns:

FieldMeaning
updateAvailableTrue only when both digests are known and differ. An unreachable registry reports false, not an update
currentDigestRegistry digest of the image this container is running
remoteDigestRegistry digest of "<channel>-latest", or null if the registry was unreachable
composeChangedtrue when this install's compose file differs from the published one; null when the local file could not be read
errorWhy a field is null, when something failed
lastUpdateOutcome of the most recent self-update (status, detail, at, targetDigest)
bash
curl -s "http://localhost:${NEST_PORT:-8000}/api/system/self-update-check" \
  -H "Authorization: Bearer <token>"

Start the update

From the interface

Confirm the dialog. The interface calls POST /api/system/self-update and then waits for the backend to come back (see Observe the update).

From the API

bash
curl -s -X POST "http://localhost:${NEST_PORT:-8000}/api/system/self-update" \
  -H "Authorization: Bearer <token>"

The response is { "started", "targetDigest", "message" }. The endpoint is authenticated but carries no additional role requirement, so any signed-in user or service caller can start it.

The Installer refuses to start when:

ConditionResult
An update is already in flight409 Conflict
A deployment job is running409 Conflict — updating now would kill the job
The target digest cannot be resolved400 Bad Request — nothing is changed
The install is already newest on the channelstarted: false, no helper is spawned

From the host

An operator can always update by hand over SSH:

bash
docker compose pull && docker compose up -d

This is also the recovery path the helper prints when both the switch and its rollback fail.

What the helper container does

The Installer cannot replace itself from inside itself — recreating the backend would kill the process that asked for the update. So the running backend does the safe part first, then hands the switch to a detached helper container named delamain-self-updater, spawned from the new image.

Pre-flight, before anything changes

Still inside the running backend (SelfUpdateService.preflight):

  1. Tags the image that is currently running as rollback — unistream.azurecr.io/delamain:rollback (the SELF_UPDATE_ROLLBACK_TAG). This both pins the rollback target and shields the old image from the Installer's own dangling-image prune.
  2. Downloads the published compose file for the channel and writes it to /var/tmp/delamain/self-update/docker-compose.yml on the host (a path bind-mounted into both the backend and the helper at the identical location).
  3. Pulls the target image.

If any of these fail, the update aborts with "nothing has changed".

Spawning the helper

The backend records a self-update action in the deployment audit trail with the target digest and the current schema migration, force-removes any leftover delamain-self-updater container, then runs:

bash
docker run -d --name delamain-self-updater \
  -v /var/run/docker.sock:/var/run/docker.sock \
  -v <work-dir>:<work-dir> \
  -v <staging-root>:<staging-root> \
  unistream.azurecr.io/delamain:<new-tag> \
  sh /usr/src/app/scripts/self-update.sh \
    <project> <work-dir> <compose-file> /var/tmp/delamain/self-update/docker-compose.yml \
    rollback <new-tag> backend

The helper mounts the host Docker socket, the compose project's working directory, and the shared staging directory. The final argument is the Installer backend's compose service name, backend, so the wait judges only that service and its dependencies.

The switch, and the one rollback

Inside the helper (scripts/self-update.sh):

  1. Backs up the live compose file to <compose-file>.pre-update.
  2. Installs the staged compose file over the live one. Local edits to the compose file do not survive an update; the previous version stays at <compose-file>.pre-update.
  3. Switches onto the new tag:
    bash
    VERSION="<new-tag>" docker compose -p <project> -f <compose-file> \
      up -d --wait --wait-timeout 300 backend
    --wait blocks on the container's healthcheck, so a container that starts and immediately dies counts as a failure. The VERSION environment variable outranks the project's .env, so an operator-pinned VERSION cannot turn the update into a silent no-op.
  4. If the switch succeeds, writes ok to <work-dir>/.self-update-status and exits 0.
  5. If the switch fails, restores <compose-file>.pre-update and runs the same command pinned to the rollback tag. Success writes rolled-back and exits 0 — the update did not stick.
  6. If the rollback also fails, writes failed and exits 1. The Installer is down and needs manual recovery over SSH: fetch the Installer compose from blob storage, then docker compose pull && docker compose up -d. The helper's output stays available with docker logs delamain-self-updater.

The escalation chain is fixed and short: one switch, one rollback, stop. There are no retry loops and no down-migrations — reverting a schema is more destructive and less tested than the forward path, so a genuinely breaking migration makes a release non-rollbackable instead.

The --wait-timeout for both attempts is 300 seconds.

The status file

The helper writes <work-dir>/.self-update-status in the compose project's working directory. Its value is exactly one of ok, rolled-back or failed.

That file is a breadcrumb for a human reading the host. It is not read back by the Installer — the project's working directory is not mounted into the backend container. So on its next boot the Installer reconciles by comparing image digests instead.

Observe the update

WhereWhat you see
The interfaceAn overlay that polls GET /api/health every 3 seconds until the backend answers; it waits for a real downtime gap (or about 20 seconds) before reloading, and gives up after 10 minutes
Hostdocker logs -f delamain-self-updater — the helper's step-by-step log
Hostdocker logs -f delamain-backend — the backend's own NestJS log
APIGET /api/system/self-update-check → lastUpdate.status, lastUpdate.detail, targetDigest
APIGET /api/version → the image reference and digest that came back
Host<work-dir>/.self-update-status

When the new build returns, the interface reports one of: updated to the new tag; rolled back and still on the previous tag (with the reason from lastUpdate.detail); or restarted on a tag that matched neither digest.

Confirm the outcome and read the reconciliation

On boot the backend looks for the last unfinished self-update action and reconciles it:

Confirm which build is live:

bash
curl -s "http://localhost:${NEST_PORT:-8000}/api/version" \
  -H "Authorization: Bearer <token>"

The digest is exact; the tag is not, because a tag moves while a container keeps running.

Roll back

An update is rolled back automatically, at most once, as described above. To roll back deliberately on the host, pin the tag and recreate the backend:

bash
VERSION="rollback" docker compose -p <project> -f <compose-file> \
  up -d --wait --wait-timeout 300 backend

rollback is the tag the pre-flight applied to the previously running image. It is protected from the Installer's dangling-image prune, so it is still on the host.

Warning

No down-migration runs on rollback. If the new version already migrated the database, going back leaves the old build running against a newer schema — the condition the reconciliation detail flags.

Pin a version

Trigger an update from the DataMind OS side

The Installer exposes the update over its own API, and the DataMind OS control plane reaches that API. The mechanism, as it appears in this repository:

So the code path from the DataMind OS control plane to the Installer's update and deployment endpoints exists and is authenticated end to end.

Note

The confirmation, the pill and the progress overlay described above belong to the Installer's own packaged interface. This repository does not contain the DataMind OS front end, so whether that interface surfaces a dedicated control that starts the Installer update cannot be confirmed from here. What is confirmed is the API path and the service-token trust boundary.

Verify the update

CheckCommandExpected
Backend came backcurl -s http://localhost:${NEST_PORT:-8000}/api/health{"status":"ok","db":"up"}
Build is the targetGET /api/versiondigest equals the targetDigest of the update
No update stuckGET /api/system/self-update-checklastUpdate.status is completed, or null if never updated
Helper finisheddocker ps -a --filter name=delamain-self-updaterExited; read its output once with docker logs