An update replaces the DataMind Installer itself. The DataMind OS stack it manages is a separate compose project and is not pulled, recreated or restarted by an Installer update.
A release can change two things, and the Installer checks both:
| What changed | How it is detected | Why it matters |
|---|---|---|
| The Installer image | The running image digest differs from the digest published on the install's channel | The code that runs |
The Installer's own docker-compose.yml | The local compose file differs from the published one | New environment variables or volume mounts cannot arrive without it |
The image is built by .azure-pipelines/build-image.yml. The Installer's own compose file is
published to blob storage by .azure-pipelines/publish-compose.yml on every push to main, so
the published file is the canonical one.
An update touches only the Installer's compose project. The DataMind OS containers carry the
label com.docker.compose.project=unistream and are a different project; the self-update helper
runs docker compose against the Installer's project only.
The Installer resolves its target from the channel, not from a hand-typed version.
| Item | Production | Development |
|---|---|---|
| Channel | prod | dev |
| Image tag the check targets | prod-latest | dev-latest |
| Registry | unistream.azurecr.io | unistreamdev.azurecr.io |
| Compose file | docker-compose.yml | docker-compose.dev.yml |
docker-compose.yml sets image: unistream.azurecr.io/delamain:${VERSION:-prod-latest}
and DELAMAIN_CHANNEL: prod; the dev compose file sets the dev equivalents.DELAMAIN_CHANNEL if set; otherwise it is the part of the running
image tag before the first - (so prod-4821 resolves to channel prod)."<channel>-latest" from the registry, whatever tag is running.In the interface, a pill in the header shows "Manager update available" when one exists; its tooltip shows the current digest and the remote digest. The check runs on load and every 30 minutes. Clicking the pill opens a confirmation dialog that states the Installer restarts, the DataMind OS stack is not touched, and to expect roughly one to three minutes without a connection.
Over the API, GET /api/system/self-update-check returns:
| Field | Meaning |
|---|---|
updateAvailable | True only when both digests are known and differ. An unreachable registry reports false, not an update |
currentDigest | Registry digest of the image this container is running |
remoteDigest | Registry digest of "<channel>-latest", or null if the registry was unreachable |
composeChanged | true when this install's compose file differs from the published one; null when the local file could not be read |
error | Why a field is null, when something failed |
lastUpdate | Outcome of the most recent self-update (status, detail, at, targetDigest) |
curl -s "http://localhost:${NEST_PORT:-8000}/api/system/self-update-check" \
-H "Authorization: Bearer <token>"Confirm the dialog. The interface calls POST /api/system/self-update and then waits for the
backend to come back (see Observe the update).
curl -s -X POST "http://localhost:${NEST_PORT:-8000}/api/system/self-update" \
-H "Authorization: Bearer <token>"The response is { "started", "targetDigest", "message" }. The endpoint is authenticated but
carries no additional role requirement, so any signed-in user or service caller can start it.
The Installer refuses to start when:
| Condition | Result |
|---|---|
| An update is already in flight | 409 Conflict |
| A deployment job is running | 409 Conflict — updating now would kill the job |
| The target digest cannot be resolved | 400 Bad Request — nothing is changed |
| The install is already newest on the channel | started: false, no helper is spawned |
An operator can always update by hand over SSH:
docker compose pull && docker compose up -d
This is also the recovery path the helper prints when both the switch and its rollback fail.
The Installer cannot replace itself from inside itself — recreating the backend would kill the
process that asked for the update. So the running backend does the safe part first, then hands the
switch to a detached helper container named delamain-self-updater, spawned from the new
image.
Still inside the running backend (SelfUpdateService.preflight):
rollback — unistream.azurecr.io/delamain:rollback
(the SELF_UPDATE_ROLLBACK_TAG). This both pins the rollback target and shields the old image
from the Installer's own dangling-image prune./var/tmp/delamain/self-update/docker-compose.yml on the host (a path bind-mounted into both
the backend and the helper at the identical location).If any of these fail, the update aborts with "nothing has changed".
The backend records a self-update action in the deployment audit trail with the target digest and
the current schema migration, force-removes any leftover delamain-self-updater container, then
runs:
docker run -d --name delamain-self-updater \
-v /var/run/docker.sock:/var/run/docker.sock \
-v <work-dir>:<work-dir> \
-v <staging-root>:<staging-root> \
unistream.azurecr.io/delamain:<new-tag> \
sh /usr/src/app/scripts/self-update.sh \
<project> <work-dir> <compose-file> /var/tmp/delamain/self-update/docker-compose.yml \
rollback <new-tag> backendThe helper mounts the host Docker socket, the compose project's working directory, and the shared
staging directory. The final argument is the Installer backend's compose service name,
backend, so the wait judges only that service and its dependencies.
Inside the helper (scripts/self-update.sh):
<compose-file>.pre-update.<compose-file>.pre-update.VERSION="<new-tag>" docker compose -p <project> -f <compose-file> \ up -d --wait --wait-timeout 300 backend
--wait blocks on the container's healthcheck, so a container that starts and immediately dies
counts as a failure. The VERSION environment variable outranks the project's .env, so an
operator-pinned VERSION cannot turn the update into a silent no-op.ok to <work-dir>/.self-update-status and exits 0.<compose-file>.pre-update and runs the same command pinned to
the rollback tag. Success writes rolled-back and exits 0 — the update did not stick.failed and exits 1. The Installer is down and needs
manual recovery over SSH: fetch the Installer compose from blob storage, then
docker compose pull && docker compose up -d. The helper's output stays available with
docker logs delamain-self-updater.The escalation chain is fixed and short: one switch, one rollback, stop. There are no retry loops and no down-migrations — reverting a schema is more destructive and less tested than the forward path, so a genuinely breaking migration makes a release non-rollbackable instead.
The --wait-timeout for both attempts is 300 seconds.
The helper writes <work-dir>/.self-update-status in the compose project's working directory. Its
value is exactly one of ok, rolled-back or failed.
That file is a breadcrumb for a human reading the host. It is not read back by the Installer — the project's working directory is not mounted into the backend container. So on its next boot the Installer reconciles by comparing image digests instead.
| Where | What you see |
|---|---|
| The interface | An overlay that polls GET /api/health every 3 seconds until the backend answers; it waits for a real downtime gap (or about 20 seconds) before reloading, and gives up after 10 minutes |
| Host | docker logs -f delamain-self-updater — the helper's step-by-step log |
| Host | docker logs -f delamain-backend — the backend's own NestJS log |
| API | GET /api/system/self-update-check → lastUpdate.status, lastUpdate.detail, targetDigest |
| API | GET /api/version → the image reference and digest that came back |
| Host | <work-dir>/.self-update-status |
When the new build returns, the interface reports one of: updated to the new tag; rolled back and
still on the previous tag (with the reason from lastUpdate.detail); or restarted on a tag that
matched neither digest.
On boot the backend looks for the last unfinished self-update action and reconciles it:
preUpdateMigration before the rollback, the detail reads
ROLLED BACK ONTO A CHANGED SCHEMA and the event is logged at error level — that build must
tolerate the new schema and needs manual verification.Confirm which build is live:
curl -s "http://localhost:${NEST_PORT:-8000}/api/version" \
-H "Authorization: Bearer <token>"The digest is exact; the tag is not, because a tag moves while a container keeps running.
An update is rolled back automatically, at most once, as described above. To roll back deliberately on the host, pin the tag and recreate the backend:
VERSION="rollback" docker compose -p <project> -f <compose-file> \ up -d --wait --wait-timeout 300 backend
rollback is the tag the pre-flight applied to the previously running image. It is protected from
the Installer's dangling-image prune, so it is still on the host.
No down-migration runs on rollback. If the new version already migrated the database, going back leaves the old build running against a newer schema — the condition the reconciliation detail flags.
VERSION in the Installer's .env; the production compose
file defaults it to prod-latest when unset.VERSION with the channel's -latest tag by design, so a
pinned value does not block an in-app update.VERSION to the tag you want
and deploy with docker compose on the host.DELAMAIN_CHANNEL selects the channel the check and the update use; the production compose file
sets it to prod.The Installer exposes the update over its own API, and the DataMind OS control plane reaches that API. The mechanism, as it appears in this repository:
unistream, where the Installer's backend is published under the alias
delamain-backend.curato.token in the Installer's secrets volume and handed to the platform's .env. It is
presented as a bearer token with x-actor-user-id and x-actor-email headers, purely so the
audit trail names a person.So the code path from the DataMind OS control plane to the Installer's update and deployment endpoints exists and is authenticated end to end.
The confirmation, the pill and the progress overlay described above belong to the Installer's own packaged interface. This repository does not contain the DataMind OS front end, so whether that interface surfaces a dedicated control that starts the Installer update cannot be confirmed from here. What is confirmed is the API path and the service-token trust boundary.
| Check | Command | Expected |
|---|---|---|
| Backend came back | curl -s http://localhost:${NEST_PORT:-8000}/api/health | {"status":"ok","db":"up"} |
| Build is the target | GET /api/version | digest equals the targetDigest of the update |
| No update stuck | GET /api/system/self-update-check | lastUpdate.status is completed, or null if never updated |
| Helper finished | docker ps -a --filter name=delamain-self-updater | Exited; read its output once with docker logs |