The DataMind Installer updates itself through a detached helper container, and reconciles the outcome on the next boot by comparing image digests. Startup failures are almost always one of four things: migrations, port conflicts, missing secrets, or a container that starts and immediately fails its health check. Each symptom below tells you where to look first.
After a self-update, GET /api/version still reports the old digest, and GET /api/system/self-update-check shows a lastUpdate whose status is failed. When this happens the helper has already tried its one automated rollback.
The self-update runs in a fixed, short chain: pull the new image and stage the compose file (pre-flight), then switch the Installer onto the new image, and if that fails, roll back exactly once, then stop. There are no retries and no down-migrations — reverting a schema is more destructive and less tested than going forward, so a genuinely breaking migration makes a release non-rollbackable instead.
The Installer reconciles what happened on its next boot by comparing the running image digest to the digest it recorded before the update. It reads the audit record, not the helper's status file (the helper's working directory is not mounted into the backend, so the new process cannot read it).
Read the recorded detail, which states one of two outcomes:
Target <digest> was not applied — running <digest>. The helper rolled back; the schema is unchanged.
ROLLED BACK ONTO A CHANGED SCHEMA: target <digest> was not applied — running <digest> — but the database migrated from "<old migration>" to "<new migration>" before the rollback. This build must tolerate the new schema; verify manually.
Check the knob that most often fails the switch:
GET /api/system/self-update-check
The response's updateAvailable, currentDigest, remoteDigest, composeChanged and error tell you whether a newer image was seen at all, and whether the compose file also drifted.
The helper logs its entire run, including why a switch failed:
docker logs delamain-self-updater
The messages it emits, in order, are:
[self-update] backing up <compose-file> to <compose-file>.pre-update [self-update] installing the new compose file from <staged-compose> [self-update] switching project <project> (service backend) onto tag <tag> [self-update] switch failed — attempting the one automated rollback [self-update] rolled back to <tag> — the update did NOT stick [self-update] rollback failed too. delamain is down and needs manual recovery over SSH: [self-update] fetch the delamain compose from blob storage, then: docker compose pull && docker compose up -d [self-update] helper logs remain available via: docker logs delamain-self-updater
It also writes a one-line status file on the host — a breadcrumb for an operator, not for the Installer:
<WORK_DIR>/.self-update-status → ok | rolled-back | failed
The previous compose file is kept beside the live one as <COMPOSE_FILE>.pre-update until the next update overwrites it.
The helper's switch did not complete, and it either rolled back or reported failure.
The switch runs docker compose up -d --wait with a 300-second wait timeout. --wait blocks on the service's health check, so a container that starts and immediately dies counts as a failure. The new compose file has already replaced the live one at this point; the rollback restores the backup.
docker logs delamain-backend) carry the reason; the switch timeout is only the symptom.<COMPOSE_FILE>.pre-update.One of:
Update preparation failed; nothing has changed: <reason> Could not start the updater; nothing has changed: <reason> Could not determine which image to update to — refusing to proceed Registry returned <status> for <tag> The running image is unknown, so there is nothing to compare Could not determine the install channel from <image>; set DELAMAIN_CHANNEL Already running the newest image on this channel — nothing to do.
These are the safe refusals. Every one happens before anything changes: the Installer pre-flights the compose download, the image pull and the rollback re-tag while it is still alive, and aborts harmlessly if any step fails. By the time the helper container starts, both the update and the rollback are guaranteed local — no network, no registry auth.
<status> / running image unknown / channel unknown — the Installer could not work out what to update to. A registry 4xx/5xx, a missing install channel (DELAMAIN_CHANNEL), or an unresolvable running image all stop the update at the check stage.DELAMAIN_CHANNEL in the environment.GET /api/system/self-update-check; when a real error remains, its error field carries it.A self-update is also refused while another job is running:
A <type> job is running (started <timestamp>) — updating delamain now would kill it. Wait for it to finish.
and A self-update is already in progress. Wait for the job to end and retry.
The Installer cannot self-update unless it was started by Docker Compose. If its container carries no compose project labels it refuses with:
This container carries no compose project labels, so its install directory cannot be found. Self-update requires delamain to have been started by docker compose.
delamain-backend does not start; docker ps shows delamain-backend-init exited with a non-zero code.
The stack runs a one-shot init container, delamain-backend-init, whose only job is to run database migrations:
npx typeorm -d dist/src/typeorm.config.js migration:run
The backend service declares depends_on: backend-init: condition: service_completed_successfully, so if the migration container does not exit 0, the backend is never started.
Read the migration container's log — it contains the database error:
docker logs delamain-backend-init
Common causes are the database not being reachable yet, or a migration that conflicts with the current schema. Correct the database, then bring the stack up again; the init container re-runs migrations.
delamain-backend is in a restart loop, or exits at once.
The backend validates its environment at boot and refuses to run on a bad one, and it requires its secrets files to exist. Two concrete exits:
Environment validation — AZURE_TENANT_ID and AZURE_CLIENT_ID are mandatory and must be non-empty; every other variable is optional but type-checked (NEST_TYPEORM_LOGGING, NEST_NODE_ENV, SWAGGER_ENABLED, BLOB_ARTIFACT_VARIANT, and the rest). A failure here aborts the boot.
Missing PostgreSQL password file — the secret bootstrap reads the password from a file on the secrets volume and refuses to continue without it:
PostgreSQL password file not found: <SECRETS_DIR>/pg.pass
The PostgreSQL service generates pg.pass into the shared secrets volume on first boot (it writes one if the file is missing or empty), and the backend reads the same file, so they agree. If the secrets volume is fresh but PostgreSQL has not run yet, or the two are not sharing the volume, the backend has no password to connect with.
AZURE_TENANT_ID and AZURE_CLIENT_ID come from the compose file and should not be removed.postgres and backend services at the same path, then start PostgreSQL first.docker logs delamain-backend.docker ps shows (unhealthy) for delamain-backend, and the deploy engine reports it in the status view.
The backend's health check hits its own endpoint:
wget -qO- http://localhost:${NEST_PORT:-8000}/api/health || exit 1The endpoint runs SELECT 1 against the database. If the database is unreachable it returns HTTP 503 with the message Database connection failed; anything other than a healthy response after the start period marks the container unhealthy.
Read the endpoint directly to see whether it is the API or the database:
curl -s http://localhost:8000/api/health
Healthy output is {"status":"ok","db":"up"}.
If the endpoint is unreachable, the API is not listening — go back to the migration and startup checks above.
If it returns Database connection failed, fix the PostgreSQL service or its credentials; the migrations and password-file symptoms on this page are the usual causes.
The stack fails to start, or a container exits immediately, because a published port is taken.
The Installer publishes its API on ${NEST_PORT:-8000}. A leftover container from a previous or bare-metal install can hold that port (the legacy migration specifically warns about old containers holding ports such as 1025, 1080 and 6333).
Find what holds the port and stop it:
docker ps --format '{{.Names}}\t{{.Ports}}'Either free the port or change NEST_PORT in the Installer .env.
If the offender is an old bare-metal container, the legacy migration is the supported way to clear it.