Updates and startup

The DataMind Installer updates itself through a detached helper container, and reconciles the outcome on the next boot by comparing image digests. Startup failures are almost always one of four things: migrations, port conflicts, missing secrets, or a container that starts and immediately fails its health check. Each symptom below tells you where to look first.

The self-update did not take effect

What you see

After a self-update, GET /api/version still reports the old digest, and GET /api/system/self-update-check shows a lastUpdate whose status is failed. When this happens the helper has already tried its one automated rollback.

What it means

The self-update runs in a fixed, short chain: pull the new image and stage the compose file (pre-flight), then switch the Installer onto the new image, and if that fails, roll back exactly once, then stop. There are no retries and no down-migrations — reverting a schema is more destructive and less tested than going forward, so a genuinely breaking migration makes a release non-rollbackable instead.

The Installer reconciles what happened on its next boot by comparing the running image digest to the digest it recorded before the update. It reads the audit record, not the helper's status file (the helper's working directory is not mounted into the backend, so the new process cannot read it).

Fix

Read the recorded detail, which states one of two outcomes:

text
Target <digest> was not applied — running <digest>. The helper rolled back; the schema is unchanged.
text
ROLLED BACK ONTO A CHANGED SCHEMA: target <digest> was not applied — running <digest> — but the database
migrated from "<old migration>" to "<new migration>" before the rollback. This build must tolerate the
new schema; verify manually.

Check the knob that most often fails the switch:

text
GET /api/system/self-update-check

The response's updateAvailable, currentDigest, remoteDigest, composeChanged and error tell you whether a newer image was seen at all, and whether the compose file also drifted.

Reading the helper's own evidence

The helper logs its entire run, including why a switch failed:

bash
docker logs delamain-self-updater

The messages it emits, in order, are:

text
[self-update] backing up <compose-file> to <compose-file>.pre-update
[self-update] installing the new compose file from <staged-compose>
[self-update] switching project <project> (service backend) onto tag <tag>
[self-update] switch failed — attempting the one automated rollback
[self-update] rolled back to <tag> — the update did NOT stick
[self-update] rollback failed too. delamain is down and needs manual recovery over SSH:
[self-update]   fetch the delamain compose from blob storage, then: docker compose pull && docker compose up -d
[self-update] helper logs remain available via: docker logs delamain-self-updater

It also writes a one-line status file on the host — a breadcrumb for an operator, not for the Installer:

text
<WORK_DIR>/.self-update-status    →  ok | rolled-back | failed

The previous compose file is kept beside the live one as <COMPOSE_FILE>.pre-update until the next update overwrites it.

The switch timed out

What you see

The helper's switch did not complete, and it either rolled back or reported failure.

What it means

The switch runs docker compose up -d --wait with a 300-second wait timeout. --wait blocks on the service's health check, so a container that starts and immediately dies counts as a failure. The new compose file has already replaced the live one at this point; the rollback restores the backup.

Fix

The self-update was refused before it started

What you see

One of:

text
Update preparation failed; nothing has changed: <reason>
Could not start the updater; nothing has changed: <reason>
Could not determine which image to update to — refusing to proceed
Registry returned <status> for <tag>
The running image is unknown, so there is nothing to compare
Could not determine the install channel from <image>; set DELAMAIN_CHANNEL
Already running the newest image on this channel — nothing to do.

What it means

These are the safe refusals. Every one happens before anything changes: the Installer pre-flights the compose download, the image pull and the rollback re-tag while it is still alive, and aborts harmlessly if any step fails. By the time the helper container starts, both the update and the rollback are guaranteed local — no network, no registry auth.

Fix

Note

A self-update is also refused while another job is running: A <type> job is running (started <timestamp>) — updating delamain now would kill it. Wait for it to finish. and A self-update is already in progress. Wait for the job to end and retry.

Warning

The Installer cannot self-update unless it was started by Docker Compose. If its container carries no compose project labels it refuses with: This container carries no compose project labels, so its install directory cannot be found. Self-update requires delamain to have been started by docker compose.

The backend never comes up because migrations failed

What you see

delamain-backend does not start; docker ps shows delamain-backend-init exited with a non-zero code.

What it means

The stack runs a one-shot init container, delamain-backend-init, whose only job is to run database migrations:

bash
npx typeorm -d dist/src/typeorm.config.js migration:run

The backend service declares depends_on: backend-init: condition: service_completed_successfully, so if the migration container does not exit 0, the backend is never started.

Fix

Read the migration container's log — it contains the database error:

bash
docker logs delamain-backend-init

Common causes are the database not being reachable yet, or a migration that conflicts with the current schema. Correct the database, then bring the stack up again; the init container re-runs migrations.

The backend starts and immediately exits

What you see

delamain-backend is in a restart loop, or exits at once.

What it means

The backend validates its environment at boot and refuses to run on a bad one, and it requires its secrets files to exist. Two concrete exits:

Fix

The container is unhealthy

What you see

docker ps shows (unhealthy) for delamain-backend, and the deploy engine reports it in the status view.

What it means

The backend's health check hits its own endpoint:

bash
wget -qO- http://localhost:${NEST_PORT:-8000}/api/health || exit 1

The endpoint runs SELECT 1 against the database. If the database is unreachable it returns HTTP 503 with the message Database connection failed; anything other than a healthy response after the start period marks the container unhealthy.

Fix

  1. Read the endpoint directly to see whether it is the API or the database:

    bash
    curl -s http://localhost:8000/api/health

    Healthy output is {"status":"ok","db":"up"}.

  2. If the endpoint is unreachable, the API is not listening — go back to the migration and startup checks above.

  3. If it returns Database connection failed, fix the PostgreSQL service or its credentials; the migrations and password-file symptoms on this page are the usual causes.

A port is already in use

What you see

The stack fails to start, or a container exits immediately, because a published port is taken.

What it means

The Installer publishes its API on ${NEST_PORT:-8000}. A leftover container from a previous or bare-metal install can hold that port (the legacy migration specifically warns about old containers holding ports such as 1025, 1080 and 6333).

Fix