Installation fails

Most installation failures surface as an error message in the DataMind Installer UI or in the response of a deploy API call. Read the message before acting: the Installer distinguishes "nothing was changed, fix the input" from "the registry or the host failed mid-run". Where a message is followed by a reason string, that reason is the real cause.

The stages that can fail, in order, are: the shared Docker network, the deployment files, the configuration values, registry access and image pulls, and the Compose topology itself.

network unistream declared as external, but could not be found

What you see

The first docker compose up fails immediately with:

text
network unistream declared as external, but could not be found

What it means

Both the DataMind Installer stack and the DataMind OS stack attach to a Docker network named unistream and declare it external. Neither creates it: that is deliberate, because Compose 2.19.1+ refuses to use a network that already exists without its own project labels, and if either project owned it, the other would fail to deploy — and a compose down on the owner would delete a network the other stack is still using.

Fix

Create the network once, before the first docker compose up, on the host:

bash
docker network create unistream

You only do this once per host. If the network exists but was created by an old project, remove the stale one and recreate it.

Missing deployment files

What you see

An API error of the form:

text
Missing deployment files: docker-compose.yml, .env.unified.template, .env — docker-compose.yml: ...

The trailing — ... part, when present, carries each file's own failure reason.

What it means

The Installer keeps three generated files in its deployment directory and restores any that are missing before it will run a deployment: the published docker-compose.yml, the .env.unified.template it downloads from artifact storage, and the generated .env. If one cannot be restored, the deployment is blocked rather than run against a partial set.

The most common underlying reason is a failed download; see the next two symptoms.

Fix

Missing required configurations

What you see

Starting a deploy or generating .env returns:

text
Missing required configurations: <CODE>, <CODE>, ...

What it means

The env template marks some configuration codes as required, or requiredIf another code is set. Any required code whose value is still empty blocks the deploy, so the platform is never started with an incomplete environment.

Fix

  1. Open the configuration list and fill every code named in the message.

  2. Or reseed the configuration rows from the template, which creates any rows that are missing:

    text
    POST /api/configs/seed
  3. Or, for a legacy install, import values from the old Jenkins configuration (requires the Jenkins volume mount):

    text
    GET  /api/configs/import/jenkins/preview   (dry run, no writes)
    POST /api/configs/import/jenkins

Values are set with PATCH /api/configs/values, passing an array of { "code": "...", "value": "..." }. Only administrators can write.

The env template or compose file cannot be downloaded

What you see

One of:

text
Failed to download docker-compose.yml from Azure: Blob request failed (<status> <statusText>) for <url>
Failed to download env template from Azure: Blob request failed (<status> <statusText>) for <url>
Failed to download the delamain compose file from Azure: ...

What it means

The Installer fetches these artifacts from Azure Blob Storage over an authenticated request. A Blob request failed (...) means the storage call itself was rejected or unreachable — usually an expired or wrong credential, or no outbound network from the host.

Fix

Azure authentication fails

What you see

Any of these messages, depending on the stage:

text
Azure client secret is not configured
Stored Azure client secret could not be decrypted (encryption key changed?). Re-enter it.
AAD token request failed (<status>): <reason>
ACR exchange failed (<status>): <reason>
ACR access token request failed (<status>): <reason>
Failed to fetch data from Azure Registry: ...
Catalog request failed: <status>
Failed to fetch tags for repository <repo>: ...

What it means

The Installer authenticates to Azure with a single stored client secret; from it it obtains an AAD token, exchanges that for a registry refresh token, and uses those for the registry and for artifact storage. Each message maps to one link in that chain:

Fix

  1. Re-enter the Azure client secret in the configuration.
  2. If you saw the could not be decrypted message, the encryption key changed. Re-enter the secret; if the underlying key material is gone, the stored value cannot be recovered and re-entry is the only path.
  3. If AAD rejects the secret, confirm it is current in Entra ID and belongs to the configured application (AZURE_CLIENT_ID / AZURE_TENANT_ID). The Installer surfaces the AAD error code and the first line of its description to help you match it.
Note

AZURE_TENANT_ID and AZURE_CLIENT_ID are provided by the compose file and must be non-empty; if either is missing the backend refuses to start (see Updates and startup).

Images are missing or a pull fails

What you see

Starting the stack before images exist:

text
Missing local images: <image>, <image>. Run Update to pull them first.

Or, during a pull, a per-service failure event:

text
Registry <host>: <reason>

What it means

docker compose up is run with --pull never, so the Installer refuses to start a service whose image is not already on the host — it will not silently pull mid-start. The second message means the registry was unreachable or unauthorised for that host while pulling.

Fix

  1. Run the update/pull action first, then start:

    text
    POST /api/deployment/pull

    The deploy action (POST /api/deployment/deploy) downloads the artifacts, pulls images, then composes — use it when you want the whole chain in one job.

  2. For Registry <host>: <reason>, treat the reason as the Azure auth or connectivity problem above. A failed pull records whether a local copy already existed, which tells you whether the service can still start from the previous image.

The topology rejects a service name

What you see

text
Unknown compose services: <name>, <name>

or, when restarting:

text
These are one-shot init containers and cannot be restarted: <name>, <name>. They re-run automatically on a full deploy.

What it means

Service names are validated against the compose file's actual services. Unknown compose services means a name you asked for is not in the compose topology (often a typo or a service renamed between releases). The one-shot message means you targeted an init container — a service that runs once and exits, like a migration step — which has no business being restarted.

Fix

A job is already running

What you see

text
A <type> job is already running (started <timestamp>). Wait for it to finish.

What it means

The Installer runs one deployment job at a time. A deploy, pull, compose-up, restart, stop or kill job already in progress blocks a new one, so two jobs cannot race over the same containers.

Fix

Wait for the running job to finish, or cancel it:

text
GET    /api/deployment/jobs           (list active and recent jobs)
GET    /api/deployment/jobs/<jobId>   (status)
DELETE /api/deployment/jobs/<jobId>   (cancel — admin only)

After a cancel that lands during the compose phase, some containers may already have been recreated; finish with a Start or remove them with a Stop.

Memory limits could not be computed

What you see

text
Memory limits could not be computed: <reason>

What it means

Before a deployment, the Installer sizes each service's memory limit from the host's total RAM and the published sizing schema. Some failures are non-blocking and only produce a warning in .env; the message above is emitted only for a schema error, which does block. Typical reasons are a host-class table with no matching entry for this machine's RAM (a catch-all entry is required), a service that references an unknown host class, or a service rule that has neither fixed_mb nor pct.

Fix

Note

If /proc/meminfo is unreadable the Installer assumes the host is not the Linux VM the limits describe and emits no limits. That is non-blocking — it is reported as a warning, not as the error above.

The deploy engine cannot reach the Docker daemon

What you see

Deploy actions fail as soon as they try to drive Docker; no containers are created.

Note

The exact wording of this failure is produced by the Docker client library and is not fixed by the DataMind Installer, so it is not quoted here.

What it means

The DataMind Installer drives the host Docker daemon through the mounted socket. On a standard (rootful) engine that is /var/run/docker.sock; a rootless engine exposes a different socket. If the wrong socket is mounted, or the process lacks access to it, every Docker call fails.

Fix

Set DOCKER_SOCK in the Installer .env to the correct host socket before starting the stack. For rootless Docker, point it at the rootless socket (find your UID with id -u):

bash
DOCKER_SOCK=/run/user/1000/docker.sock
Warning

Socket access is effectively host root. Mount the real socket and expose it only to the Installer, not to untrusted workloads.