Deployment¶
How the stack is assembled, what it refuses to start without, how to check that it is healthy, and how to upgrade it. For the meaning of individual settings see configuration.md.
The compose stack¶
docker/docker-compose.yml
is the production
shape. docker/docker-compose.dev.yml is an overlay that swaps the frontend to
next dev and supervises the worker so source edits take effect; make dev-up
applies both.
| Service | Image | Purpose |
|---|---|---|
postgres |
postgres:16-alpine |
Relational store. Healthchecked with pg_isready. |
redis |
redis:7-alpine |
Queue, events, rate-limit counters. Password-protected; the healthcheck asserts an actual PONG. |
qdrant |
qdrant/qdrant:v1.18.2 |
Vector store, pinned to match the client version. |
minio |
minio/minio:latest |
Sample storage, with its own console. |
ghidra-mcp |
built from external/ghidra-mcp |
Static analysis engine. Capped at 6 GB memory and swap, because the JVM's own limit does not bound the container. GHIDRA_JAVA_OPTS, GHIDRA_MEM_LIMIT and GHIDRA_RESTART in docker/.env change the JVM options, the cap and the restart policy; see Running Ghidra lighter in configuration.md. |
migrate |
maljan-backend |
One-shot alembic upgrade head. Runs to completion before the API and the worker start; on an upgrade, migrate before restarting either, since a stored built-in team written before a seeding revision loads under a renamed key until the revision has run. |
backend-api |
maljan-backend |
The FastAPI service. |
backend-worker |
maljan-backend |
The arq worker. Capped at 8 GB memory and swap; WORKER_RSS_RESTART_MB makes it exit between jobs before it gets there, and restart: unless-stopped brings it back. |
frontend |
built from docker/Dockerfile.frontend |
The console. |
The three backend services share one build and one image tag, so Compose
builds maljan-backend once rather than once per service.
Ordering is explicit: migrate waits for a healthy Postgres, the API waits for
migrate to exit successfully and for Postgres, Redis and Ghidra MCP to be
healthy, and the worker waits for the API's own healthcheck on top of that.
Required secrets¶
Every secret Compose substitutes is declared with :?, so the stack refuses to
start — with the variable named — while any is missing. There is no baked-in
default for any of them.
| Variable | Used by |
|---|---|
POSTGRES_PASSWORD |
The database, and the DATABASE_URL Compose assembles. |
REDIS_PASSWORD |
--requirepass on Redis, and the API and worker URLs. |
QDRANT_API_KEY |
QDRANT__SERVICE__API_KEY on Qdrant. |
MINIO_ROOT_USER / MINIO_ROOT_PASSWORD |
MinIO, and the API's MinIO credentials. |
GHIDRA_MCP_AUTH_TOKEN |
The bearer token the Ghidra MCP container requires. |
JWT_SECRET_KEY |
Signs API session tokens. |
SETTINGS_ENCRYPTION_KEY |
Encrypts secrets in the settings store. |
Optional knobs with defaults: BIND_ADDRESS (127.0.0.1), the published
ports, RUN_MIGRATIONS_ON_STARTUP, CORS_ORIGINS, COOKIE_SECURE,
WORKER_RSS_RESTART_MB, MALJAN_EMBED_THREADS and MALJAN_EMBED_BATCH.
Every published port binds to BIND_ADDRESS, which the example file sets to
127.0.0.1. The stack is unreachable from the network until that is changed
deliberately, behind a firewall or a reverse proxy you control.
The Ghidra samples path¶
The worker copies each sample into data/samples/.work/ on the host, and the
ghidra-mcp service mounts ../data/samples at /data/samples inside its
container (read-only). GHIDRA_CONTAINER_SAMPLES_PATH is the path INSIDE the
Ghidra container, /data/samples with this Compose file, and never the host
directory. The API and worker services set it; a worker started outside
Compose (bootstrap.env, a supervisor, a shell) needs it set to
/data/samples too, or left unset, since that is the default.
Set to the host path, Ghidra answers every load with File not found: ...,
the Ghidra agent stops before its first model turn with "Ghidra could not open
the job's sample: ...; check GHIDRA_CONTAINER_SAMPLES_PATH / the container
mount", and the rest of the team runs without it. The worker's first lines
say which path it uses and where it came from:
Rotating the JWT signing secret¶
Three steps and one wait. The wait is the part a rotation is usually forgotten in, so the end of it is something you write down.
- Move the current
JWT_SECRET_KEYtoJWT_PREVIOUS_SECRET_KEYandJWT_KEY_IDtoJWT_PREVIOUS_KEY_ID, put the new secret inJWT_SECRET_KEYand a new id inJWT_KEY_ID, and setJWT_PREVIOUS_SECRET_NOT_AFTERto the moment the old secret stops being accepted — an ISO-8601 timestamp, read as UTC when it carries no offset. Make it a little longer than the refresh-token lifetime (jwt_refresh_token_expire_days), because that is the oldest token still in circulation. Restart the API and the worker. The moment is optional and a previous secret without one is accepted with no end — the startup check warns about that at every start until you set it or clear the previous secret — but a value that does not read as an ISO-8601 moment is refused at startup rather than taken as "no end". - Wait. New tokens are signed with the new secret and carry the new
kid; tokens already issued keep working until they expire.GET /api/v1/system/statusreports the rotation to an admin caller underjwt_grace_secret— the previouskid, when it lapses, whether it is still accepted and whether an end was written down at all (bounded) — and the API logs the same line at every start. - After the moment passes, a token signed with the old secret is refused and
jwt_grace_secret.acceptedreads false; the startup check then warns until you clearJWT_PREVIOUS_SECRET_KEYandJWT_PREVIOUS_SECRET_NOT_AFTER.
Nothing rotates on its own. The point of the timestamp is that a secret you forget stops being honoured rather than being honoured for the life of the deployment. See security.md for where signing sits among the rest of the auth posture.
Health¶
GET /health and GET /healthz are the same endpoint under two paths, so both
a bare and a Kubernetes-style liveness probe work without extra configuration.
- Without parameters it performs no I/O: it returns
status, the service name and version, and aconfigblock reportingbootstrapandencryption, both of which the process already satisfied before it could serve anything. A liveness probe must not restart the API because Postgres blinked. GET /health?deep=trueprobes Postgres, Redis, MinIO and Qdrant concurrently, each with a three-second budget, and reports every result as data.statusbecomesdegradedwhen Postgres or Redis is unreachable anddegraded_optionalwhen only MinIO or Qdrant is, so use the deep form for readiness and the bare form for liveness.
Operational detail — throttle state and dropped audit rows — lives on
GET /api/v1/system/status, where those fields are returned to admin callers
only; the health endpoint publishes only the public readiness bit.
Outside Compose¶
The API and the worker need nothing but the bootstrap contract in their process environment, so any orchestrator that can inject environment variables can run them.
- API:
uv run uvicorn app.main:app --host 0.0.0.0 --port 8000from the repository root, without--reload. - Worker:
uv run arq app.worker.analysis_worker.WorkerSettings. - Kubernetes: put the contract in a Secret and a ConfigMap and inject it as
environment variables.
SETTINGS_ENCRYPTION_KEYmust be identical for the API and the worker deployments, or the worker cannot open the credentials the console stored. Point the liveness probe at/healthzand the readiness probe at/health?deep=true. Run migrations as a Job before the rollout and leaveRUN_MIGRATIONS_ON_STARTUPfalse, so replicas do not race the same upgrade. - systemd: an
EnvironmentFile=pointing at a root-owned copy ofbootstrap.envis enough for both units; nothing else is read from disk.
Behind a reverse proxy, set CORS_ORIGINS to the console's real origin, keep
COOKIE_SECURE true, and populate the trusted_proxy_ips setting so the rate
limiter honours X-Forwarded-For only from your proxy.
VirusTotal over stdio (optional)¶
The virustotal built-in needs nothing installed: it is VirusTotal's own
server over HTTP, and registering an agent token from Settings → Setup guides
is the whole setup. Install the package only for submit_local_file, the
upload-by-path tool that exists solely when the server runs on the host that
holds the sample:
uvx --python 3.12 vt-mcp==0.8.4 runs the same release without installing it.
The command is stdio-only and takes no flags; it reads the agent token from
VTAI_TOKEN (or VTAI_TOKEN_FILE) in its own environment. Add it as a second
tool server with transport stdio and put the variable name in env_allow, so
the token reaches the child from the worker's environment rather than being
written into a setting the console echoes back.
A server added this way is started with exactly the names it lists. A built-in
is the one case where that field has a floor: analysis and network are
always passed the sample roots and the staging directory, whatever the stored
registry holds, because a child without them refuses the run's own sample. See
configuration.md.
Upgrades and migrations¶
The API does not migrate on startup unless RUN_MIGRATIONS_ON_STARTUP is set,
which it should not be in production.
- On the compose stack the
migrateservice reruns on the nextupand exits immediately when there is nothing to apply:docker compose up migrate. - Elsewhere,
make migratefrom the repository root. It sourcesbootstrap.envwhen present and runsalembic upgrade headfromapps/api;DATABASE_URLmust be reachable from where you run it, so the compose-internalpostgres:5432will not do.
After a pull that adds a revision, migrate before restarting the services. The
production stack bakes the frontend into an image and runs the worker under
plain arq, so neither picks up a source edit on its own: make fe-rebuild
rebuilds and replaces the frontend container, and make worker-restart
restarts the worker.
There is no automatic import of a previous .env deployment. An instance
upgrading from one enters its settings once in Settings → Configuration, or
imports a JSON export from another instance.