Operations¶
Day-to-day running of a deployed instance: what the logs say, what is recorded, what is throttled, where samples live, what to back up, and what the common failures look like.
Logs¶
The API and the worker log through apps/api/app/logging_config.py. Production
output is one JSON object per line — timestamp, level, logger, message,
correlation id, module, function and line, plus the extra fields the request
middleware attaches (method, url, status_code, duration_ms,
client_ip, user_id) and a structured exception block when one is
attached. A human-readable coloured format is used in debug mode.
Every request carries a correlation id: the middleware reuses an inbound
X-Request-ID or generates one, puts it in every log line for that request and
echoes it back in the X-Request-ID response header.
SQL_ECHO adds every SQL statement; it is deliberately
independent of DEBUG, because turning it on drowns pipeline stages in
duplicated queries.
What a run's metrics mean¶
Every finished report carries a run_summary, and five of its keys are about
the run rather than the sample. They are the numbers to read before the
verdict, because each of them says how much the verdict is standing on.
| Metric | Healthy | What a bad number means, and what to do |
|---|---|---|
validation.retries |
A handful | Every retry is one extra model round trip. A run spending dozens of them is a model that keeps answering in the wrong shape; check by_code for which rule it keeps breaking, and the agent's prompt or its model choice. |
validation.unresolved |
Empty | A producer was told what was wrong, got its turn, and did not fix it. The finding is on the record with the agent that owns it. A recurring code from one agent is a prompt problem; a recurring technique_id code across agents usually means the ATT&CK cache is stale. |
sections_without_evidence |
0 |
A report section that can name neither a ledger entry nor the finding it came from. Nothing in the console can resolve it, so a reader has no way to check it. Not an operator setting — it is a defect in whichever builder produced the section. |
evidence.failed |
Low | Tool calls that raised. A few are normal (a parser that cannot open a container); a large share against one server means that server is misconfigured or down, and the analysts spent the run guessing. |
evidence.trimmed |
0 |
Entries whose output was dropped to keep an agent inside report.evidence_budget_bytes. The call and its result still stand and are still citable; the text is gone. Persistent trimming means the budget is too small for the tools this deployment runs, or one tool is answering with a log file. |
stages |
Every stage ran |
A stage that declined records the condition that turned it off. A team where the same stage always declines is a condition that does not match the samples this deployment sees, not a broken run. |
fp_warnings is the post-run linter and changes nothing: C6 flags a section or
a TTP row with nothing citable behind it, C7 a technique id the validation loop
could not get resolved.
Staged samples¶
A tool server reached over HTTP does not share the worker's filesystem, so the
sample is delivered to it rather than named. The worker writes each delivery
into the staging directory (MALJAN_STAGING_DIR, a maljan-analysis-mcp directory under the
system temp directory by default) and the sidecars sweep it on a TTL
(MALJAN_STAGING_TTL_HOURS, 24 hours by default; 0 disables pruning).
Staging is per job. The configured directory is the base, and each job gets
one directory of its own inside it, job-<the job's id>: the sidecar's
put_sample uploads land there, so does the carved/<sha256>/ tree of
everything that job carves, and so do the sandbox captures fetched for it, in a
captures/ child. The worker removes the whole directory when the run ends —
on success, on failure and on an operator's cancel — and what a removal misses
is taken by the same TTL, which prunes a job directory whole once the newest
file in it is past the cutoff. So a nameable residue means a worker that was
killed, and it goes on its own within the TTL.
A run that lasts longer than the TTL without staging anything new keeps its directory: the job touches it while it refreshes its own owner heartbeat, which is how a sidecar sweeping a base two workers share can tell a long job from an abandoned one.
The sandbox capture. A capture is the full record of a detonation, and the
path to it is written by the model rather than by the platform — the network
analyst is told where it is and passes it to pcap_summary. It used to be
fetched into one directory under the system temp directory, shared by every job
and every worker on the host, named as a readable directory for every sidecar
started afterwards, and removed by nothing. It now lands in the job's own
captures/ directory, at 0600 inside a 0700 tree, is named as a readable
directory for that job's sidecars only, and goes when the job does. Whatever an
earlier release left in maljan-cape-pcap is swept on the same TTL; nothing
writes there any more.
Everything in it is live malware, on the same terms as SAMPLES_DIR: exclude it
from on-access scanning and keep it off shared storage. A staging write that
fails costs that server its tools for the run and is logged; it does not fail
the job.
On upgrading: the flat uploads and the single carved/ tree of the previous
release stay where they are and are swept by the same TTL. Nothing has to be
migrated or moved; a host you want clean at once can have that directory
emptied while no job is running.
Audit trail¶
Security-relevant actions are written to the audit_log table by
apps/api/app/services/audit.py: registrations and logins, sample uploads and
deletions, job creation and cancellation, sandbox-report uploads and deletions,
and every settings write, including imports. Each row carries the action, the
resource, the acting user, the client address and a details object.
Two properties are deliberate. An audit row is written on a session of its own,
so a row recording a refused action is not rolled back with the request that
was refused. And writing one is best effort: a failure is logged at ERROR and
counted rather than turning a handled 4xx into a 500. The count is visible to
admin callers on GET /api/v1/system/status.
Admins read the trail through GET /api/v1/audit/logs and one row at
/audit/logs/{log_id}.
Rate limits and login protection¶
RateLimitMiddleware applies a Redis-backed sliding window per client address
and path. It is configured from the settings store, under the API group, and
every change takes effect immediately: rate_limit_enabled,
rate_limit_requests, rate_limit_window_seconds, rate_limit_whitelist and
trusted_proxy_ips, which decides whether X-Forwarded-For is honoured.
Login attempts are bounded separately by login_max_attempts and
login_lockout_seconds.
The limiter needs Redis. When the counter store is unavailable the deep health
check reports throttle_degraded, and GET /api/v1/system/status shows the
same state.
Sample storage¶
Sample bytes live in MinIO, in the bucket named by MINIO_BUCKET
(maljan-samples by default); metadata, jobs and reports live in Postgres.
Uploads are streamed through UPLOAD_TEMP_DIR rather than the system temp
directory, and the worker mirrors each downloaded sample into SAMPLES_DIR
so the Ghidra container can read it through its read-only bind mount. Both
directories hold live malware: exclude them from any on-access scanner and
keep them off shared storage.
upload_max_bytes and upload_allowed_mime_types bound what the API accepts.
Backups¶
Two things carry state that cannot be rebuilt: the Postgres volume (pgdata)
and the MinIO volume (minio_data). The Qdrant volume holds long-term memory,
which is derived from past analyses; the ATT&CK cache volume is purely derived.
Back up MinIO by copying the bucket out with any S3 client, or by snapshotting
the minio_data volume while the stack is down.
A database dump is not enough on its own: every secret in the settings store is
encrypted with SETTINGS_ENCRYPTION_KEY, which is not in the dump. Store the
key with the backup, or the restored instance comes up with every credential
unreadable. A JSON export (see configuration.md) is a
useful complement — it is safe to keep next to the dump precisely because it
carries no credentials.
Troubleshooting¶
| Symptom | Cause and fix |
|---|---|
The process exits with bootstrap: ... |
The environment failed validation. The line names every problem at once: set the variables it lists. MINIO_SECRET_KEY left at minioadmin and a short or placeholder JWT_SECRET_KEY are refusals outside debug. |
| Compose refuses with "set X in docker/.env" | A :? variable is missing from docker/.env. |
| Every API request fails on a fresh stack, tables missing | The migrate service did not run or did not succeed. docker compose up migrate, then check its logs. |
| The worker takes no jobs | It waits on the API's healthcheck, on migrate and on Postgres, Redis and Ghidra MCP being healthy. Check docker compose ps for the one that is not healthy, and that REDIS_URL matches the password Redis was started with. |
| Jobs run but the console shows no progress | Events are published on Redis and relayed over /ws/analysis/{job_id}. The job row is still written; check Redis and the WebSocket connection rather than the job. |
| A probe fails with a connection error | Catalog defaults are localhost-shaped and wrong inside the Docker network. Use http://ghidra-mcp:8089, http://qdrant:6333, and host.docker.internal for a service on the host. |
| A probe fails with 401 or 403 | The credential in the settings store does not match the one the container was started with — most often GHIDRA_MCP_AUTH_TOKEN or QDRANT_API_KEY. |
| Secrets show as set but the provider behaves as unconfigured | SETTINGS_ENCRYPTION_KEY changed. The service logs one warning per row it cannot open and falls back to the default. Restore the old key or re-enter the secrets. |
Storage service unavailable on upload |
MinIO is unreachable, or UPLOAD_TEMP_DIR is not writable. GET /health?deep=true distinguishes the two. |
| The worker container is killed and restarts mid-analysis | It hit its 8 GB ceiling. The startup sweep repairs the job row it was holding. Lower concurrency, or raise mem_limit and WORKER_RSS_RESTART_MB together. |
| The first analysis after a rebuild is very slow | The ATT&CK embedding cache was discarded. It lives in the attck_cache volume; keep it across recreations, or prebuild it. |
| A cancelled job goes on calling the model, or the worker ignores SIGTERM | Fixed: a cancel stops the pipeline at the next node, refuses every model call and cancels the ones in flight, and SIGTERM ends the worker within 10 s, the teardown's WORKER_TEARDOWN_TIMEOUT, 2 × 10 s for its connections and 10 s at exit. A thread still blocked in a call that cannot be cancelled is left at exit, named in the log, and the process ends with status 1. |
| A source edit has no effect | The production stack bakes the frontend and runs the worker under plain arq. make fe-rebuild and make worker-restart, or use make dev-up. |