Maljan¶
Multi-agent malware analysis in which every claim in the report resolves back to the tool call it came from.
Maljan connects LLM agent teams to malware-analysis tools: signature scanners, rule engines, binary and capture readers, a sandbox, an ATT&CK catalogue. The output is a report against MITRE ATT&CK and a STIX 2.1 bundle.
-
Quick start
From a clean checkout to a first finished analysis on the compose stack: the two configuration files, the first account, the first run.
-
Console
The web console: submitting samples, following a run as a conversation, reading the evidence ledger and the report.
-
LLM providers
OpenAI-compatible endpoints (DeepSeek included), Anthropic, Gemini, Ollama and a local llama.cpp server, per agent if you like.
-
Tools
Ghidra, radare2, capa and YARA, CAPE and Triage sandboxes, VirusTotal, the built-in analysis sidecars and any MCP server of your own.
Key capabilities¶
- Teams are configuration, not code. A team is an ordered list of stages — triage, the passes the sample is worth, a debate, a verdict and a report — each with a condition, so a team applies to a sample rather than being written for one. See Teams and profiles.
- An evidence ledger with citable ids. Every
tool call is written to the ledger as
ev_NNNN, and every section of the report names the ids it was built from. See the evidence ledger. - The agent decides, the code says what is wrong. Nothing rewrites a claim, a technique id, a confidence or an attribution behind the producer's back: a problem goes back to the producer as feedback, with one turn to fix it. See validation loops.
- Format-agnostic routing. The sample's own bytes decide its route (PE, ELF, Mach-O, APK, DEX, documents, scripts, archives); a format nothing recognises gets the neutral path and a report that says so. See format routing.
- Reports that other systems can act on. Markdown, HTML, PDF, the full JSON document, STIX 2.1, an IOC feed with a publish decision on every row, and detection drafts. See Reports and exports.
- Configured from the console. A small bootstrap environment, and every other setting in a store edited from the web console, with setup guides and connection tests. See Configuration.
Multi-agent architecture in brief¶
A run is one team executing over one sample. The deterministic triage pack establishes the facts first — what the file is, its hashes and signature, what the rule engines and the reputation services say — and writes each to the evidence ledger. The analysts of each stage then call the tools they were given, the debate puts their positions against each other, the judge states the verdict, and the report is assembled from the ledger and the claims that cite it.
The components, the job lifecycle and every stage are described in Architecture.
A quick example¶
Upload a sample and start an analysis with the default team.
- Open the console at
http://localhost:3000and sign in. - Upload the sample and start an analysis.
- Watch the CONVERSATION tab: every agent's messages, its tool calls and the verdict, with each stage filling in on the analysis header.
- Read the report from the evidence up: every section carries the ledger ids it was built from, and each opens that call on the EVIDENCE tab.
TOKEN=$(curl -s -X POST http://localhost:8000/api/v1/auth/login \
-H 'Content-Type: application/json' \
-d '{"email":"you@example.com","password":"..."}' | python -c 'import json,sys; print(json.load(sys.stdin)["access_token"])')
curl -X POST http://localhost:8000/api/v1/samples/upload \
-H "Authorization: Bearer $TOKEN" -F 'file=@sample.exe'
curl -X POST http://localhost:8000/api/v1/jobs \
-H "Authorization: Bearer $TOKEN" -H 'Content-Type: application/json' \
-d '{"sample_id":"<id from the upload>"}'
Analyse malware only in isolated, authorised environments
Maljan handles live malware. Run it, and any sandbox it submits to, only on hosts and networks you are authorised to use for that purpose, isolated from production systems and data. A sandbox detonates the sample; the mock sandbox, the default, executes nothing. Uploading a sample to an external service such as VirusTotal publishes it; that tool ships unticked. See Security.
All documents¶
Each document is written against the code in this repository: the bootstrap
contract in apps/api/app/bootstrap.py, the settings catalog in
src/maljan/core/, the routers under apps/api/app/api/v1/ and the compose
stack in docker/.
| Document | What it covers |
|---|---|
| Quick start | Prerequisites, the two configuration files, starting the stack, first login and first analysis. |
| Console | The navigation tree, the analysis tabs, the Conversation view and how the console follows a run. |
| Running an analysis | Samples, jobs, the per-job overrides, the sandbox choice and following a run. |
| Teams and profiles | The teams that ship, stages and conditions, the measurement baseline, writing a team of your own. |
| Reports and exports | The report, its renderings, STIX, the IOC feed, detection drafts and comparing two runs. |
| LLM providers | The four provider backends, per-agent models and fallbacks, and a page per provider. |
| Tools | Static providers, sandboxes, VirusTotal, the built-in sidecars and generic MCP servers. |
| REST API | Router groups, the evidence endpoint, the run-summary fields, the stage events on the WebSocket, authentication and pagination. |
| MCP servers | Transports, remote sample delivery, and the conventions that make a tool server a good citizen of a run. |
| CI | Driving Maljan from a pipeline with an API key. |
| Configuration | The bootstrap environment contract, Settings → Configuration, the setup guides, connection probes, JSON export and import, secret storage. |
| Architecture | Components, the request and job lifecycle, teams as stages, agents and their tools, the evidence ledger, validation loops and report assembly. |
| Security | Authentication, roles, API keys, secret encryption, what an export leaves out, CORS and cookie flags, vulnerability reporting. |
| Deployment | Compose services and healthchecks, required secrets, health endpoints, Kubernetes and systemd notes, upgrades and migrations. |
| Operations | Logs, what a run's metrics mean, staged samples, the audit trail, rate limits, backups and a troubleshooting table. |
| Development | Repository layout, make targets, the test suites, CI jobs and the branch workflow. |
| Benchmark | The benchmark: samples, models, the scoring key, results across iterations and against a human report, the findings beyond it, the false-positive control and limitations. |
| Paper | Where the published evaluation, its harness and its fixtures live. |
Images used by these documents and by the top-level README.md live in assets/.
Historical design record¶
The specs, plans and superpowers notes that once lived under docs/specs/,
docs/plans/ and docs/superpowers/ are no longer carried in the working
tree; read them with git log -- docs/plans docs/specs docs/superpowers.