Local llama.cpp¶
A llama.cpp server — or a fork such as ik_llama.cpp — is reached through the
OpenAI-compatible provider, pointed at the server's
/v1 address. It is how the Benchmark's local model,
Qwen3.6-35B-A3B on ik_llama.cpp, was served.
Settings¶
api_key is needed only if the server was started with one. Under compat =
auto, a loopback, link-local, .local or private host is taken to be
llama.cpp; a local server reached through a public name needs compat =
llama_cpp set explicitly.
What a llama.cpp endpoint is sent¶
| Setting | Sent as | Default |
|---|---|---|
core.llm.openai.repetition_penalty |
repetition_penalty |
1.0 (no effect). Around 1.15 breaks a small reasoning model out of ATT&CK id-recall loops. |
| The request's output cap | n_predict as well as the standard field |
always, because llama.cpp ignores max_completion_tokens |
core.llm.openai.disable_thinking |
chat_template_kwargs.enable_thinking=false |
false |
core.llm.openai.dry_multiplier, dry_base, dry_allowed_length, dry_penalty_last_n |
the DRY sampler's fields, each only when set | unset |
The DRY sampler penalises a token that extends a sequence already repeated in the context.
On a constrained host, consider disable_thinking
A reasoning model can spend its whole decode budget inside its thinking,
which the server strips into reasoning_content, leaving an empty answer.
What is learned from the server¶
- The context window.
GET /propsreports the window the server was started with, and that window sizes how much of a tool answer a model may read and its output caps.core.llm.openai.context_sizestates it instead for a server that does not report it. - The slots.
/propsalso reports how many requests the server serves at once. Undercore.llm.parallel_analysts = auto, a local server with one slot runs the analysts one after another, because concurrent analysts on a single-slot server clobber each other's per-slot state and every step re-processes its prompt. More than one slot runs them in parallel. - Its own timings. The prompt and generation rates llama.cpp reports on every answer are recorded, and the call deadlines are sized from them.
Structured output is never attempted against a custom base_url: a local
server handles it badly, and the plain text path does the same job.
Two servers, two agents¶
A per-agent entry can name its own base_url, so two agents can sit on two
different local servers while core.llm.openai.base_url stays the fallback
for everything that sets none. The llama.cpp extras and the structured-output
decision follow the endpoint the agent will actually call. See Per-agent
models and fallbacks.
The server the measurements used¶
For the record, the local model server behind this repository's measurements ran ik_llama.cpp on loopback port 8080 as:
llama-server -m Qwen3.6-35B-A3B-IQ3_K_R4.gguf -c 131072 -t 16 -fa on \
-ctk q8_0 -ctv q8_0 -ngl 999 \
-ot 'blk\.([1-3][0-9])\.ffn_(up|gate|down)_exps=CPU' \
--context-shift on --jinja
Host-specific helpers — a launcher, a memory guard — live outside the repository; see Development.