JESVS

One Word Per Intent: A Launcher That Knows Why You're Asking

Or: how a benchmark campaign’s most durable output turned out to be seven lines of shell.

A terminal window running –role quality, with seven labeled cables fanning out to the role icons: daily (speedometer), quality (crown), write (pen), agent (robot arm), rp (theater mask), dense (tower), dense-unc (open padlock)

Somewhere around the fifteenth model, my launcher had a knowledge problem.

serve-qwen.sh started life the way these things do: a table of model files and their best tensor split, a handful of flags, exec llama-server. Every model I benched got a key. That worked while “best config” meant one number — a split ratio. Then the optimization campaign added MTP drafts and KV-cache rules. Then the agent-model round added models that think. Then the coding shootout proved that whether a model thinks can be worth six points on the same benchmark, and that some models burn their entire token budget doing it wrong.

The launcher knew what to run. All the knowledge about how — which mode, which sampler, which failure to avoid — lived in markdown tables I would have to re-read before every launch. Which is a polite way of saying it lived nowhere.

The footgun that broke the camel

The specific incident: I have an uncensored variant of my daily-driver MoE that scores a perfect 35/35 on my test suite — with thinking on. With thinking off, it scores 29/35. Same weights, same quant, same speed. The difference is one field in the request: chat_template_kwargs: {"enable_thinking": true}.

Now open the server’s web UI — the built-in one on port 8080, which is genuinely nice for poking at a model — and ask it something. The browser sends no chat_template_kwargs, because browsers send what the chat widget tells them and nothing else. The model runs no-think. 29/35 mode, silently, forever. Nothing errors. Nothing looks wrong. The model is just quietly six points worse than the one I benchmarked.

That’s the gap between a model and a behavior. My launcher was serving models. What I actually wanted to serve was behaviors.

The mechanism that makes it possible

The fix hinges on a llama-server detail that doesn’t get enough credit: CLI sampling flags are defaults, not overrides. --temp, --top-p, --top-k, --presence-penalty — and the one that matters most here, --chat-template-kwargs '{"enable_thinking": true}' — apply whenever the client omits the parameter. Any client that sends its own values still wins.

That inverts the usual configuration problem. Zero-config clients — web UI, a lazy curl, an SDK with defaults — stop being the risk case and become the correct case: they get exactly what the benchmark said this model needs. Opinionated clients keep their opinions. You get sane defaults and full override in the same mechanism, with zero middleware.

Roles: presets with a thesis

So the launcher grew a second, smaller table:

# role|model key|extra server flags (sampling/template defaults)|note
ROLES=(
  "daily|moe||fastest all-round: 34/35 @ ~118 tok/s, no-think by design"
  "quality|ablmoe|THINK|best score (35/35) — thinking defaulted ON"
  "write|hemm|--temp 0.5 --top-p 0.8 --top-k 20 THINK|prose: thinking + temp 0.5 (measured best)"
  "agent|glm47|THINK|agentic coding MoE, reasoning on"
  "rp|joyfox||roleplay MoE; no-think only — its family dies thinking past 1024"
  ...
)

--role quality resolves to the model and its behavior: thinking baked on, server-side, for every client that doesn’t say otherwise. --role write serves the creative finetune at the settings that measured best for prose (temp 0.5 with thinking — a combination yesterday’s matrix proved out at 35/35, and which happens to be the vendor’s own thinking-mode recipe with the temperature dropped). --role rp serves the roleplay model without THINK, because its entire base generation exhausts small thinking budgets and returns empty strings — a cliff the coding post documented in detail. The role encodes the finding. The finding stops being trivia.

Two small design decisions worth keeping if you build one of these:

The THINK sentinel. The flag’s value is JSON, and JSON in a pipe-delimited config row is a quoting crime scene. The table stores the word THINK; the expansion loop turns it into the flag plus its JSON as a single argv element. No eval, no escaped pipes, and the table stays readable.

Flag precedence is a feature. Model-table flags (like my 80B model’s -ncmoe 48 --threads 8) apply first, role flags second, anything after -- last. Anything the user types explicitly beats both. Presets that can’t be overridden aren’t presets; they’re cages.

What it looks like now

$ ./serve-qwen.sh --role write
role  : write — prose: thinking + temp 0.5 (measured best)
model : .../hemmingway-1-q4_k_m.gguf
rflags: --temp 0.5 --top-p 0.8 --top-k 20 --chat-template-kwargs {"enable_thinking": true}
cfg   : -sm layer -ts 5,3 -c 16384 -fa on -ctk f16 -ctv f16 ...

Fifteen model keys still work exactly as before — roles are a layer, not a replacement, and the benchmark harnesses still hit the model keys directly because scripts shouldn’t have intents. But for the human standing at the terminal at some hour where nobody re-reads markdown, daily / quality / write / agent / rp is the whole interface. One word per intent, and each word is load-bearing: it carries a model, a mode, a sampler, and a failure mode somebody already hit so you don’t have to.

The part I like most

Every benchmark campaign produces tables. Score tables, speed tables, mode matrices. And tables are dead knowledge unless something reads them at decision time — usually a human, usually from memory, usually wrong at 2am. Roles are just a compiler from “what I measured” to “what happens by default.” The benchmark findings didn’t change. Their activation energy did.

There’s a version of this idea that ends in a TUI with VRAM gauges and live tok/s sparklines, and I’ll probably build it. But the interesting layer isn’t the interface — it’s the data model. The moment your config says why instead of just what, every interface you build on top of it gets smarter for free: a picker can show scores, a launcher can warn when a combination measured badly, a monitor can badge which mode it’s in. Intent is the schema. The terminal is just where it prints.

If your own launcher serves more than a couple of models, steal the pattern: a second table, one sentinel, defaults that lose politely to explicit requests. The hardest part is already done — you know what your models need. You benchmarked them, after all. This is just letting the launcher read your notes.

← volver a posts
↑