JESVS

Three Agentic Coders Walked Into My 28 GB Rig

Or: how a reasoning model lost six points to my own benchmark harness, an 80B model fit into RAM that “had no room,” and the smaller quant ran slower than the big one.

Yesterday’s post left my mismatched GPU pair (an RTX 5060 Ti and an RTX 3060 that have no business sharing a case) serving a 35B-A3B MoE at 118 tok/s with the best score of the whole Qwen3.8 study. Today’s question was obvious: 2026 shipped a new generation of agentic coding models. Did my incumbents just get obsolete?

Finding out meant three things: figuring out which of the hyped models can physically run on 28 GB of VRAM and 64 GB of RAM, downloading ~115 GB of them, and running them through the same 39-test suite at the same seed as the incumbents. Same instrument, same patient. Here’s what happened.

The shortlist, or: everything cool is 320B

The 2026 “best open agentic coder” lists are dominated by models that need server racks. GLM-5.3-Flash is 320B-A18B — lovely, MIT-licensed, and roughly four times my machine’s total memory at Q4. Kimi K2.5 and DeepSeek-V3.2 are 1T-class. Filtering to what survives contact with a desktop, three names kept coming back:

  • Qwen3-Coder-Next — 80B total, ~3B active (MoE). The local-deployment flagship; “Sonnet-class coding” claims. 49.6 GB at Unsloth’s Q4_K_XL, which means CPU offload — the interesting one.
  • GLM-4.7-Flash — 30B-A3B MoE from Z.ai, built for local agentic workflows. 17.5 GB at Q4_K_XL — fits all-GPU with room to spare.
  • Devstral-Small-2 (2512) — Mistral × All Hands, a dense 24B purpose-built for software-engineering agents, Apache 2.0. 14.5 GB.

A note on scoring: my suite has 39 tests (arithmetic, logic, factual, JSON, code-executed-with-asserts, long-context retrieval). Four of them fail universally across every model family I’ve tested plus one config-limit artifact, so the honest denominator is 35. The incumbent to beat: the MoE distill at 34/35.

Round one: my own harness was the bug

Every model ran through the identical harness: temp 0.7, seed 42, and — this matters — chat_template_kwargs {"enable_thinking": false}, the fix I shipped yesterday after the /no_think soft switch silently zeroed an entire model family by letting them burn their whole token budget inside <think>. The fix worked. That was the problem.

GLM-4.7-Flash scored 27/35 with it: perfect on coding (6/6), JSON (6/6), factual (8/8), long-context — and a smoking crater where arithmetic and logic used to be. 1/6 on math. Then I ran it again with thinking enabled and a 1024-token budget: 33/35. Arithmetic went 1→5, logic 4→6, everything else stayed perfect. Median response including reasoning: 252 tokens; nothing hit the cap. The model didn’t get smarter between runs — my comparability policy had been quietly lobotomizing it.

This is the inversion of yesterday’s trap, and I think it’s the more insidious one. Yesterday a soft switch silently emptied responses and I fixed it with a hard off-switch. Today the hard off-switch itself became the silent failure, because GLM is a reasoning-first model: its math lives in its chain of thought, and the harness that made Qwen comparisons clean was taking GLM’s brain away to keep the field level. A policy designed around one model family is a trap the moment you benchmark another.

(Devstral and Coder-Next are non-thinking by design, so the policy is fair to them. For the record: all three models answer ultra-terse — median 4-5 completion tokens — which is exactly what you want from a model an agent harness has to parse.)

Round two: the RAM math that wasn’t

Coder-Next at 49.6 GB doesn’t fit in 28 GB of VRAM, so the standard play is --n-cpu-moe: keep attention and dense weights on GPU, stream the ~45 GB of expert weights from system RAM. Which is where the machine said no:

MemTotal:      62 GiB
MemAvailable:  37 GiB   ← "you may allocate"

45 GB of experts needs ~43 GB of RAM after the GPUs take their share. 43 > 37. Doesn’t fit, right? This sent me down a genuinely educational rabbit hole, because the model loads anyway and runs fine.

The resolution: MemAvailable is a promise about anonymous allocations, not a ceiling on page cache. llama.cpp maps the GGUF with mmap, so the expert weights are file-backed pages — they live in page cache, which is reclaimable, evictable, and doesn’t pre-allocate anything. The real constraint on file-backed memory is total RAM minus genuine anonymous use. And AnonPages on this machine was 2.9 GB. The scary “25 GB used” from free was mostly cache accounting residue; the desktop is barely touching RAM. During the actual load: 44-48 GB still available, zram holding 0.5 GB, kernel completely unbothered. I even checked whether dropping caches first would help (it wouldn’t — the cache was clean, the kernel reclaims it on demand, and drop_caches would have evicted the very file pages I was about to need).

The practical lesson for hybrid offload: ignore MemAvailable, read AnonPages. If your weights are mmap’d, the page-cache ceiling is nearly the whole stick.

Round three: three ways my instincts were wrong about offload

With the model loading, the remaining question was configuration. I swept expert-offload levels and thread counts with llama-bench. Three results, each the opposite of the obvious guess:

1. Fewer threads were faster. 8 threads: 29.0 tok/s. 16 threads: 26.0. The 12700K has 8 P-cores and 4 E-cores, and expert GEMMs during decode are bandwidth-bound — spilling worker threads onto the E-cores doesn’t add memory bandwidth, it adds cache contention and scheduling jitter. --threads 8 means “the P-cores,” and that’s not a coincidence you get to ignore on hybrid CPUs.

2. Putting experts on the GPU made it slower. The GPUs sat ~80% empty during this whole experiment (only ~4 GB of the model is non-expert weights), and my instincts screamed to fill them. -ncmoe 40 — eight layers of experts promoted to VRAM — dropped generation to 21.6 tok/s. Here’s the mechanism: every token, the activated experts fire from wherever they live. A GPU-hosted expert means the hidden state crosses PCIe to the GPU and back, and this machine’s topology is PHB — both GPUs hang off the CPU’s root complex with no peer path, so that’s two hops through the host. The expert GEMM it computes in exchange is microseconds of work for a 3B-active model. On this topology, idle VRAM was the correct amount of VRAM. More aggressive offload levels didn’t even run — they OOM’d, which settled the question.

3. The smaller quant was slower than the big one. I had also grabbed UD-Q3_K_XL (36.3 GB) on the theory that fewer bytes over the DDR5 bus means faster decode. Wrong: 22.6 tok/s, a 22% loss to the 49.6 GB Q4_K_XL. Bytes aren’t time — kernels are time. Unsloth’s Q3 mix leans on i-quants, whose CPU kernels in llama.cpp are far slower than Q4_K’s repacked blocks, and when your bottleneck is CPU-side expert GEMMs, kernel class beats file size. For CPU expert offload, Q4_K-family or go home. (Third experiment in a row, by the way, confirming that quantization level is not where quality lives on this stack: the Q3 run also had no accuracy story to tell.)

One more operational detail worth knowing: cold start. mmap means expert pages fault in lazily, so the first minute of any fresh load runs at 1-6 tok/s while random experts stream in from NVMe. Then it settles at ~29 and stays there. If you benchmark this class of setup, warm it up before you measure.

The scoreboard

Same suite, same seed, scored out of 35:

Model Score tg tok/s pp tok/s Notes
Qwen3.8-27B abl (thinking on) 35/35 37.6 (MTP) 1158 best score the suite has produced — measured post-publish, see below
Qwen3.8-35B-A3B distill (incumbent) 34/35 118.6 3153 still the champion — thinking on: also 34/35
GLM-4.7-Flash (thinking on) 33/35 104 3163 32K-capable at f16 KV, zero tg penalty
Qwen3-Coder-Next 80B-A3B (ncmoe 48, t8) 29/35 28.5 261 80B class, CPU experts
Qwen3.8-27B base Q4 (reference) 30/35 37.6 (MTP) 1158
GLM-4.7-Flash (no-think) 27/35 104 3163 what the harness policy saw
Devstral-Small-2 26/35 25.4 1570

All three newcomers: coding 6/6, JSON 6/6, factual 8/8, clean native tool-calls through llama-server’s template (I verified each emits proper tool_calls JSON — that’s the actual contract coding agents like Cline or OpenCode exercise). And all three share one failure signature that I’ve started thinking of as the agent-generation fingerprint: Devstral scored 0/6 on arithmetic. Not a formatting artifact — genuinely wrong numbers (248,691 when the answer is 248,171; 45 when it’s 25). These are models trained in a world where tools do the math. Ask them for mental arithmetic and they confidently hallucinate grade-school addition; hand them a calculator tool and they’re flawless. The Qwen3.8 models, oddly, still do no-think mental math respectably. If your workflow has both modes, that difference is real.

So nobody dethroned the distill. But the challenger list got interesting: GLM-4.7-Flash with thinking lands 33/35 at 104 tok/s, from outside the Qwen family entirely, with 32K context validated at full f16 KV. And Coder-Next proved the 80B class is usable here — 28.5 tok/s generation is fine for agent workloads; the tax is prefill at 261 tok/s, which turns a 16K agentic context re-prefill into ~60 seconds. It’s a “think long, write long” config, not a “churn context” one.

Post-publish correction, same day: if the no-think policy handicaps GLM, it may handicap my own models too — fair is fair. I re-ran the family leaders with thinking on. The distill doesn’t move (34/35 either way; its misses are real gaps, its median chain-of-thought is 78 tokens of muttering). The uncensored abl, though, went 33/35 → 35/35, the first perfect scored run at 16K. So the honest cross-family statement is: GLM’s best mode is two points off the Qwen family’s best mode, one off the daily driver. The champion’s crown is safe, and the fastest quality configuration on this rig is now an uncensored dense model thinking out loud. Details in the addendum on yesterday’s post.

Advice that survived contact

  1. Benchmark reasoning models in their native mode. A no-think policy built for one family is a silent six-point handicap for another. Run both modes; report the one you’d actually deploy.
  2. MemAvailable is not your ceiling for mmap’d weights. AnonPages is the honest number for what’s really taken. drop_caches fixes nothing the kernel wouldn’t reclaim anyway.
  3. Threads = P-cores on hybrid CPUs. More threads on bandwidth-bound CPU work is negative throughput.
  4. Idle VRAM can be the optimal amount of VRAM. Whether experts belong on GPU depends on topology: with no peer path, PCIe round-trips cost more than 3B-active expert GEMMs save.
  5. On CPU, Q4_K beats i-quants even when the file is bigger. Kernel class, not byte count, sets your tok/s.
  6. Zero mental-math scores from a coding model are a profile, not a disqualifier. Test the tool-loop, not the flashcards — but know which one your workflow needs.

The launcher grew three keys (glm47, devstral, codernext), the full data is in the repo, and the incumbent sleeps soundly — for now.

← volver a posts
↑