JESVS

Squeezing 2.7x More Tokens Per Second Out of Two Mismatched GPUs

Or: how I spent a weekend benchmarking local LLM inference, caught my optimization guide lying to me four times, and learned that the boring measurements beat the exciting flags.

My desktop is an awkward machine for local LLMs. An i7-12700K, 64 GB of RAM, and two GPUs that have no business sharing a case: an RTX 5060 Ti with 16 GB (Blackwell, PCIe 5.0, fast) and an RTX 3060 with 12 GB (Ampere, stuck negotiating PCIe 3.0 x4, slow). 28 GB of VRAM total, connected through the CPU’s root complex with no peer link between them. The goal was simple: run Qwen3.8-27B and its MoE cousins entirely on GPU, with useful context, as fast as this hardware allows.

This is the story of what worked, what didn’t, and how I know the difference.

First, the guide was wrong

I started with a methodology document that explained all the tunables: split modes, tensor splits, flash attention, KV cache types, P2P, CUDA launch queues. It looked thorough. Then I ran a 30-second smoke test before the big benchmark matrix and it caught four silent failures:

  • -fa on silently parses as flash_attn: false in llama-bench. It accepts the string without complaint and disables flash attention. You need -fa 1. The server accepts on just fine, which is why nobody notices.
  • -ts 4,3 doesn’t mean what you think in llama-bench. Comma is the separator for repeated test values, so it runs two single-GPU tests. A two-GPU split is -ts 4/3. I only caught this because I read the parsed JSONL instead of trusting the exit code.
  • -ngl all is invalid in llama-bench (fine in the server). Use a number.
  • Tensor split mode and CUDA_SCALE_LAUNCH_QUEUES don’t exist in my build at all. The guide recommended benchmarking both.

The lesson that shaped everything after: read back what the tool actually recorded (-o jsonl and inspect the fields) instead of trusting that your flags were honored. This bit me twice more later — a -ot tensor-override regex that double-escaped its backslashes and silently matched nothing, producing four “results” that were all just the default config wearing a costume.

There was also a fun detour where the models refused to load at all. The GGUFs use the qwen35 architecture; the installed binaries were five months older than the source tree they sat in and predated that architecture’s support. An unknown architecture that manifests as a one-second load failure sends you down a very different debugging path than a crash.

Asymmetric GPUs want asymmetric splits

The obvious split for 16 GB + 12 GB is proportional: -ts 4/3, or 57/43. That balances capacity. It doesn’t balance compute, because the 5060 Ti is much faster than the 3060.

Measured prompt processing on the dense 27B at Q4_K_M:

4/3 → 1080 tok/s
3/2 → 1133 tok/s
5/3 → 1144 tok/s   (62.5/37.5)

Generation stayed flat around 21 tok/s regardless. So the win is in prefill, and it’s real: 6% for shifting work toward the faster card. Pushing further (68%, 71% via explicit layer overrides) made things slightly worse and would have blown past VRAM at 16K context anyway. 5/3 is the plateau.

Why doesn’t generation speed care? Roofline math. With 5/3, the 5060 Ti holds ~10.7 GB of weights at 448 GB/s and the 3060 holds ~6.4 GB at 360 GB/s. Add those serially and you get 41.7 ms per token, or ~24 tok/s theoretical. Measured: 21.2, which is 88% of the ceiling. Decode is memory-bandwidth-bound and the two GPUs execute their layer stacks in series — the hidden state that crosses the PCIe link per token is tens of kilobytes, microseconds of transfer. The scary x4 link on the 3060 is simply not the bottleneck for generation, which the P2P experiment confirmed: enabling peer-to-peer access changed nothing (≤0.3%), as you’d expect on a topology with no peer path.

I did test the row split mode, which parallelizes each op across GPUs instead of pipelining layers. Prompt processing collapsed to 228 tok/s — five times worse. Every op pays an allreduce over that x4 link. Layer mode or nothing on this machine.

The MoE plot twist

Then I tested a 35B-A3B MoE distill, and the picture changed completely:

dense 27B Q4 MoE 35B-A3B Q4
generation 21.2 tok/s 111.8 tok/s
prefill (2048) 1144 tok/s 3026 tok/s

Same VRAM ballpark, five times the decode speed. With only ~2.6B parameters active per token, the model isn’t fighting the memory wall the same way. The bandwidth roofline for it is somewhere north of 250 tok/s, which means the MoE is overhead-bound — kernel launches, expert routing, cross-GPU sync — not bandwidth-bound. That distinction predicted everything that came after: tricks that help bandwidth-bound models do nothing here, but concurrency is nearly free, because idle overhead can absorb more sequences.

The concurrency numbers are the sleeper hit of the whole project. llama-batched-bench with parallel sequences on the MoE:

1 stream  → 117.8 tok/s aggregate
2 streams → 186.9
4 streams → 239.4
8 streams → 284.4

2.4× aggregate throughput at 8 streams, and llama-server’s default n_parallel=4 already collects it if your client actually sends concurrent requests. If you run agents with parallel tool calls on a rig like mine, this is the free lunch.

Speculative decoding: measure the fine grid

The Qwen3.8 release ships an MTP head — a small draft model (~1.3 GB at Q4_0) that guesses the next couple of tokens, which the big model verifies in one batched pass. My first sweep tested draft lengths 3, 5, and 8, found 3 best at 27.1 tok/s (+28%), and shipped it.

An external review pushed me to sweep with step 1 instead of step 2. n_max=2 gives 31.3 tok/s. The coarse grid had skipped straight over the optimum. The curve falls off a cliff after that: 31.3, 27.9, 25.3, 21.7, 18.9, 16.8 for drafts of 2 through 7.

Why do longer drafts lose so badly on this architecture? I measured the verify cost directly — forward passes for batches of 1 through 8 tokens:

p1: 47.9 ms    p2: 49.3 ms    p4: 57.4 ms    p8: 72.7 ms

The second token in a verify batch costs 1.4 ms extra. The eighth costs about 3.5 ms each. Meanwhile acceptance is 0.977 on code-edit workloads (the server logs it per request — draft acceptance = 0.977 (338 accepted / 346 generated)), so a 2-token draft captures most of the achievable value at nearly zero marginal cost. My running hypothesis for the steep marginal cost is the hybrid architecture itself: this is a Gated DeltaNet design where most layers carry recurrent state, and rejecting a draft means rewinding that state.

Two more spec-decode findings worth knowing. First, acceptance is workload-dependent: on prompts that ask the model to rewrite a file (output largely copies the context), the dense model hit 46.9 tok/s with MTP versus 37.6 on free-form generation. If you benchmark with generic prose, you’ll understate your coding throughput by a quarter. Second, n-gram speculation — drafting by matching prefixes in your own context, no draft model needed — is a no-op for free-form generation but gave +59% on those same edit workloads. On the MoE, both MTP and n-gram made things worse: verification activates enough extra experts that high acceptance doesn’t pay.

The cheapest win: rebuilding llama.cpp

My build was five months old, which in llama.cpp years is geological. I built current master in a separate worktree (so the known-good binaries stayed untouched), explicitly targeting both GPU architectures:

cmake -DGGML_CUDA=ON -DCMAKE_CUDA_ARCHITECTURES="86;120" ...

Before rebuilding I checked whether the old build was even leaving performance on the table — cuobjdump showed native SASS for both sm_86 and sm_120, so no PTX JIT tax. The gains came from five months of real code improvements, not architecture flags:

May build master
dense + MTP n2 31.3 37.6
MoE generation 111.8 118.6
MoE prefill (4096) 3581 3812

Now the important part: how do you know a new build didn’t change your model’s behavior? I ran the same quality suite at temperature 0 on both builds. Identical scores, identical failure sets, three answers differing in wording out of twenty-nine. That’s the regression test I trust; temp 0.7 runs on both builds differed by three tests purely from sampling-path divergence. Which leads to a methodology confession.

39 tests at temp 0.7 is a vibes detector, not a ruler

Every quality number in this project comes from a 39-question suite — arithmetic, logic, facts, JSON schema checks, executable Python against asserts, needle-in-haystack retrieval at 8K/16K/32K — at fixed seed and temp 0.7. It’s excellent for catching catastrophic problems. It cannot resolve differences of two or three questions, and I stopped pretending it could. When the Q5 quant of the MoE scored 30/39 against Q4’s 34/39, the honest read was “no evidence of improvement, possibly noise” — and the speed cost (−10%) and lost vision headroom made the decision regardless.

That’s now three separate experiments telling the same story: quantization level is not where quality comes from on this stack. Dense Q4→Q5: 30→30. MoE Q4→Q5: 34→30. Model choice moved scores far more (base 30 → best variant 33 → MoE distill 35). Spend your time picking models, not bits.

Two methodology traps I hit that others should know about:

The /no_think trap. Qwen models have a soft switch to disable chain-of-thought. Qwen3.8 templates honor it; Qwen3.6 templates ignore it. If your harness appends /no_think, sets a modest max_tokens, and checks message.content — every response comes back empty with exactly max_tokens generated, because the model spent the whole budget inside <think>. My first uncensored-MoE scores were 10-14/39 and looked like abliteration had lobotomized the models. The fix is chat_template_kwargs: {"enable_thinking": false} in the request. Current master made this worse by ignoring /no_think even where it used to work.

Fixed seeds don’t make configs comparable. Any numeric change — different quant, different build, different split — shifts floating-point results, and at temp 0.7 those small shifts compound into entirely different token streams. Same seed, different trajectories. For build-vs-build and quant-vs-quant questions, deterministic (temp 0) A/B runs are the honest instrument; sampled runs are for “does it feel right” questions.

The endpoint: one GPU, no hop

After all the split tuning, the dense model was still paying for pipeline serialization — two GPUs executing in series, each waiting for the other. The escape hatch: shrink the quant until the model fits on the fast GPU alone. Unsloth’s UD-IQ4_XS (14.3 GB) squeezes onto the 5060 Ti with 8K of q8 KV cache:

llama-server -m Qwen3.8-27B-UD-IQ4_XS.gguf -ngl 99 -sm none -mg 0 \
  -c 8192 -ctk q8_0 -ctv q8_0 -fa on \
  --spec-type draft-mtp -md mtp-draft-Q4_0.gguf \
  --spec-draft-ngl all --spec-draft-n-max 2 --device-draft CUDA1

The draft model lives on the 3060, which otherwise sits idle with 11 GB free. Plain generation: 28.0 tok/s, which is 96% of that GPU’s roofline — no serialization, no hop. With MTP: 51.4 tok/s free-form, 58.1 on code edits, and the quality suite puts the smaller quant above the Q4_K_M baseline (31 vs 30, inside the noise band, but certainly not worse). The trade is honest: 8K context instead of 32K, no VRAM for the vision projector. It’s the sprint config; the split is the marathon config.

Where it all landed

dense:  21.1 → 21.2 (split) → 31.3 (MTP n2) → 37.6 (master) → 51-58 (single-GPU)
MoE:    111.8 → 118.6 (master), 239-284 aggregate with parallel streams

And an equally valuable list of things that measurably do nothing on this hardware: P2P, row split mode, splits past 5/3, draft placement, draft quant, p-min tuning, disabling CUDA graphs, forcing MMQ (no change — the heuristic already picks it), forcing cuBLAS (down 19% on dense prefill, down 60% on MoE prefill; don’t), n-gram speculation on free-form text, and every KV quantization below q8 (q8 itself costs nothing — zero speed penalty, zero measurable quality change, and a 27,710-token needle retrieved from the middle of a 32K context to prove it).

If I had to compress the whole thing into advice:

  1. Smoke-test one tiny run and read the recorded JSONL before any big matrix. Silent flag misparse is the default failure mode, not the exception.
  2. Roofline first. Two divisions tell you whether you’re bandwidth-bound (buy nothing, split toward the faster card) or overhead-bound (chase concurrency and launch costs).
  3. Sweep with step 1. The difference between 27 and 31 tok/s was hiding between my grid points.
  4. Temp-0 A/B for build and quant changes; temp-0.7 only for vibes.
  5. Upstream moves fast. A five-month-old binary left 20% of my spec-decode speed on the table.
  6. Write down the negative results. They’re what kept three rounds of external AI review productive — the reviewers couldn’t waste my time re-proposing doors I’d already measured shut.

The machine that started at 21 tok/s now serves a 35B MoE at 118 (284 in parallel), a vision-capable dense model at 37.6 with 32K context, and a 58 tok/s sprint config — all on the same mismatched pair of GPUs I was told couldn’t be balanced.

← volver a posts
↑