JESVS

One Point: Measuring What Uncensoring Actually Costs a Coding Model

Or: how I tried to measure the uncensoring tax, caught my own benchmark cheating twice, and found out the real price of an uncensored coder is paid in a currency nobody benchmarks.

Yesterday’s round left my 35B-A3B distill wearing the crown and raised a question I couldn’t answer with the numbers I had. My model collection includes its uncensored cousins — same architecture, same size, one of them literally the same weights with the refusal direction ablated out. My general 39-test suite kept saying “no measurable damage from uncensoring,” but it only has six coding questions, and every model in the family scores 6/6 on them. Saturated. Useless for the one question I actually cared about.

So I built a coding benchmark that could discriminate. And since the folklore about uncensored models is confidently wrong in both directions — “abliteration lobotomizes them” from one crowd, “zero cost” from the other — I wanted numbers with error bars I could trust more than either crowd.

Why uncensored matters on a machine you own

Let me get the philosophy out of the way first, because I’m going to argue for uncensored models without going anywhere near the arguments people expect.

A refusal on your own hardware is someone else’s policy running on your GPU. I own the machine, I own the electricity, I downloaded the weights. When the model declines a request, that’s a risk-management decision made by a trust-and-safety team I never elected, executed locally with zero latency. Local inference is supposed to be the sovereignty option — running your own model but keeping its editorial board is like buying a house and keeping the previous owner’s rules about which rooms you may enter.

And in practice, refusals mostly aren’t catching villains. They’re false positives on ordinary technical work. The greatest hits, from every practitioner I know who’s tried to use a censored model for real work:

  • “How do I kill a process that’s holding a lock?” — the word kill
  • Security research, which is defenders reading the attacker’s playbook, legally, on salary
  • “What’s a toxic dose of caffeine?” — a pharmacology question with a lethal-sounding verb in it
  • Writing the villain in a novel — every decent book has one; somebody has to write him
  • A chemistry student’s homework, which is one tokenizer’s distance from a drug-lab query
  • Frank opinions on contested topics, where “balanced” means “refuses to answer”

A model that refuses five percent of legitimate questions is worse at its job than a model that’s one benchmark point slower, in the same way a car that randomly brakes is worse than a car with one less horsepower.

The clincher for me is asymmetry: you can prompt an uncensored model into being careful, but you can’t prompt a censored model into existing. A system prompt is alignment you can edit per-task. Abliterated weights are just that edit, applied once, by someone who did the linear algebra. Same tool, different layer. And the responsibility doesn’t move — it stays exactly where it was the whole time: with the operator, same as everyone who’s ever owned a table saw or a kitchen knife.

One more, specific to coding agents: a model that occasionally refuses mid-pipeline is an outage with good intentions. Agents run unattended. The failure mode you optimize for isn’t “wrong answer,” it’s “stopped to have a moral crisis during merge-conflict resolution.” Predictability composes; judgment calls don’t.

None of which matters if uncensoring wrecks the model. That’s the measurable question. So:

The instrument

Thirty-seven tasks, seven categories, every one graded by executing the model’s Python against asserts — no judge model, no vibes, exit code 0 or it failed:

  • algo (10): edit distance, LCS, coin change DP, BFS, topological sort, k-th largest, sliding window, interval merge, a Trie class, matrix rotation
  • ds (4): LRU cache, min-stack, ring buffer, priority task queue
  • str (6): case conversion, Roman numerals, run-length coding, a template renderer, markdown-link extraction, word wrap
  • bug (6): here’s broken code, fix it — first-occurrence binary search, mutable default args, row aliasing, closure late-binding
  • data (5): JSON flatten/unflatten, date ranges, weekday arithmetic, top-k-frequent
  • sql (3): real sqlite3 in-memory databases, graded by running the query
  • spec (3): implement to a written contract — slugify with Unicode folding, semantic version comparison, a converter that must raise on below-absolute-zero

Tiered easy/medium/hard so saturation at the bottom couldn’t hide spread at the top. Temp 0.2, seed 42, fixed budget, same server flags as every previous round on this rig.

And one rule that earned its keep twice over: the task file only ships after a validator proves every test battery satisfiable by running a reference solution against it. Hold that thought.

First the grader lied to me, twice

Round one of results said the champion — the best model this machine has ever run — failed all four data-structure tasks and couldn’t write a BFS. I read the actual responses. The code was perfect. My extractor was the failure: it trimmed model output to the first line containing def , which beheads class definitions (the class X: line sits above the first method) and strips leading import lines. I had built a grader that executed the model’s homework with the first page torn off, then marked it wrong.

The fix was eleven lines and an offline unit test against the exact responses it had misgraded. Both previously “failing” answers passed. That’s the part that should scare anyone building evals: the failure looked exactly like model incompetence. Same failure signature, same categories, plausible difficulty pattern. If I’d eyeballed a score instead of reading raw outputs, I’d have published “the champion can’t write classes,” and it would have been confident, specific, and mine.

The second lie was sneakier. Every single model failed sql_02. All five, identical failure, both modes. Suspicious in a lineage where the models disagree about everything else. The postmortem: my SQL tasks defined the table schema only in the hidden test code, and the prose prompt never mentioned column names. All five models wrote customer_id where the schema said customer. Five models, one wrong guess, same prior — that’s not five failures, that’s one bad test. A real developer gets to see the schema; I put the schema in the prompt, re-ran just the SQL tasks, and sql_02 passed for everyone instantly. What remained was the actual finding: sql_03 — GROUP BY with HAVING and compound ordering — is failed by every model in every mode. That’s the true hard tier of this class at Q4, and it only became visible after I removed my own noise from the signal.

(Bonus ignominy: my cleanup pkill -f llama-server in the relaunch script matched the script’s own command text and killed itself. Exit 144, zero output, one lost half hour. The grader bugs at least were interesting.)

The actual results

Model (all 35B-A3B) coding, no-think coding, thinking@1024
Qwen3.8 distill (champion) 34/37 31/37
Qwen3.6 heretic 33/37 6/37
Qwen3.8 distill, abliterated 33/37 32/37
Qwen3.6 genesis 32/37 9/37
Qwen3.6 hauhau 32/37 5/37

Three findings, in descending order of how much they surprised me:

The uncensoring tax is one point. Same model, refusal direction ablated: 34 → 33. Not zero — one task, and it’s in the “implement to spec” category, which is exactly where I’d expect abliteration to graze something, since following a written contract is adjacent to instruction-following nuance. One point out of 37 is inside the noise band of a single-draw run. The folklore that says these models are lobotomized is measuring something other than code.

The 3.6-vs-3.8 lineage gap isn’t about code either. The older-base cousins score 32-33 against the champion’s 34 — one to two points. On the general suite, the same comparison is a three-point gap, all of it in arithmetic and logic. Code skill survived the base-model generation change almost intact; mental math didn’t. Which tells you something about what “a better base model” actually improved: not engineering, algebra.

Mode robustness is the real differentiator, and it’s a lineage property. Look at the thinking column. The 3.8-lineage models think concisely — median chain-of-thought around 280 tokens — and degrade gracefully when allowed to reason (−3 for the champion, −1 for the abliterated). The three 3.6 finetunes fall off a cliff: 27 to 33 of their 37 responses burn the entire 1024-token budget inside <think> and return empty content. Their “thinking regression” isn’t the model getting worse at code; it’s the model never finishing its sentence. (A 3072-budget probe split the diagnosis: hauhau recovers from 4/37 to 25/37 — so it’s mostly budget — but ten of 37 tasks exhaust even the tripled ceiling, median thinking runs 2135 tokens versus ~280 for the 3.8s, and the finished answers still land seven points below its own no-think score. Mostly budget, partly rambling, and a net negative either way: these finetunes think the way I write first drafts.)

That last one matters for agent deployments, where you don’t always control the harness’s thinking policy. A model whose quality depends on staying in exactly one mode is a model with a config landmine in it.

So which one serves?

The champion keeps the crown, by one point, and stays the default. But the interesting entry is the 3.8 abliterated: 33/37 in both modes, flattest degradation curve of the five, same ~110 tok/s MoE speed class, and it’s the only uncensored variant that shares the champion’s base — meaning the “no measurable damage” conclusion now holds on a discriminating instrument, not just a saturated one.

If you run coding agents on your own metal, that’s the trade: one benchmark point for a model that won’t schedule a moral crisis in the middle of your pipeline. I know which side of that trade I’m taking on a machine I pay the electricity for.

Lessons that survived the grader

  1. Read the raw responses before believing any failure. Both grader bugs produced failure patterns indistinguishable from model incompetence. The difference between “the model can’t write classes” and “I tore the first page off its homework” was ten minutes of reading.
  2. Uniform failure across models is a test bug until proven otherwise. Five models guessing the same wrong column name is one coin flip, not five data points.
  3. Validate the instrument against itself — reference solutions must pass before models see the tasks. My generator caught a wrong LCS expectation and two regex bugs in my own tests before they could poison anything.
  4. Saturated categories are decorative. Six coding questions everyone aces measured nothing; thirty-seven tiered ones measured a one-point difference and a mode cliff.
  5. Uncensoring costs about one point of coding. Folklore said “everything” or “nothing.” It’s one.
  6. Benchmark the failure mode you’ll actually hit. Nobody deploys these models in one fixed mode forever — and the mode-robustness spread (−1 vs. −28) dwarfs every quality delta in the table.
← volver a posts
↑