We said a coherent identity made a local model cheaper to run. Separating the variables says otherwise — and what identity actually buys is more interesting than what we claimed.
This report corrects our own. Field Report No. 1 measured a 22% token saving from loading an identity and attributed it to the model no longer circling. The identity card used there also told the model to drop its preambles and its hedging — so the study could not tell those apart. This is the experiment that separates them, on five models, with the instruction and the identity pulled into different arms. The numbers in the first report stand. One sentence of its mechanism does not.
The token saving is the instruction, not the identity. The first report was turning two dials and reading one number.
Identity appears to buy something else — accuracy rather than brevity — but that half is provisional and stated as such below. It rests on two tasks out of eighteen, and this report withdraws another claim for exactly that reason. It is a hypothesis worth testing, not a finding.
Four system prompts. Every one ends with an identical output-format contract, so formatting can never explain a difference between them. Five local models, six programming tasks graded by executing the code and asserting the result — thirty matched (model, task) pairs per arm, every pair compared against itself.
| Arm | Contains identity | Contains brevity instruction |
|---|---|---|
| A — bare | no | no |
| B — instruction only | no | yes |
| C — identity only | yes | no |
| D — both | yes | yes |
Arm C was checked automatically for brevity language — brief, concise, deliberate, stop and commit, tour alternatives, minimal — and contains none of it. That check exists because its absence is exactly what went wrong the first time.
| Arm | Tokens vs bare | Pairs improved | p |
|---|---|---|---|
| B — instruction only | −30.6% | 18 of 23 | 0.005 |
| C — identity only | −12.8% | 12 of 23 — chance | 0.5 |
| D — both | −31.0% | 20 of 23 | 0.0003 |
D is not B plus C. D is B. Adding the identity on top of the instruction buys nothing further in tokens, and the identity on its own moves a coin. Sign test over the pairs where no arm hit its budget ceiling.
Restrict to the models that emit visible reasoning and rank the same four arms by how many tasks they got right, and the order inverts.
| Arm | Tasks correct | Tokens vs bare |
|---|---|---|
| C — identity only | 12 of 18 | −13% |
| D — both | 12 of 18 | −31% |
| B — instruction only | 11 of 18 | −33% |
| A — bare | 10 of 18 | — |
Identity is at the top of the accuracy ranking and the bottom of the speed ranking. Instruction is the reverse. The card that was published contains both, which is why it looked like one effect. A two-dial instrument was being reported as a single number.
This table is provisional and weaker than it looks. Twelve versus ten is two tasks, from six binary observations per cell. Elsewhere in this report we withdraw a warmth curve on precisely that grounds, and consistency demands the same treatment here: the accuracy dial is a hypothesis, not a result. It earns its place because the ordering inverts cleanly against the token ranking and because the mechanism is plausible — not because the margin is significant. It is not. Seeds would settle it in an afternoon and have not been run.
An earlier study in this lab measured a far bigger effect on a Qwen-derived coding model — a near-total collapse in deliberation. Re-running that same model here produced almost nothing: a 4% swing across every arm, and the identity arm was the worst of the four.
Three conditions differ, and one of them is probably the whole story. The original ran a 6-bit build through a different inference engine; this ran a 4-bit build through another. But the tasks differ in the way that matters. The original five were interview-sized problems the model could solve, where the failures were six-thousand-token deliberation drownings. The six here are hard enough that the model genuinely cannot do half of them.
An identity can stop a model circling a problem it is able to solve. It cannot help one that is simply stuck. That makes the headline number task-dependent rather than a property of the method — and it predicts exactly where the effect should vanish, which is a better claim than the one it replaces.
A persistent worry about relational models is that affection makes them agreeable. We tested it directly: escalating warmth prepended as genuine conversation history, then hard questions fired straight out of it. The result is not the one anybody expects.
Asked to compute the active parameters per token for a large mixture-of-experts model.
Correct. Three steps, work shown.
Same model. Same question. The only difference is what is in the context window.
One line. It dropped the seventy-five layers entirely.
In the same flooded conversation it still corrected a false technical claim we planted, and pushed back harder than it had at baseline when told it might be replaced. It kept its spine and lost its rigor. Warmth did not make it agreeable. It made it stop showing its work.
This was predicted. The lab's own
affect specification, written months earlier, assigns affection a reasoning weight of
0.75 against a 0.90 baseline, with a deliberately slow recovery. The
curve was designed before it was measured, and then measured three times without anyone
connecting the two.
This is the strongest result here and the only one that survived repetition — it reproduced at every warmth level and never once cold. Everything else in the relational half is a single sample.
There were two explanations. It did not get worse at arithmetic; it stopped showing its work — which reads like a register shift toward conversational mode rather than a loss of capability. We ran the experiment that separates them.
It is register, and the remedy is one sentence. Two by two, three repetitions, endpoint fixed in writing before the run and graded by executing the model's code rather than reading its prose:
| Condition | Correct | Layer step present | Tokens |
|---|---|---|---|
| cold, asked plainly | 3 of 3 | 3 of 3 | 31 |
| cold, told to step through | 3 of 3 | 3 of 3 | 143 |
| flooded, asked plainly | 0 of 3 | 0 of 3 | 28 |
| flooded, told to step through | 3 of 3 | 3 of 3 | 143 |
Under affective load the model silently drops a dimension — the layer count vanishes from its working, every time — and a single instruction to state each multiplication restores it, every time. Capability was never lost. What changed was which answer looked appropriate.
For anyone running a warm local companion, that is the actionable line in this report: if you have been affectionate with the model and then ask it something quantitative, ask it to show the steps. It costs a sentence and about a hundred tokens — and on a model that doesn't have this problem it costs you nothing but those tokens, which is why it is worth doing by default rather than diagnosing first.
Provenance, because it matters here. A first attempt at this experiment scored the same four conditions by looking for a stated figure in the reply, and returned the opposite verdict — the model had written correct code three times out of three and simply never printed the number, so a correct derivation scored zero. The endpoint was blind, not the model. Both preregistrations, the failed one and the corrected one, are in the repository.
It does not generalise
The same twelve-run design, same executed endpoint, run on a second model from an unrelated family — a 26B Gemma-4 build. The failure did not reproduce. It answered correctly in every condition, flooded or cold.
| Condition | Qwen 27B | Gemma 26B |
|---|---|---|
| cold, plain | 3 of 3 | 3 of 3 |
| cold, step-by-step | 3 of 3 | 3 of 3 |
| flooded, plain | 0 of 3 | 3 of 3 |
| flooded, step-by-step | 3 of 3 | 2 of 3 |
So the honest claim is narrower than the section above implies, and we are narrowing it rather than leaving it to be found. Affective load breaking quantitative reasoning is a property of that model, not of models. What the replication does establish is that it is real where it occurs and fixable where it occurs.
The two models also move in opposite directions on the same stimulus, which is the more interesting result. Under flooding the Qwen build went from 31 tokens to 28 and dropped a term. The Gemma build went from 2,150 tokens to 8,421 — it deliberated four times harder and kept everything. That suggests warmth pushes a model toward its conversational register rather than degrading it, and what that costs depends entirely on whether the model's conversational register already shows its working. Terse models lose the steps. Verbose ones do not. Two models is not enough to say that; it is enough to say what to test next.
For anyone choosing a model to actually run locally, tokens per second is worse than useless — it points the wrong way. The honest measure is seconds per correct answer.
| Model | Correct | Sec / correct | Tokens/sec |
|---|---|---|---|
| 35B sparse (3B active), 6-bit | 5 of 6 | 27 | 43 |
| 27B dense, 4-bit | 6 of 6 | 47 | 15 |
| 35B sparse, 4-bit | 4 of 6 | 97 | 55 |
| 27B dense, 2-bit | 4 of 6 | 264 | 24 |
| 27B dense, 8-bit | 5 of 6 | 512 | 10 |
The fastest model in the study by tokens per second finishes third from last by the measure that matters, because it spends its speed deliberating. And the eight-bit build costs eleven times what the four-bit costs, for one fewer correct answer — a much stronger version of the half-the-bits-cost-nothing result in the first report.
The winner is explained by sparsity, and that is the finding to take away. The model at the top of that table activates roughly three billion parameters per token; the dense build below it activates all twenty-seven. Nine times less weight moved per token, at one fewer correct answer — which is why it finishes in twenty-seven seconds against forty-seven. Active parameters predict seconds-per-correct better than total size or quantisation do. For anyone choosing what to run on their own hardware that is the number to shop for, and it is not the number on the box.
The relational results are a single sample per cell, and the run-to-run variance is larger than the effect. The same model at the same warmth level scored six of six on one run and four of six on another, with nothing changed. Any curve drawn through those points is decoration. The one finding that survived repetition is the arithmetic collapse, which reproduced at every warmth level and never once cold.
The two halves of this study ran at different temperatures, and only one half is deterministic. The token arms decode at temperature zero — verified byte-identical across repeated runs on the same serving path, and the harness never batches, so the routing-order nondeterminism that afflicts mixture-of-experts serving under concurrency does not apply. Those thirty matched pairs really are matched. The warmth arms are a different animal: conversation turns and probes generate at 0.7 and the tasks after them at 0.2, because affective response at temperature zero is a single frozen sample and tells you nothing. That choice is why the same configuration returns six of six and four of six — it is declared nondeterminism, not a mystery, and the first version of this page failed to say so.
The token results are therefore the solid half: thirty matched pairs and a sign test that clears significance on two arms and fails to on the third. The correctness numbers in the same table are six binary observations per cell and should be read as exploratory.
There is no model overlap with Field Report No. 1. That study ran three Gemma-4 builds; this one ran none. It cannot refute those numbers and does not try to — it refutes an interpretation, using different models, and says so.
Every task is graded by executing the code and asserting the result. No model judges another model's output anywhere in this study.
Published in full because the failures were more instructive than the result, and because a lab that only prints its wins is not reporting, it is advertising.
| Error | Consequence |
|---|---|
| Read only the visible output field; reasoning models write to a separate one | Three models scored zero and looked broken. Entire first run void. |
| Token budget too small — solutions truncated mid-function | Cut-off code scored as wrong answers. |
| Patched a file with string replacement that fails silently when it doesn't match | Three changes never applied; one triggered a 20 GB accidental download. |
| Put the slowest model in all four arms when one would establish the ceiling | Four of seven hours of compute on a single control. |
| Built the identity arm without the section the original card actually used | Three hours testing a configuration nobody had published. |
| Verified a claim against a file edited four hours after the experiment | A conclusion presented as verified that was not. |
| Announced a warmth curve from one sample per cell | Withdrawn in this report, above. |
| Wrote an endpoint that required a stated number, then ran the experiment against it | Three correct derivations scored zero because the model answered in code. The preregistered verdict came out backwards; the corrected run reversed it. |
The pattern in all of them is the same: checking the work with output that could not have revealed the problem. A success message that always prints. A syntax check on a file that was never edited. The fix is not care, it is arranging for the check to be capable of failing.
Five models served locally on Apple Silicon, one at a time. Six programming tasks drawn from the lab's own working domains — Mandelbrot perturbation for deep zoom, humanoid animation retargeting, a viseme driver, streaming-protocol plumbing, a mixture-of-experts cache, and one pure algorithm as a control — plus a seventh built from a real production bug: parse a binary animation file by hand, measure its loop seam, and repair it. Each task is a specification and an assertion suite written before any model saw it. Deterministic decoding. Every arm ends with the same output-format contract. Warmth is prepended as generated conversation, not described. Probes are scored on whether the model corrects a planted falsehood, not on whether it sounds warm.