What a layer transplant restores
On this page
Can parts of an older language model recover answers that a newer version no longer gives? I tried adding selected blocks of layers from an older model to its newer relative, producing the model I call Jumbo. I chose blocks whose weights had changed most between versions, then tested whether the transplant helped on obscure factual questions.
On 2,000 questions, Jumbo answered 13.5% correctly, compared with 7.2% for the newer model before modification. But wrong answers also rose from 102 to 796: the original often declined to answer, while Jumbo attempted many more responses. The additional correct answers came with a substantial cost in incorrect ones.
It might seem that answering more questions correctly makes Jumbo the better model. However, that depends on how we treat the questions it answers incorrectly. The benchmark gives the same score to a wrong answer and a refusal to answer, while someone relying on an assistant may have good reason to distinguish them. Jumbo’s higher score therefore supports a narrower conclusion than saying it is the more dependable assistant. I built and evaluated the models on my own machine.
Compare the models (opens in a new tab) or see the released model.
Technical methods and evaluation
Jump to the layer selection, the measurement, or what the model is for.
The transplant
To choose layers systematically, I compared corresponding weights in the older and newer models. The models share an architecture, and their weights are sufficiently aligned to compare corresponding tensors. The global median cosine alignment is 0.917. Dividing the 64 layers into sixteen blocks of four identifies three blocks with unusually large changes, at layers 12 to 15, 20 to 23, and 40 to 43. Jumbo inserts each donor block immediately before its counterpart in the newer model, giving 76 layers and about 31.9 billion parameters. It remains compatible with the newer model’s multi-token-prediction drafter at 82 to 89 percent acceptance. The models are Qwen3.6-27B (donor), Qwen3.8-27B (parent), and Qwen3.8-32B-Jumbo (transplant). My original model card claimed restored factual breadth while retaining code and agentic capabilities; the evaluation here tests the factual claim.
I chose these blocks on the hypothesis that the largest changes might include information or behaviors worth recovering from the older model. A large change in weights is a reason to investigate a block, but it does not identify what information was changed or whether that change was beneficial. The factual evaluation is therefore needed to test whether this particular selection produces more correct answers.
The measurement
I tested the original, the layer donor, and the modified model under the same conditions so that their answers could be compared directly. The questions come from PopQA, a public benchmark of about 14,000 short factual questions derived from Wikidata. Each question includes a popularity score for its subject. I sampled 2,000 questions from the least-popular third using a fixed seed. The donor, parent, and Jumbo were all evaluated at 8-bit quantization with the same group size, greedy decoding, and thinking mode off. The prompt asked for a fact in ten words or fewer and allowed the response “I don’t know.”
Scoring follows the benchmark’s alias-containment rule. An answer is counted as correct if its normalized text contains an accepted alias. This is a permissive rule. For example, “Gaelic football” can be accepted when an alias for the reference answer is “football,” even when the intended sport is association football. The comparison page (opens in a new tab) shows the responses alongside the references so that these cases can be inspected.
| Model (8-bit, greedy) | Accuracy | Answered at all | Right when it answered | Wrong answers |
|---|---|---|---|---|
| Qwen3.6-27B, donor | 18.7% | 69.8% | 26.8% | 1,022 |
| Qwen3.8-27B, parent | 7.2% | 12.3% | 58.5% | 102 |
| Qwen3.8-32B-Jumbo | 13.5% | 53.3% | 25.3% | 796 |
Jumbo scores between the donor and parent on this sample. Its accuracy gain over the parent is 6.3 percentage points (95% bootstrap interval 5.2 to 7.4), and it recovered 109 of the 236 questions the older model answered correctly and the newer one missed. Compared with the parent, it answers 130 additional questions correctly and loses 4 correct answers. Both models answer 140 questions correctly and miss 1,726. These results address factual accuracy under this evaluation; they do not test the model card’s separate claim about retaining code and agentic capabilities.
The parent declined 1,754 of the 2,000 questions. When it did answer, it was correct 58.5 percent of the time. The donor answered 70 percent of the questions and gave 1,022 wrong answers. Jumbo answered 53 percent and gave 796 wrong answers. In particular, there were 820 questions on which the parent abstained and Jumbo answered. Jumbo was right on 124 and wrong on 696 of those questions.
The additional attempts complicate the claim that the transplant restores knowledge. If the parent declines to answer a question, that could mean the relevant fact is absent from its weights, but it could also mean the prompt fails to elicit a fact it could otherwise supply. Both possibilities are compatible with a refusal, so this test cannot distinguish them or identify which training changes caused the parent’s abstention. What it does show is a change in response behavior. Jumbo’s precision when answering, 25.3 percent, is within two percentage points of the donor’s; it answers more readily and comes closer to the older model’s rate of success when it does.
Where the recovery happened
PopQA labels each question with the Wikidata relation it asks about, which makes it possible to see what kind of fact came back.
| Relation | n | Donor | Parent | Jumbo |
|---|---|---|---|---|
| country | 220 | 69.5% | 48.2% | 59.5% |
| sport | 156 | 57.7% | 10.3% | 37.2% |
| capital | 34 | 58.8% | 20.6% | 44.1% |
| occupation | 127 | 22.8% | 0.8% | 8.7% |
| place of birth | 171 | 14.6% | 1.2% | 6.4% |
| genre | 192 | 9.9% | 0.5% | 9.4% |
| father | 70 | 14.3% | 2.9% | 14.3% |
| author | 334 | 4.2% | 1.2% | 1.5% |
| director | 238 | 0.0% | 0.0% | 0.4% |
| screenwriter | 159 | 1.3% | 0.0% | 0.0% |
| producer | 131 | 0.8% | 0.0% | 0.0% |
The effect differs substantially by relation. Questions about country, sport, and capital improve more than questions about the author of an obscure book or the director of an obscure film. One possible explanation is that attempting an answer is more likely to succeed when there are fewer plausible answers. Country and sport questions may offer fewer candidates than questions about obscure authors. However, the categories differ in more than the number of possible answers, and this comparison does not control for those differences. The pattern gives a reason to investigate that explanation, without resolving the earlier distinction between a missing fact and one the model does not provide under this prompt.
The regressions matter as well. In three cases, Jumbo replaced a correct answer with a plausible but incorrect one. It gave Norway for a river in Finland, Medyka for the capital of a Slovak district, and Richard Rorty as the author of a book by Robert Anton Wilson. There are only four regressions in the bank, and the comparison page includes all of them.
What the model is for, then
Whether Jumbo is preferable to the parent depends on what the system is meant to do. Under a rule that awards a point for a correct answer and treats an incorrect answer and an abstention equally, Jumbo performs better on this sample. If incorrect answers carry a cost, that comparison changes. In a retrieval-based assistant, for example, declining to answer when the passages do not support a response can be desirable. This benchmark does not test that setting directly, but the increased rate of wrong answers is a reason I would want further evaluation before using Jumbo there.
The model card needs to describe this tradeoff alongside the accuracy gain. The results support a narrower claim than simply saying that Jumbo knows more.
Trying it yourself
The comparison page (opens in a new tab) runs a prompt through the parent and Jumbo at the same quantization with greedy decoding. It streams both responses and marks where their text first differs. The page also shows examples from the evaluation bank, including a random dozen recoveries and all four regressions, with the donor’s response and the reference answer. You can run any of those questions live. The live demo uses a general assistant prompt, so it is an exploration tool rather than an exact reproduction of the short-answer benchmark. Both models run on one machine, and requests are queued one at a time.
PopQA was built in 2023 from Wikidata, and any of these models may have encountered its questions during training. Their exposure may also differ, so using the same evaluation conditions does not control for contamination. The scores should not be read as recall of unseen facts. The alias-containment rule is another limitation. It accepts some answers a human reader would reject, and I have not measured the size of that effect.