What Actually Happens When You Merge Models
You load two fine-tunes into memory, point your merge tool at them, and wait for output. The first few samples look fine. Then you notice the perplexity climbing. The model starts repeating phrases it shouldn't know. Words get scrambled in ways that look almost intentional, like it's trying to say something but can't find the path through its own weights. This is Merge Rot, and it's been around since people started doing model arithmetic without understanding what they were actually merging. I spent about six months dealing with this during a project where we tried blending three medical-domain adapters onto a base LLaMA-3-8B. Two of them converged cleanly. The third one introduced what I now call semantic drift cascading — the model would generate perfectly coherent English sentences until roughly token 40, then suddenly switch to a register that looked like it was hallucinating training data from a completely different domain. Not random noise. Structured nonsense. The kind of output that makes you wonder if the model is lying to you.
Why Merge Rot Happens
Model weights are high-dimensional surfaces. When you merge two sets of weights using simple averaging or Slater interpolation, you're assuming those surfaces are convex in the region you're working in. They're not. Fine-tuning pushes weights toward local minima that correspond to specific behaviors. Average two minima together and you land somewhere in the valley between them where the gradient doesn't point toward either behavior anymore. The architecture makes this worse. Modern LLMs have LayerNorm, RMSNorm, attention patterns, and MLP blocks that evolved to work together as a system. Change the weight distribution in one block and the downstream effects compound. A small perturbation in an attention head's query matrix might look harmless in isolation, but by the time it reaches the output layer, the semantic space has warped enough that the model generates fluent-but-nonsensical text. There's also the issue of activation entropy collapse. When you merge models trained on different data distributions, the merged model's internal activations can collapse toward lower entropy states. The model becomes overconfident. It starts assigning high probability to incorrect next tokens because the merged weight space doesn't properly represent the uncertainty it should have. You see this as the model confidently generating wrong answers instead of hedging or staying neutral.
The Practical Workarounds
Simple arithmetic merging works when models are closely related — same base, same task, similar data. It fails when they diverge. Here's what I've learned the hard way. Task-vector merging is the most reliable approach I've used. Instead of averaging raw weights, you compute the difference between the fine-tuned model and the base model, then add that delta to a fresh base. The math looks like this: merged = base + × (model_A - base) + × (model_B - base)
Get the Full Details

The coefficients and control how much of each adapter you're injecting. This usually keeps the model anchored to the base's general capabilities while importing task-specific knowledge. I found that setting = 0.7 and = 0.5 worked for my medical project. Going higher caused the semantic drift I described earlier. SLERP merging (spherical linear interpolation) is another option. It interpolates along the geodesic between weight vectors rather than a straight line. This preserves the norm of the weight space better and reduces the chance of landing in pathological regions. The trade-off is computational cost. SLERP takes about 3x longer than simple averaging for large models, which matters when you're iterating. For my medical domain merge, I ended up using a hybrid approach. Task-vector merging for the base structure, then SLERP on the attention layers only. The MLP blocks stayed with task-vector. This cut the semantic drift from visible after token 40 down to invisible until token 120 or so. That gave the model enough coherence for most practical use cases.
Edge Cases That Break Everything
I ran into a specific problem with LoRA-based merges. If you're merging models that used different rank values in their LoRA adapters, the weight matrices don't align. A rank-16 adapter and a rank-32 adapter operating on the same layer produce incompatible shapes. You can't just average them. The merge tool will either crash or silently produce garbage. The workaround is to pad the smaller adapter to match the larger rank before merging. Zero-pad the weight matrices, or better yet, retrain the smaller adapter at the higher rank if you have the compute budget. I chose padding for speed. It works reasonably well for ranks up to 64. Beyond that, the quality degradation becomes noticeable. Another gotcha: quantization state mismatch. If one model is stored in FP16 and another in INT8, merging them directly introduces quantization error into the result. The merged model will have worse accuracy than either input. Always convert both models to the same precision before merging. I learned this when a teammate merged a quantized distillation model with a full-precision fine-tune and got outputs that looked like they were running through a bad denoiser.
When Merging Is the Wrong Tool
Not every problem needs a merged model. If you're trying to combine two models that solve fundamentally different tasks, the result will usually be worse than just running them separately and routing queries to the appropriate one. Task-aware routing through a lightweight classifier or even a simple prompt template often outperforms any weight merge. There's also the question of compute. Merging models sounds cheap — you do it once and ship the result. But validation is expensive. You need test sets that cover the overlap region between the source models' capabilities. For my medical project, I spent more time building evaluation benchmarks than I did on the actual merge. The merge itself took about 15 minutes on a single A100. The validation took three days. If you're dealing with models that have diverged significantly in their training data or objectives, consider consensus decoding instead. Run both models on the same input, compare their outputs, and use a voting mechanism or a secondary model to pick the best response. This avoids weight-space pathologies entirely. The latency cost is higher — you're running two forward passes instead of one — but the output quality is usually superior.

Download and Tools
For practitioners looking to experiment, the most stable merge pipeline I've found is based on the mergekit library. It supports task-vector merging, SLERP, and several interpolation schemes. The current version handles most common architectures including LLaMA, Mistral, and Qwen variants. mergekit on GitHub I also recommend keeping a copy of the base model separate from your merge workflow. When things go wrong — and they will — you need a clean reference point to determine whether the merge introduced the problem or whether it was already there.
The short version: merge models carefully, validate thoroughly, and don't trust the first output you see. Merge Rot is real, but it's manageable if you understand what you're doing to the weight space.