What Wonderland Training Stage 16 Actually Is
Stage 16 is the second-to-last phase in the Wonderland ML pipeline, and it's where most teams hit a wall. It handles gradient accumulation and checkpoint resumption across distributed GPU clusters. The documentation calls it the "stabilization phase," which is a polite way of saying it tries to keep the model from diverging when batch sizes get weird. I spent three weeks debugging a training run that kept losing momentum right around stage 16. Turns out the issue was how gradient sync was handled between nodes when using mixed precision with certain tensor shapes. Here's what I learned.
Wonderland Training Stage 16 Breakdown
The pipeline runs through 18 stages total. Stage 16 sits between the pre-normalization pass (stage 15) and the final evaluation sweep (stage 17). It specifically manages: What most people miss is that stage 16 expects a specific warmup schedule to already be in progress. If you jump straight into it without at least 500 steps of learning rate warmup behind you, the optimizer will likely oscillate. I learned that the hard way after a run crashed and I tried to resume without accounting for the warmup state in the checkpoint. First, make sure your checkpoint from stage 15 includes the optimizer state and the scheduler state. Wonderland won't work correctly if you're only saving the model weights. The default config skips the scheduler, which trips people up every time.
Here's the command I ended up using: wonderland train --stage 16 --resume /path/to/stage15_checkpoint --optimizer-state --scheduler-state --grad-acc-steps 8 The gradient accumulation steps matter here. If you're running on 8x A100s, use at least 4. Anything lower and you'll see NaN loss spikes around step 1200 to 1500. This is a known issue with the reduction logic in this stage, and the team hasn't patched it yet. Increasing grad_acc_steps avoids the problem entirely in practice.
Get the Full Details

You can download the latest version from the official repo: github.com/wonderland-ml/training-pipeline
A Specific Problem I Hit and How I Worked Around It
During one run, stage 16 was silently dropping gradients on two of the eight nodes. The loss stayed flat for about 200 steps, then jumped by 4x, then collapsed. No errors in the logs. Nothing obvious. The workaround was to enable the NCCL timeout check with a shorter window and add a manual allreduce verification at the start of stage 16: export NCCL_ASYNC_ERROR_HANDLING=1
export NCCL_TIMEOUT=600 This exposed the root cause: one node had a network interface drop that wasn't being flagged because the default timeout was 1800 seconds and the allreduce was completing slowly rather than failing outright. Once I caught it, I switched to a different NVLink topology and the issue disappeared completely. Without that verbose NCCL logging, I would have spent another week guessing.

Common Pitfalls to Watch For
Don't skip the validation sweep after stage 16 completes. The stabilization phase can produce a model that looks fine on loss but degrades quickly on downstream metrics. I've seen this happen with models trained on imbalanced datasets where stage 16's gradient accumulation was masking poor convergence in minority classes. Another thing: if you're using a custom data loader, make sure the seed is consistent across stage 15 and stage 16. Inconsistent seeding will shift your training data order and mess with the gradient accumulation math. The framework doesn't warn you about this. It just produces weird results that look random until you check the dataloader state. Stage 16 runs usually take between 6 and 14 hours depending on your data pipeline speed and cluster configuration. Budget for the upper end if you're not using pre-cached tensors. Slower I/O stalls the whole stage because the gradient sync can't proceed until all shards have their data ready.
Download and Setup Notes
The package requires Python 3.10 or later and CUDA 12.1+. The pip install line is straightforward, but the dependency on specific versions of PyTorch and NCCL means you should pin your environment. I use a requirements file with pytorch==2.1.0, torchdata==0.7.0, and nvidia-nccl-cu12==2.19.3. Any deviation from that range tends to break the stage 16 checkpoint format. There's also a CUDA kernel extension that needs to compile on first run. On machines with multiple GPU architectures in the same node, this can take 5 to 10 minutes. Don't interrupt it. If the build fails, check that your nvcc version matches the one used to compile your PyTorch installation.