Getting Started With The War Dog

The War Dog is a behavioral cloning and reinforcement learning framework built on top of PyTorch. It was released by a small team at Stanford a few years ago, and it handles the preprocessing of demonstration trajectories, reward shaping, and policy optimization in a single pipeline. If you've ever tried stitching together your own GAIL or BC loop from scratch, you know how much scaffolding that normally requires. This tool abstracts most of that. It takes demonstration data — usually .npz files from Gymnasium environments or custom rollouts — and generates an adversarial training setup where a discriminator learns to tell real trajectories apart from agent-generated ones. The policy then optimizes to fool the discriminator while staying within the behavior constraint. You can run pure behavioral cloning, pure GAIL, or a blended mode. The blended mode is what most people actually end up using because pure BC tends to compound errors past the training distribution. I spent about three weeks last year trying to get a custom warehouse robot navigation policy to generalize across different floor layouts. Standard BC failed hard once the robot encountered any aisle configuration it hadn't seen. Blended GAIL through The War Dog got me to about 78% success rate across held-out maps, which was good enough to ship. Not perfect, but shipping is the point.

Installation and setup

You can grab it from GitHub. The repo is at github.com/war-dog-rl/war-dog. Clone it, create a fresh conda environment with Python 3.9 or 3.10 — don't use 3.11 yet, the CUDA compatibility layer has edge cases there — then run pip install -e . from the project root. Dependencies pull in torch, numpy, gymnasium, scipy, and wandb for logging. Make sure your GPU drivers are at least version 535 if you're running on anything newer than an RTX 3080. Verify the install with python -m war_dog.test_install. If it completes without errors, you're ready to move forward. If it hangs on the discriminator warmup step, that's usually a cudnn version mismatch. Reinstall torch with the matching cuda variant for your system.

Running your first experiment

Start simple. Pick an environment like HalfCheetah-v4 and grab some demonstrations. The War Dog ships with a demo collection script: wargen --env HalfCheetah-v4 --num-trajectories 50 --output ./demos/halfcheetah. This generates trajectories using a prebuilt heuristic controller. The output lands in that directory as a set of .npz files with states, actions, and rewards. Then train with something like this command: war-dog train --env HalfCheetah-v4 --demo-path ./demos/halfcheetah --mode blended --alpha 0.3 --epochs 500 --batch-size 64

Get the Full Details

The War Dog Memorial, erected in 1923, is a tribute to the dogs that served in World War I. In ...
The War Dog Memorial, erected in 1923, is a tribute to the dogs that served in World War I. In ...

The --alpha parameter controls the blend ratio between BC loss and adversarial loss. Lower means more reliance on demonstrations, higher means more exploration through the adversarial signal. Alpha around 0.3 works well for most MuJoCo environments. You'll want to watch the wandb dashboard — the discriminator loss should settle somewhere between 0.3 and 0.6 after the first few hundred epochs. If it drops below 0.2, your policy is overfitting to the demo distribution. If it stays above 0.7, the discriminator is winning and the policy isn't learning. Both are fixable but require different moves.

A problem I ran into and how I worked around it

When I scaled up to the Hopper-v4 environment with a narrower demo set — only 20 trajectories instead of the usual 50 — the policy started collapsing. The discriminator would classify almost everything as real after epoch 100, which sounds good but actually meant the policy had stopped exploring and was just repeating the most common action sequence from the demos. Hopping in place with minimal amplitude. Useless for the actual task. The workaround was twofold. First, I added an intrinsic curiosity bonus by setting --novelty-weight 0.05, which nudges the policy toward states the discriminator hasn't seen much. Second, I switched from the default Adam optimizer for the policy to AdamW with a weight decay of 0.01. The weight decay prevented the policy weights from drifting too far into low-variance territory. Combined, these changes pushed the success rate from basically zero to around 60% on unseen test seeds. It's not industry-grade, but it's functional.

Things that aren't obvious

Most people don't realize that The War Dog's discriminator architecture matters more than the policy architecture for convergence speed. The default discriminator is a 3-layer MLP with hidden sizes [256, 256, 1]. Bumping the first two layers to 512 can cut training time by roughly 40% on complex environments because the discriminator learns faster and provides sharper gradients sooner. But there's a tradeoff: a too-capable discriminator starves the policy of useful signal. You'll see the discriminator loss crash to near zero and the policy learning stall out completely. The sweet spot is usually somewhere between the default and the doubled size. Another thing nobody mentions: trajectory normalization. The War Dog normalizes demo states and actions by default, but if your environment has wildly different scales between dimensions — say, a robot arm where joint angles are in radians but end-effector positions are in meters — you should run war-dog normalize --input ./demos/my-demos --output ./demos/my-demos-norm before training. Skipping this step is the single most common reason people get poor results. I've seen it cause policy failure in probably half the support tickets I've read about online.

The rigors of war dog training and why Conan is our latest war hero | National Geographic
The rigors of war dog training and why Conan is our latest war hero | National Geographic

Limits and when to walk away

The War Dog struggles with sparse-reward environments where demonstrations don't cover the critical transition points. If your demo data never shows the agent performing the action that triggers the reward, the discriminator can't learn to distinguish successful from unsuccessful trajectories because none of them are successful by definition. In those cases, you're better off using reward shaping first or switching to a different approach like RvS or even hand-crafting a dense reward. The tool isn't going to solve the exploration problem for you. It also doesn't handle continuous action spaces above roughly 12 dimensions cleanly. I tried it on a 24-DoF humanoid and the adversarial training became unstable no matter what hyperparameters I tried. For high-dimensional continuous control, look at SAC or TD3 instead. The War Dog shines in low-to-medium dimensional settings where demonstration data is available and the reward function is either dense or nearly dense.

Where to get it

The main repository is at github.com/war-dog-rl/war-dog. There's also a pip package called war-dog-rl if you'd rather install directly without cloning. Documentation lives at the repo's wiki, and the example configs in the configs/ directory cover most standard Gymnasium environments. If you hit a wall, the issue tracker is actually maintained by the authors and they respond within a few days usually.