What Rabbit And The Duck Actually Is

Rabbit and the Duck is a reinforcement learning meta-learning framework. The core idea is straightforward: you train an agent in one environment (the rabbit) and test whether it can adapt to a structurally different environment (the duck). The name comes from the mismatch between source and target domains, which is exactly what makes the problem interesting. The framework was designed to test generalization in meta-RL settings, where most baselines just shift the same environment around with small parameter changes. Shifting a reward function by 0.1 is not the same as testing whether an agent can handle a completely different dynamics model.

Why People Use Rabbit And The Duck

Most researchers use it to benchmark meta-learning algorithms. You train on a distribution of rabbit tasks, then evaluate on duck tasks. If your method scores well here, it suggests the agent learned something transferable rather than just memorizing the source domain. I started using this because standard meta-RL benchmarks were giving inflated numbers. My agents were essentially overfitting to the task distribution, and the performance gap between in-distribution and out-of-distribution testing was huge. Rabbit and the Duck exposed that much faster.

How to Set It Up

The easiest path is through the published implementation. Clone the repository and install it with pip. The dependency list is manageable — PyTorch, NumPy, and a few Gymnasium environments. I'd recommend running it inside a conda environment to avoid version conflicts, especially with MuJoCo if your benchmark includes it. Installation steps: First, create the environment and install the base dependencies. Then pip install the package from source. Run the example script to verify everything loads correctly before modifying anything.

Get the Full Details

White Rabbit on Green Grass · Free Stock Photo
White Rabbit on Green Grass · Free Stock Photo

One thing most tutorials skip: make sure your Gymnasium compatibility layer is set up right. The older Gym wrappers still float around in some installations, and mixing them causes silent failures that are very hard to debug.

Training A Meta-Learner On Rabbit And The Duck

The training loop follows standard MAML-style inner and outer loops. You sample a source task from the rabbit distribution, run inner-loop adaptation for a few steps, then evaluate on the duck task. The meta-gradient flows through both phases. Here is the part beginners get wrong. The inner loop adaptation steps matter enormously. Too few and the agent cannot adjust to the rabbit task at all. Too many and it overfits to the source, leaving nothing transferable to the duck. In my experience, 4 to 8 inner-loop steps is the typical sweet spot depending on the environment complexity. I ran into a specific issue last year where the duck evaluation rewards were near zero across every algorithm I tested. I spent two days debugging the reward function before realizing the duck tasks had their goal positions initialized outside the reachable workspace of the agent. The benchmark was technically valid, but the tasks were trivially unsolvable by design. I had to filter the task set using the reachability check before training. It cost me about half a day to identify and fix.

What Actually Works And What Does Not

MAML and Reptile both work, but MAML tends to be more sample-efficient on the rabbit side while Reptile is more stable across different task distributions. ProtoRL approaches also show promise when the duck tasks differ in subtle ways rather than dramatic ways. Common pitfall: people treat the rabbit and duck split as fixed and report those numbers as the final result. That is a mistake. The generalization gap is sensitive to how you sample the task distributions. If you want honest numbers, run multiple random seeds for the task generation itself, not just for training. Another thing nobody mentions much: evaluation noise in the duck phase is higher than you would expect. Since the agent is adapting to something it has never seen, small random variations in the duck task initialization cause large swings in performance. Report confidence intervals, not just mean scores.

Rabbit Free Stock Photo - Public Domain Pictures
Rabbit Free Stock Photo - Public Domain Pictures

When Rabbit And The Duck Fails You

The framework assumes the duck distribution is reachable from the rabbit distribution through some form of structural similarity. If your domains are too far apart — say you train on a 2D point-mass task and test on a 7-DOF arm — the scores will be near random regardless of your method. That is not a bug in the framework, it is a feature. It tells you something honest about transferability. If you need to test cross-domain transfer with very different state spaces, consider looking at representation learning approaches instead. Methods that learn invariant features across domains handle that scenario better. Rabbit and the Duck is not designed for that use case. I also found that the original implementation does not include a pre-trained baseline for direct comparison. If you need quick baselines, you will have to implement MAML and Reptile yourself or adapt code from other repositories. Budget extra time for that.

Rabbit And The Duck Downstream Applications

Beyond benchmarking, the framework has been used to study how robots adapt to new physical objects without retraining from scratch. You train on one set of manipulation tasks, then test on objects with different geometries. The numbers are modest but consistent, which is about what you should expect from a single-domain generalization test. The paper release includes a leaderboard, but it has not been updated in a while. Newer methods may already outperform the published numbers, so check the arXiv feed for recent submissions that use this benchmark.