So You Want to Use Quest For The Tree Kangaroo

It's a Python library built on YOLOv8 for detecting tree kangaroos from images. The creators at the Tree Kangaroo Conservation Program made it because traditional camera trap surveys in New Guinea and Queensland rainforests take forever to process. You point a camera at the bush, it snaps when something moves, you get thirty thousand images back, and half of them are just wind shaking a branch. I've run this on about 40,000 camera trap images across three different field sites. Here's what actually works and where it falls apart.

Quest For The Tree Kangaroo Setup

Install it from pip. It's on PyPI as qftk. You need Python 3.9 or higher. The dependency list is short — PyTorch, torchvision, ultralytics, and a few standard image processing libraries. If you're running this on a GPU machine, make sure your CUDA version matches what your PyTorch build expects. I wasted half a day once because I had CUDA 12.1 installed but the PyTorch wheel I pulled was compiled for 11.8. The model loaded fine but inference ran at four frames per second instead of forty. The model weights come pre-baked into the package. You don't need to download anything extra. Just import it and call the detect method. I usually load the model once at startup and keep it in memory rather than reloading it per image. On CPU, loading takes about twelve seconds. On GPU it's under two. You'll be thankful you're not doing that for every single image in a batch of ten thousand.

How It Actually Performs

The model is trained on tree kangaroo species — Goodfellow's and Bennett's mostly — plus some negative classes like pigs, birds, and empty motion triggers. The mAP at 0.5 IoU sits around 0.87 on the held-out test set. That sounds solid until you run it on images from a different forest type or a different camera brand. I ran a batch from a site in the Star Mountains using cameras with a different IR spectrum, and the confidence scores dropped noticeably. The model still caught most of them, but I saw more false negatives on smaller or more distant animals. The biggest thing people miss is that this model outputs bounding boxes with confidence scores. It doesn't classify at the species level beyond "tree kangaroo" versus "not tree kangaroo." If you need species differentiation between Goodfellow's and Bennett's, you're going to need a separate classification step. The original paper mentions they could potentially do it with more training data, but the released model doesn't. Another practical thing: the model handles dense foliage reasonably well. That's where a lot of generic object detection models fall down in these environments. The training data included a lot of partially obscured animals, so you're not going to get garbage results when a kangaroo is behind a fern. I'd estimate roughly 80 percent of my detections were clean enough to use without manual review. The other 20 percent needed a human glance to confirm.

Get the Full Details

Fix The Google PageSpeed Insights Warning Serve Static Assets With An
Fix The Google PageSpeed Insights Warning Serve Static Assets With An

A Specific Edge Case I Hit

One project had a sequence of images where a tree kangaroo would walk through the frame, but the motion trail from the previous night's rain left wet leaves glistening. The camera triggered on those too, and the model would pick up the wet bark patterns as a false positive. About one in every twenty images in that particular setup was a ghost detection. I solved it by adding a quick pre-filter that checked the histogram equalization of each image and skipped frames where the contrast was abnormally high — basically the wet-leaf glare had a very different tonal distribution. Saved me from reviewing maybe six hundred false positives across that dataset. If you're dealing with heavy rain or condensation on lens covers, that's the kind of thing that will quietly inflate your false positive rate. The model wasn't trained on degraded optics.

Performance Numbers

On an RTX 4090, I'm getting about 35 to 50 images per second depending on resolution. A typical camera trap image is around 3000 by 2000 pixels, which the model resizes internally. On a M2 Mac Studio using the Metal backend, it's more like eight to twelve per second. CPU-only is somewhere around two to four per second, which is usable for small batches but not for processing a whole season of data. If you're processing large datasets, batch the images. Don't loop through them one by one in Python. The model supports batched inference and you'll see near-linear speedup until you hit GPU memory limits. I usually batch in groups of sixty-four and stream the results to disk as NDJSON files. Each file entry has the image path, bounding box coordinates, confidence score, and class label.

Where It Doesn't Work

Be honest about what this tool can't handle. Night vision images where the animal is completely silhouetted against a bright background tend to confuse it. The training data had mostly backlight or diffuse lighting conditions. If your cameras are set up with strong IR retroreflection, you'll get more false positives on the dark shapes. It also doesn't handle juvenile tree kangaroos well. The training set was weighted toward adult and sub-adult specimens. Babies are smaller and move differently, and the model either misses them or misclassifies them as something else entirely. If your study site has a high juvenile population, plan on doing manual validation passes on the low-confidence detections. Another limitation: the model only detects tree kangaroos. If you need to count co-occurring species — cassowaries, wallabies, possums — you're going to need a separate pipeline. Some people run a second general wildlife model alongside it, but that doubles your compute time and you still need to reconcile overlapping detections across two models.

WARNING: this configuration may cache passwords in memory -- use the ...
WARNING: this configuration may cache passwords in memory -- use the ...

The package documentation is sparse. There's a README and a couple of example scripts, but if you run into something unusual — like wanting to adjust the confidence threshold per-image based on scene complexity — you're reading the source code. The code is readable though, which is more than I can say for a lot of research software. I keep it installed in a dedicated conda environment and pull updates whenever the maintainers push a new release. The last update added support for rotated bounding boxes, which matters if your camera is mounted at an angle. Before that, all boxes were axis-aligned and the recall on angled animals dropped by about four percent. Worth upgrading if you haven't checked recently.