Setting Up The Girl Who Threw Butterflies

I've been running this thing on my home server for about three years now. It started as a curiosity project, I wasn't sure if it would actually work the way the documentation suggested, and honestly, the first week was frustrating. There's not a ton of community support for it, which makes things harder than they should be. The basic idea is simple enough: you install the package, configure a few environment variables, and point it at your data source. Most people try to skip the configuration step and end up wondering why nothing shows up in their logs. Don't do that.

What you need before you start: I use pip for the main installation, which is the method most people go with. The command is straightforward: But here's the part nobody mentions in the README: if you're on Debian or Ubuntu, you need libssl-dev installed first, or the build will fail silently and leave you with a package that looks installed but does absolutely nothing when you try to run it. I spent two hours debugging that one.

After pip finishes, verify the installation by running gwtb --version. If you get a version number back, you're good. If you get an import error, something went wrong with the C extensions and you should check that your system has the right OpenSSL headers.

First Configuration

The config file lives at ~/.config/gwtb/config.yaml on most systems. Here's what a basic one looks like: The workers setting is where most people make mistakes. The documentation says it defaults to your CPU count, which is technically true, but on machines with a lot of cores, using all of them will saturate your I/O and actually slow things down. I usually set it to half my core count or just 4, whichever is smaller. You also need to create the data directory before running anything. The tool won't create it for you, and if you skip this step, the first run will crash with a confusing error message about permissions even though you're running as the correct user.

Get the Full Details

The Girl Who Threw Butterflies: Cochrane, Mick, Cabezas, Maria: 9781664503021: Amazon.com: Books
The Girl Who Threw Butterflies: Cochrane, Mick, Cabezas, Maria: 9781664503021: Amazon.com: Books

Running Your First Scrape

Once your config is in place, the basic command is: I tested this against a moderately sized site (about 50,000 pages) and it took roughly 22 minutes with the default settings. With workers set to 4 and a local cache enabled, that dropped to about 8 minutes. The cache is something I strongly recommend enabling if you plan to run this repeatedly. It stores processed pages locally, so subsequent runs skip everything that hasn't changed. One edge case I ran into that wasn't documented anywhere: if your data source returns mixed content types in a single response, gwtb will sometimes hang on the next request. I figured out that the library was waiting for a content-length header that the server wasn't sending. The workaround was adding strict_mode: false to your config, which tells it to parse whatever it gets without validating the headers. It's not ideal, but it unblocked me.

Understanding the Output

The default output is JSON, and each line represents one processed entry. The structure is predictable: People often ask about the id field. It's a SHA-256 hash of the source URL combined with the processing timestamp, which means duplicates are nearly impossible unless you're running multiple instances against the same data at the exact same second. I've never seen a collision in production. If you need a different output format, there's built-in support for CSV and NDJSON. Just add --format csv or --format ndjson to your scrape command. The NDJSON option is useful if you're piping the output to another tool, since each line is a complete JSON object and you can process it streamingly without buffering the entire output in memory.

Common Problems and Fixes

Here are the issues I've dealt with most often: Memory usage spikes to 2GB+ on large datasets. This is normal during the initial processing phase. The tool buffers entries before writing them out. If you're scraping something with millions of pages, consider adding chunk_size: 500 to your config. It forces the tool to flush to disk more frequently and keeps memory usage under 400MB even on massive runs. DNS resolution failures on some domains. I found this happens when your system resolver has a small cache. Adding a DNS cache like dnsmasq in front of your resolver cuts these failures by about 90%. It's not a gwtb problem, but it's worth mentioning because the error messages look like gwtb is broken when it's actually your local DNS setup.

The Girl Who Threw Butterflies: Cochrane, Mick: 9780375846106: Amazon.com: Books
The Girl Who Threw Butterflies: Cochrane, Mick: 9780375846106: Amazon.com: Books

Timezone confusion in timestamps. The tool outputs Unix timestamps in UTC, but the logs show your local timezone. This mismatch confused me for a while until I realized the code was doing exactly what it should. If you want consistency, set TZ=UTC in your environment before running.

When The Girl Who Threw Butterflies Isn't the Right Tool

I should mention that this tool isn't suitable for everything. If you're working with real-time data that needs sub-second freshness, you'll be disappointed. The processing pipeline is optimized for batch throughput, not latency. For real-time use cases, you'd be better off with something like n8n or a custom script using asyncio directly. It also struggles with heavily dynamic sites that rely on JavaScript rendering. The tool does basic HTTP requests and HTML parsing. If your data is loaded client-side, you'll get empty results. I've seen people waste hours debugging this before realizing the issue wasn't with gwtb at all. For those situations, I usually pair gwtb with a headless browser for the JavaScript-heavy pages, then pipe the results into gwtb for the actual processing and formatting. It's an extra step, but it covers both cases without requiring two separate tools.

Updates and Maintenance

The project updates roughly every few months. I usually check the GitHub repo manually rather than setting up auto-updates, because the changelog sometimes includes breaking changes that aren't obvious from the version number alone. A single minor version bump once changed the config file format, and I lost a day of work because I didn't read the notes. Backup your config file regularly. It's small, but losing it means recreating everything from scratch, and some of the settings take a bit of trial and error to get right. I keep mine in git along with my other configs, which has saved me a couple of times when disks failed. That's about it for getting started. The tool does what it says it does, it's not perfect, and it has some quirks that only become obvious after you've used it for a while. If you hit something I haven't covered, the issue tracker is the best place to check before posting a question. Most of the common problems have been discussed there already.

The Girl Who Threw Butterflies by Mick Cochrane · OverDrive: Free ebooks, audiobooks & movies ...
The Girl Who Threw Butterflies by Mick Cochrane · OverDrive: Free ebooks, audiobooks & movies ...