Getting Started With Gubby
Gubby is a lightweight Python-based tool designed for automating repetitive UI interactions and web scraping workflows. It sits somewhere between Selenium and Playwright in capability, but the API is intentionally simpler. I started using it about two years ago when a project required low-maintenance browser automation without the overhead of a full framework. At its core, Gubby is a wrapper around headless Chromium with some built-in abstractions for element handling, session management, and retry logic. You write Python scripts, define selectors, and it handles browser lifecycle management. Unlike Selenium, you don't manage WebDriver instances manually. Unlike Playwright, there are no async APIs or TypeScript bindings — it stays strictly Python. Installation is straightforward. You can install it via pip:
pip install gubby It pulls in its own Chromium binary on first run, so you won't need to separately download ChromeDriver or set PATH variables. That alone saves me about 15 to 20 minutes per project setup. The initial Chromium download is roughly 130MB, which is larger than I'd like, but it's cached after the first run. For production environments where you can't download binaries on demand, Gubby supports pointing to an existing Chrome installation. Set the environment variable GUBBY_CHROME_PATH to your Chrome executable location, and it will skip the bundled binary entirely.
Basic Usage
Here is a minimal example of how I structure my scripts:
from gubby import Browser, Selector
with Browser(headless=True) as browser:
page = browser.open("https://example.com")
title = page.find(Selector.css("h1"))
print(title.text)
links = page.find_all(Selector.css("a[href]"))
for link in links[:5]:
print(link.href)
The context manager handles browser startup, teardown, and cleanup automatically. If a script crashes mid-execution, the browser process gets killed. This has saved me from debugging phantom zombie processes multiple times.
Get the Full Details
Session Persistence and State
One thing beginners often miss is how Gubby handles sessions. Unlike basic scraping tools, Gubby maintains cookies, local storage, and even authenticated states across multiple page.open() calls within the same Browser instance. I use this for workflows that require login followed by data extraction. After one successful auth flow, I can persist the session to disk and reload it on subsequent runs. The session files are stored in ~/.gubby/sessions/ by default. To save and restore a session:
browser.save_session("my_login_session")
... later ...
browser.load_session("my_login_session")
This usually cuts repeated login overhead from about 30 seconds down to under 5 seconds because the auth tokens are reused.
Handling Dynamic Content
Gubby includes a built-in wait system. You don't typically need explicit waits like in older Selenium workflows. The Selector classes accept timeout parameters, and Gubby polls for elements automatically:
result = page.find(
Selector.css(".dynamic-content"),
timeout=15,
poll_interval=0.5
)
The default timeout is 10 seconds. I've found that raising poll_interval to 0.5 or even 1.0 second significantly reduces CPU load on pages with heavy JavaScript rendering, without noticeably increasing total wait time.

A Real Problem I Encountered
Early on I ran into an edge case where a target site used anti-bot detection that flagged Gubby's default user agent and fingerprint. The site would return a CAPTCHA or simply blank out after a few requests. The fix was straightforward but not immediately obvious from the docs: Gubby supports custom fingerprint profiles. I created a configuration object with randomized browser properties and passed it at instantiation:
from gubby import Browser, FingerprintConfig
config = FingerprintConfig(
user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) ...",
viewport=(1366, 768),
locale="en-US",
platform="Win32",
)
with Browser(headless=False, fingerprint=config) as browser:
...
Running with headless=False was actually critical here because some detection systems differentiate between headless and non-headless browsers. Pairing that with the custom fingerprint dropped my block rate from roughly 40% to under 5% on that particular target. It's not foolproof — if a site uses aggressive behavioral analysis like mouse movement tracking, no fingerprint tweak will fully protect you. But for most standard scraping tasks, this combination works reliably.
Advanced: Chaining and Data Extraction
Gubby supports method chaining, which keeps scripts readable for longer workflows:
results = (
page
.find(Selector.css(".product-grid"))
.find_all(Selector.css(".product-card"))
.map(lambda card: {
"name": card.find(Selector.css(".product-name")).text,
"price": card.find(Selector.css(".price")).text,
"url": card.find(Selector.css("a")).href,
})
)
The .map() method operates on the list of found elements and returns a new list. It's essentially a built-in list comprehension for DOM traversal, which cuts down on boilerplate considerably.
Known Limitations
Gubby does have clear limitations worth noting before you commit to it for a large project. First, it has no mobile device emulation. If you need to test or scrape mobile layouts, you're out of luck — stick with Playwright or Puppeteer for that. Second, the documentation is sparse beyond the basics. The GitHub repo has examples, but there's no formal API reference, so you end up reading source code to understand edge cases. Third, concurrency support is limited. You can run multiple Browser instances in parallel threads, but there's no built-in pool or semaphore management like you'd find in more mature frameworks. For small-scale scripts (under 20 concurrent browsers), this isn't a problem. Beyond that, you'll need to roll your own orchestration. Another practical issue: error messages are often unhelpful. A failed element lookup might throw a generic TimeoutError without indicating whether the selector was malformed, the element existed but was hidden, or the page simply didn't load. I've spent hours debugging what turned out to be a typo in a CSS selector because the error output gave zero hints.
When to Use Gubby vs. Alternatives
Use Gubby when you need a quick, Python-only solution for routine scraping or UI automation and don't want to manage WebDriver binaries or learn a new framework. It's suitable for scripts that run intermittently, handle straightforward pages, and don't require heavy concurrency or mobile testing. Avoid it if your project demands: high-concurrency browser pools, mobile emulation, robust anti-detection features, or enterprise-grade error handling. For those scenarios, Playwright is the better investment despite its steeper learning curve. It handles many of Gubby's shortcomings natively.
Where to Get It
The official package is available on PyPI at https://pypi.org/project/gubby/. The source code and issue tracker live on GitHub under the repository name gubby. There is no paid tier or commercial license — it's open source under the MIT license. If you run into issues, the most active discussion happens in the GitHub issues tab. The maintainer responds reasonably quickly, usually within 24 to 48 hours, which is better than average for tools at this scale. Community contribution is low though, so bug fixes tend to come solely from the primary developer.