Setting Up Clickers for Browser Automation Tasks

Clickers is a Python framework that sits between an LLM and a browser, letting the model take actions like clicking buttons, filling forms, and reading page content through a structured interface. It emerged as one of the more practical approaches to the "AI agent that can actually use a browser" problem, and it has held up better than a lot of the competing solutions because it keeps the action space simple. At its core, Clickers gives an AI model access to a set of atomic browser operations: navigate, click, type, scroll, extract text, and so on. The model doesn't see a screenshot and guess at coordinates. Instead it sees a DOM representation and a list of available actions, then picks one. That single design decision matters more than most people realize. Coordinate-based UI automation degrades rapidly as screen resolutions change. Action-based selection stays stable. The framework is lightweight enough to run locally without heavy dependencies. You install it with pip, point it at a browser instance, and feed it prompts the same way you would any other agentic tool. It supports both Chrome and Firefox out of the box, and the API surface is small enough that you can read the source in an afternoon and understand exactly what happens at each step.

Installation and Basic Configuration

Get it from PyPI. The standard command works fine: pip install clickers. If you want the full version with support for headless mode and extended selectors, pull in the extra dependencies as well. You will need a running browser instance. Clickers connects to it through a debugger port, so either launch Chrome with remote debugging enabled or let Clickers spin one up for you. The headless flag is useful in production but can break certain page interactions on sites that detect headless browsers. I found that out the hard way. Here is a minimal setup script that demonstrates the basic pattern:

from clickers import BrowserAgent, ClickerAction agent = BrowserAgent(headless=False) agent.navigate("https://example.com")

Get the Full Details

Ludo Latte Clickers - Latte Fidget Clicker - 3d Printed Ludo Fidget toy - 3D model by LUDO on Thangs
Ludo Latte Clickers - Latte Fidget Clicker - 3d Printed Ludo Fidget toy - 3D model by LUDO on Thangs

actions = agent.get_available_actions() agent.take_action(actions[0]) This gets you from zero to clicking something on a page in roughly five minutes. The real work starts after this point.

A Realistic Use Case

Let me walk through something that actually came up in production. I was building a workflow to automatically fill out a government benefits portal. These sites are notoriously hostile to automation: they use dynamic element IDs, have aggressive rate limiting, and sometimes reload pages in ways that break session state. A naive script would crawl through and fail on the second form field. Using Clickers, I structured the automation as a loop. The model reads the current form state, decides which field to fill next, types the value, and submits. Between steps, I added a check that re-examines the page and confirms the previous action actually took effect. This meant the agent could recover from unexpected redirects or modal dialogs instead of silently proceeding with stale data. The specific problem I hit was that one page used a dropdown that appeared as a custom div rather than a standard select element. Most clickers-style frameworks handle native selects fine. This custom component responded only to mouse position events, not click actions. The workaround was to use Clickers' execute JavaScript capability to programmatically open the dropdown and trigger the click event with the correct parameters.

agent.execute_js(""" const dropdown = document.querySelector('.custom-dropdown'); dropdown.dispatchEvent(new MouseEvent('mousedown', {bubbles: true}));

🧲 CocoSnaps: 3d-printed magnetic clickers without springs・Free STL File for ・Cults
🧲 CocoSnaps: 3d-printed magnetic clickers without springs・Free STL File for ・Cults

""") This took about twenty minutes to diagnose and implement. Without that fallback capability, the whole automation would have stalled at that one form.

Common Pitfalls and What They Cost You

The biggest mistake I see is over-relying on the model's ability to plan ahead. Clickers is designed for single-action-at-a-time loops, but users often expect the agent to execute a full multi-step process in one shot. The model will hallucinate a sequence that looks correct but fails silently because it cannot observe intermediate states between steps. Break everything into small, verifiable actions. Each action should be followed by a state check before the next one. Another issue is action list bloat. As the page grows, the list of available actions can become enormous. On a complex dashboard I tested, the action list exceeded four thousand entries. Feeding that many options to a model introduces latency and increases the chance of selecting the wrong action by confusion. The solution is aggressive filtering: trim the action list to only interactable elements in the viewport, group related actions, and discard elements behind modals or overflow containers. Session management is also underdocumented. Clickers maintains browser state between calls, but if the page navigates externally or the session times out, the agent may continue operating on a stale DOM view. Always verify the current URL before assuming the session is intact. A three-line check at the start of each loop iteration saved me from hours of debugging on a separate project.

Performance and Scaling

Clickers is single-threaded by design. Running multiple instances in parallel requires spawning separate processes, each with its own browser instance. This adds overhead but is necessary if you are processing more than a few pages simultaneously. I typically run four parallel agents on a machine with 16GB of RAM, and each one gets about 4GB. Beyond that, the system starts swapping and performance degrades sharply. For low-volume automation where you need reliability over speed, Clickers is hard to beat. It handles complex pages better than most alternatives in its weight class. For high-throughput scenarios where you are processing thousands of forms per day, you are better off with a dedicated RPA platform or a custom Selenium-based solution that you can optimize specifically for your target pages.

3d Printed Clickers - Etsy
3d Printed Clickers - Etsy

When Clickers Is the Wrong Tool

There are legitimate cases where you should not use it. If your target site requires CAPTCHA solving, Clickers does not include that capability and you would need a third-party service on top of it. If you need visual regression testing or pixel-perfect screenshot comparisons, use a proper testing framework instead. If the site uses heavy client-side rendering with frequent AJAX updates, Clickers may report stale DOM content while the model makes decisions based on outdated information. In those cases, adding a wait-and-verify loop between actions is essential, but it slows everything down considerably. The framework also does not handle authentication flows well out of the box. Single sign-on, OAuth redirects, and multi-factor authentication require manual handling. I built a wrapper around Clickers that pre-seeds cookies from a stored session, which works for about eighty percent of cases. The remaining twenty percent involve active login flows that need scripted interaction.

A Quick Alternative Comparison

If Clickers does not fit your situation, here are the main alternatives and when they make sense. AgentQL excels at query-based extraction from complex pages but lacks the broad action repertoire for true automation. Playwright offers deeper integration and better performance for high-volume tasks but requires more infrastructure. LangChain-style agent frameworks add reasoning layers on top but introduce their own failure modes around tool selection and state tracking. Pick the one that matches your bottleneck, not the one with the most documentation. Clickers is a solid choice when you need a simple, readable automation framework that gets out of the way. It will not solve every problem, and the limitations are real. But for the middle ground between a raw Selenium script and a heavy enterprise automation platform, it occupies a useful space. Test it against your target pages before committing to it as the backbone of a production system. The few hours you spend on a proof of concept will save you days later. Documentation lives at the official Clickers GitHub repository, and the PyPI page has installation instructions that cover the basic setup. The source code itself is the best reference material available. Most of the edge cases I described above are only obvious after you have spent time reading through the action dispatch logic yourself.