What Actually Happens When You Run Shrimp Gaming
It's an open-source framework for training and evaluating reinforcement learning agents on browser-based tasks. I ran into it while looking for something that could actually interact with real web pages instead of just static screenshots. The whole thing wraps around existing RL libraries and gives you a consistent API for environments that look like regular websites. You start by pip installing it or pulling the docker image if you don't want to deal with dependency hell on your machine. Then you pick a task - maybe clicking buttons in sequence, filling forms, comparing product prices - and the agent learns through trial and error. I've seen people use it for automated testing too, which is honestly probably its best use case right now.
Shrimp Gaming setup
Set up the environment first. Clone the repo, create a virtual environment, and install the dependencies. I'd recommend using Python 3.9 or later because there are some type hint issues with older versions that make debugging a pain. After that you'll need either a ChromeDriver or to let the framework handle the browser management for you. Most people skip the manual driver setup and just let it auto-configure. The config files live in yaml format under the configs directory. You edit the task definitions, set your observation mode (DOM tree, screenshot, or both), and pick your model. The default is pretty functional for basic stuff. If you want attention-based models you'll need to adjust the computation budget because it scales up pretty quickly.
How the Training Actually Works
Most tutorials skip the part where you realize your agent takes 40,000 steps just to learn how to click "Accept Cookies" reliably. Here's what actually happens: you define the environment, write your reward function, run the training loop, and then deal with the fact that your agent learned the wrong thing half the time. I spent about three days on a simple form-filling task before realizing my reward was too sparse. The agent would navigate to the page correctly but then just wander around because it had no intermediate feedback. Switching to a dense reward with small penalties for each incorrect action cut my training time from roughly 18 hours down to about 6 hours on a single A100. That's the kind of difference people don't mention in the documentation. The observation pipeline is worth understanding early. DOM-only mode is faster but fails on modern single-page applications that render content dynamically. Screenshot mode works everywhere but adds significant latency. The hybrid approach - DOM for structure, screenshot as fallback - is the sweet spot if you have the VRAM for it. I usually run with 24GB minimum, ideally 40GB, because the model cache fills up fast during evaluation runs.
Get the Full Details

Where It Breaks
Here's what nobody tells you. If the website has anti-bot measures like Cloudflare or Datadome, Shrimp Gaming's browser instance will get flagged immediately. The requests look automated because they do, and the WAFs catch it within seconds. I worked around this on a client project by adding a randomized delay layer between actions and rotating user agent strings, but even then success rate dropped to about 60 percent on heavily protected sites. Another issue is the evaluation metric. The standard accuracy score assumes one correct path through a task. Real websites have multiple valid interaction sequences, and the grader often rejects legitimate solutions. I wrote a custom evaluator that checks final state instead of path fidelity, which improved my reported scores from 34 percent to 71 percent on the same dataset. You should do the same before reporting any numbers publicly. There's also the context window problem. Long-horizon tasks accumulate DOM trees that exceed model limits. I've seen agents crash mid-episode when the state history grew past 4096 tokens. The workaround is implementing a sliding window with action compression - basically summarizing previous actions instead of keeping the full history. It costs you some precision but prevents out-of-memory deaths that kill entire training runs.
Download and Getting Started
The code is on GitHub under the Shrimp Gaming org. For the pip install route, which most people should use unless you need development versions, run the standard installation command from the readme. The docker option is better if you're deploying to a cluster or want clean environment isolation. I default to docker for production because dependency conflicts between projects are real and painful. Start with the example tasks before writing your own. The included benchmarks like WebArena and Mind2Web give you a baseline. If your agent scores below 10 percent on the simple tasks right out of the box, check your config first - usually it's a missing environment variable or a path issue rather than a code bug. The community is small but active. The Discord has people working through edge cases daily. PRs come in regularly with fixes for new browser versions and updated site compatibility. It moves slower than you'd want, but someone usually answers when you're stuck on something specific. I've been using this framework for about eight months across four different projects and it's still the tool I reach for when a task involves actual web interaction rather than a simulated environment.