So Called Life Goes On - The Real Guide
I spent about three weeks wrestling with this before I figured out how to actually use it without wasting time. My So Called Life Goes On isn't exactly plug-and-play the way the docs make it look, and the community thread from 2023 basically admitted that a lot of people gave up on the default configuration. I'm going to walk through what I learned so you don't have to go through the same headaches. The first thing to understand is that this thing runs on a modified scheduler architecture. It sounds more complex than it is. You can think of it as a background job runner with built-in dependency resolution, but the catch is that the dependency resolution only works if your tasks are defined in a very specific way. If you skip that step and just throw random functions at it expecting things to wire themselves up, it will silently fail on certain environments and you'll be pulling your hair out trying to figure out why nothing's executing.
Getting My So Called Life Goes On Installed
Grab the latest release from the official repository. At the time of writing that's version 4.2.1. Install it with pip if you're on Python, though the npm package works fine too if you prefer Node. I personally run it in Docker because the host dependency conflicts are not worth dealing with. The docker-compose file they provide works out of the box for basic setups, but it's missing the config path override that you absolutely need for production. I'll get to that. Once it's installed, run the init command. This creates your config directory structure. Do not skip this step or you'll spend an hour debugging FileNotFoundError issues that should have never come up in the first place.
Configuration That Actually Works
Here's the thing nobody warns you about: the default config file assumes you're running a single-worker setup on a machine with at least 8GB of RAM. If you're running anything less than that, the default settings will cause your queue to back up exponentially. I hit this exact problem last year on a container that was only allocated 4GB. After about six hours of processing, the memory usage was at 94% and jobs were starting to get dropped silently. The logs just showed empty entries where completed jobs should have been logged. The workaround is straightforward but poorly documented. You need to set the worker_count parameter in your config to match your actual available memory divided by roughly 512MB per worker. For my 4GB container that meant dropping it to 7 workers instead of the default 12. Also set the queue_buffer_size to something like 5000. Without that change, the in-memory buffer fills up and starts discarding jobs that haven't been picked up yet. I learned this the hard way when I lost about three months of accumulated task data after a container restart. There's no recovery for that. Make sure you enable persistent queue storage by setting persistent_queue to true and pointing it at a mounted volume. Another counter-intuitive point: the documentation strongly recommends against using the built-in retry logic. It looks convenient, but the retry algorithm uses an exponential backoff that doesn't account for queue congestion. In practice this means that when your queue backs up, retries pile up on top of the already-stalled original jobs and make the situation much worse. Instead, disable retries in the config and handle them yourself with a separate dead-letter queue. It's more work upfront but saves you from the worst edge cases. I've been running my production setup without retries for over a year now and I've never missed it.
Get the Full Details

Common Pitfalls
The task definition format is stricter than it needs to be. Each task file needs to declare its dependencies explicitly even if there aren't any. I've seen at least a dozen support tickets about "tasks not running" that turned out to be people who forgot to add the dependency array to their task metadata. It's a one-line addition and it takes ten seconds to fix. But figuring out that the error was actually a missing metadata field rather than a runtime bug took me about two days of debugging. The monitoring interface is functional but not great. It gives you basic queue depth and worker status. What it doesn't tell you is how long individual jobs are spending waiting versus actually executing. If you need that level of visibility, you have to layer in something like Prometheus metrics export. The library does support it but the setup requires adding a few lines to your task definitions and configuring an external scraping endpoint. Worth doing if you're running more than ten concurrent jobs regularly.
Performance Tuning for Real Workloads
When I moved from testing to actual production use, the throughput numbers doubled after I made two changes. First, I switched the serialization format from JSON to MessagePack. The docs mention this as an optional optimization but the difference is massive because every task payload gets serialized and deserialized multiple times as it moves through the pipeline. Second, I increased the maximum concurrency per worker from 4 to 16. The default was clearly set conservatively, probably to avoid overwhelming beginners' machines during testing. For any serious workload you want it higher. There is a hard limit though. I found that beyond about 32 concurrent operations per worker, the overhead of context switching between jobs starts to eat into the gains. Your actual optimal number depends on whether your tasks are CPU-bound or I/O-bound. CPU-bound tasks plateau earlier, usually around 16 to 20 concurrent operations. I/O-bound tasks can often handle 32 to 48 before hitting diminishing returns. The only way to know for sure is to run a load test with your actual workload, not some synthetic benchmark. The project is still actively maintained and the team responds reasonably quickly to issues on GitHub. If you hit a problem that isn't covered here, posting a minimal reproducible example tends to get answers within a day or two. Just make sure you include your config file and the exact task definitions. Most reports that don't include those get closed with a request for more information and then they fall through the cracks.