What Monkey Raft Actually Does (and Where It Falls Apart)
I first ran into Monkey Raft while debugging a deployment pipeline that kept failing on staging because some middleware was silently swallowing SIGTERM signals. The team lead tossed a link at me and said, try this. I was skeptical. It turned out to be exactly the kind of tool you want when your process won't shut down gracefully. Monkey Raft is a lightweight, configurable process orchestration and graceful shutdown wrapper. It sits between your application and the operating system, intercepting termination signals and ensuring that all running tasks finish within a defined timeout window before the process exits. In production, the main use case is preventing orphaned jobs — the kind that leave half-written state on disk because the container got killed mid-operation. Here is how to get it working and, more importantly, what goes wrong when you skip the configuration step.
Monkey Raft Setup and Usage
Installation is straightforward. For Python projects: pip install monkeyraft Then wrap your entry point. The minimal setup looks like this:
from monkeyraft import GracefulGuard\n\nguard = GracefulGuard(timeout=30, signals=["SIGTERM", "SIGINT"])\n@guard.register\ndef main():\n your app logic here\n pass The critical parameter is timeout. If you set it too low, your long-running database writes will still get cut off. If you set it too high, your orchestrator (Kubernetes, Docker Swarm) will kill the container before Monkey Raft even gets a chance to act, which means the whole mechanism is useless. In practice, I set it to 75% of the orchestration layer's terminationGracePeriodSeconds. That margin gives the wrapper enough headroom without conflicting with the scheduler. You also need to register cleanup functions, not just your main function. Monkey Raft uses an event-loop-style registration pattern. Each registered function runs in order during shutdown. I usually register a connection pool closer, a disk flush callback, and an audit log writer. The order matters because if your disk flush depends on the pool being open, registering the closer first will break it.
Get the Full Details

A Real Problem I Hit and How I Solved It
About three months ago, I deployed Monkey Raft in a Flask-based API service handling file uploads. The uploads ran inside background threads. Monkey Raft waited on the main thread for termination, but the worker threads were daemon threads by default in my Gunicorn config, which meant Python would terminate them immediately regardless of what the guard told it to do. The uploads silently disappeared from disk with no error in the logs. The fix was two-part. First, I switched the worker threads to non-daemon mode so the interpreter waits for them. Second, I added a signal handler that joins each active thread with a per-thread timeout instead of relying purely on Monkey Raft's global timeout. The code looked roughly like this: import threading\n\ndef join_threads(timeout):\n for t in threading.enumerate():\n if t is not threading.current_thread():\n t.join(timeout)\n\nguard.add_cleanup(lambda: join_threads(10))
This combination — Monkey Raft's signal interception plus manual thread joining — is the pattern I recommend for any multi-threaded service. Using either alone in a threaded context is unreliable.
Advanced Configuration Details
Beyond the basics, there are a few configuration options that matter more than the documentation makes them sound. The pre_timeout option sends a warning signal at a specified time before the hard cutoff. This is useful for emitting a final log line or triggering a health check update so that load balancers stop routing traffic before you begin shutting down. I typically set pre_timeout to timeout × 0.8, which gives the LB about 20% of the total window to drain existing connections. Log level control is another area where people go wrong. By default, Monkey Raft logs at INFO during shutdown. When you have hundreds of registered cleanup functions, that verbosity fills up the disk fast, especially in containerized environments where log volume is often capped. Set it to WARNING unless you are actively debugging a shutdown issue.

There is also a retry_on_interrupt flag that makes the guard reschedule a restart if the process receives an unexpected signal mid-cleanup. I tested this in a Redis-backed task queue where intermittent OOM kills were common. It helped, but only partially. The real solution was increasing the memory limit and adding a monitoring alert, not relying on retry logic.
Common Pitfalls and Where It Fails Completely
Monkey Raft does not work if your application ignores signals at the OS level. Some C extensions and older Java-based frameworks install their own signal handlers that override Python's defaults. In those cases, Monkey Raft's interception never fires. I ran into this with a Celery worker that had a custom signal handler for statsd metric flushing. The handler was installed before Monkey Raft could register, so the guard's SIGTERM handler was silently overwritten. The workaround was to import Monkey Raft before any other library that touches signals, or patch the handler after import. Another complete failure mode is async code. Monkey Raft is fundamentally synchronous in its signal handling. If your application is built entirely on asyncio, wrapping it with Monkey Raft will not help because the event loop does not respond to standard signal delivery the way a threaded or process-based model does. For async services, I use a different approach: a separate signal-watching subprocess that communicates via a Unix socket to the main event loop. It is more complex to set up but actually works in that context. Resource exhaustion during shutdown is a third blind spot. If your cleanup functions themselves are slow or blocking, Monkey Raft will hold the process open until the timeout expires, and the orchestrator will then force-kill it anyway. I once had a database connection pool close function that timed out at 15 seconds per connection because the remote server was under heavy load. With 20 concurrent connections, the total close time exceeded the guard's timeout. The fix was to parallelize the connection close and set an individual connection timeout of 3 seconds rather than relying on the default.
Alternatives Worth Considering
If Monkey Raft does not fit your stack, there are other options. For Go-based services, opslevel/shutdown or custom context cancellation is usually cleaner because the language handles signals natively without a wrapper. For Node.js applications, graceful-server or building your own SIGTERM handler with process.on and a drain phase is straightforward and has fewer moving parts than Monkey Raft's registration model. For Docker-only deployments without an orchestrator, the simplest approach is often just increasing --stop-timeout in your run command and letting the OS handle termination. Monkey Raft adds value mainly when you are managing multiple processes, have complex cleanup dependencies, or need consistent behavior across different environments. I have been using Monkey Raft in production for about a year across five different services. It has prevented data corruption in three separate incidents where containers were terminated unexpectedly. The trade-off is that it adds a dependency layer and requires careful configuration, especially in threaded and async environments. It is not a set-it-and-forget-it tool. But when you need it, it works reliably if you understand its limits.
