Getting Your Project Working Without Losing Your Mind
I spent about three weeks fighting with this last year before I actually figured out what was going wrong. The documentation is sparse, the community is basically nonexistent outside of a couple Discord servers, and half the tutorials you find online are copy-pasted from each other and flat-out wrong. I'm writing this because someone needs to explain it plainly. It's a modular orchestration framework built around event-driven state machines. That's the textbook definition. What it actually is in practice is a system that lets you define complex multi-step workflows as configurable JSON schemas, then execute them across distributed nodes with at-least-once delivery guarantees and built-in retry logic. The name comes from an inside joke in the original dev team that somehow stuck on GitHub. The core concept is the thunder engine. Each thunder instance tracks state transitions independently. You configure workflows by declaring states, transitions, conditions, and side effects in a manifest file. When you run the orchestrator, it reads that manifest, initializes the event bus, and starts consuming from the configured transport layer. Most people default to Kafka or RabbitMQ for the transport, though Redis pub/sub works fine for smaller setups.
Here's how I actually set one up. First, you install the CLI tools and the runtime separately. The runtime goes on your infrastructure nodes, the CLI stays on your dev machine. Then you create a project directory, run the init command, and it generates a skeleton manifest and a sample workflow. The important part most people skip is configuring the state store correctly. By default it uses SQLite, which is fine for development but completely falls apart when you hit more than about fifty concurrent workflows. I switched mine to PostgreSQL early on and haven't looked back. The execution model is declarative. You describe what should happen, not the imperative steps to make it happen. The engine compiles your manifest into a DAG, validates it, and then schedules it across available workers. Workers pull tasks from a shared queue, execute them, report results back, and the engine transitions states accordingly. If a worker dies mid-execution, the task gets requeued and another worker picks it up. That's the at-least-once guarantee at work. I ran into a real problem with nested workflows about six months ago. I was trying to compose a parent workflow that called three child workflows sequentially, with conditional branching between them. The framework supports this natively, but there's a known issue where the state checkpointing doesn't properly serialize nested context objects. The parent would complete successfully but the child workflow IDs would get lost in the state store, making it impossible to query or debug later. I spent two days on this before I found the workaround.
The fix was to disable automatic nested context serialization and instead pass workflow IDs explicitly through the input payload. You do this by setting the nested_context_mode parameter to manual in your manifest, then flattening your child references into the root context object. It's ugly but it works reliably. There's an open pull request about this on GitHub that's been sitting for four months as of my last check. Let me address something beginners usually get wrong. People think the manifest syntax is the hard part. It's not. The hard part is understanding idempotency. Every side effect in your workflow should be idempotent, meaning running it twice produces the same result as running it once. The engine will retry failed tasks, and sometimes it'll process the same task twice during a state recovery. If your workflow triggers an external API call or sends an email or writes to a database without idempotency guards, you will have duplicate records and angry customers. I learned this the hard way when a production workflow duplicated about forty payment processing calls during a recovery window. Another thing nobody mentions in the docs: the default timeout values are far too long for most use cases. A standard workflow step defaults to a thirty-minute timeout. In production, that means a hung task sits there burning resources for half an hour before the engine retries it. I've seen entire clusters slow to a crawl because a single misbehaving worker wasn't killed fast enough. Set your timeouts explicitly, and I'd recommend starting with something like two minutes for external API calls and five minutes for database operations. You can always increase them if you hit edge cases.
Get the Full Details

For monitoring, the framework ships with basic Prometheus metrics out of the box. You get workflow duration histograms, state transition counts, error rates per step, and queue depth gauges. That's enough for a basic Grafana dashboard but not much more. The plugin system for custom observability is decent but underdocumented. I wrote a custom OpenTelemetry exporter that ties workflow events to trace spans, which made debugging significantly easier once we started running production loads. Scaling is where this tool actually shines. I ran a test deploying the orchestrator across twelve workers handling roughly two hundred concurrent workflows. The system held steady with average task latency around 80 milliseconds and zero state corruption incidents over a forty-eight-hour stress test. The bottleneck wasn't the engine itself, it was the PostgreSQL connection pool. I had to tune max_connections and enable connection multiplexing to keep the queue moving smoothly. Worth noting that the connection pool settings aren't exposed in the main config file. You have to pass them as environment variables, which the docs mention exactly once in a footnote. There are legitimate downsides you need to know about before committing to this. The learning curve is steep compared to alternatives like temporal.io or AWS Step Functions. The documentation has significant gaps, especially around advanced features like compensating transactions and distributed locking. The community is small, so when something breaks, you're often on your own unless you can reproduce it and file a detailed issue. Performance-wise, the framework introduces noticeable overhead for very high-throughput scenarios. If you're processing thousands of workflows per second, you'll want to look elsewhere.
For smaller to medium projects, though, it's quite capable. My team uses it for internal ETL pipelines and event-driven microservice orchestration. We run maybe sixty to eighty workflows per day across all environments combined. It handles that comfortably. The declarative approach means our manifests are version-controlled and reviewable, which is valuable for compliance work. If you want to get started, grab the latest release from the official GitHub repository. The project has a MIT license so you're free to modify and distribute. You'll need a Go runtime installed for the CLI, Node.js for some of the scaffolding tools, and whichever message broker you choose for transport. Docker images are available on Docker Hub if you prefer containerized deployments. Start simple. Build a single workflow with two states and one side effect. Get it running locally with the SQLite backend. Once you understand the flow, gradually add complexity. Don't try to migrate your entire infrastructure into it on day one. I wish someone had told me that before I wasted three days debugging a deployment that was fundamentally flawed from the start.