What Zizikis Actually Is and Why It Exists

Zizikis is a lightweight orchestration framework for managing batch data pipelines that need to run across distributed workers without requiring a full Kafka or Airflow deployment. People tend to overcomplicate this space. Most teams don't need a heavyweight scheduler. They need something that can retry failed tasks, track job state, and hand off work to worker nodes without constant manual intervention. I built a system on top of Zizikis roughly two years ago for a log aggregation pipeline. We were processing about 400 GB of structured logs per day, routing them through five transformation stages before landing in a columnar store. The pipeline kept falling apart on edge cases—specifically, when one worker node became slower than the others and the main scheduler's timeout logic kicked in, clearing out entire partitions before the straggler finished. That was a real headache.

Getting Started with Zizikis

Download the current release from the official repository and run the installation script. The binary drops into /opt/zizikis by default, and the configuration files live in /etc/zizikis/. You'll need a coordinator node and at least two worker nodes to make the system do anything useful. The coordinator handles job scheduling and state persistence. Workers pull tasks from the queue and execute them. Config file example for the coordinator: node_type = coordinator
bind_address = 10.0.1.5
port = 8920
state_store = /var/lib/zizikis/state
max_concurrent_jobs = 16
task_timeout = 1800

Worker config looks similar but sets node_type = worker and includes a coordinator_url pointing at the coordinator node. Keep the versions identical across all nodes. Mismatched versions caused me exactly one production incident last year. It was not a fun debugging session involving stale task states and phantom completions.

How the Execution Model Works

Zizikis uses a DAG-based task graph. You define your pipeline as a series of nodes with explicit dependencies. When you submit a job, the coordinator resolves the graph, determines which tasks are ready to run, and pushes them to available workers. Tasks complete asynchronously. The coordinator tracks progress through a durable state store, which defaults to an embedded RocksDB instance. Here's the counter-intuitive part most people miss: Zizikis does not use a FIFO queue for task distribution. It uses a priority queue weighted by dependency depth. Tasks higher in the DAG get scheduled first, which prevents a scenario where workers idle waiting for upstream tasks to finish. This behavior is controlled by the scheduling_policy parameter in the config. Setting it to depth_first gives you better throughput on wide graphs. Setting it to breadth_first is better for narrow, deep pipelines. I learned this the hard way after running a pipeline with 80 tasks across 12 dependency layers. With the default breadth-first policy, my workers sat idle 40% of the time because deep leaf tasks had no chance of being picked up before shallow independent tasks filled all the slots. Switching to depth_first cut wall-clock execution time from about 47 minutes down to 19 minutes on the same hardware.

Common Pitfalls and What to Watch For

The state store is the single point of failure in a basic Zizikis setup. If the coordinator goes down and the RocksDB directory gets corrupted, you lose all in-flight job state. There is no automatic failover in the community edition. Some teams run a second coordinator as a hot standby with a shared state volume, but that introduces consistency questions that Zizikis doesn't fully solve for you. Another issue I ran into involves large payload serialization. When you pass complex objects between tasks using Zizikis's built-in result bus, everything gets serialized to JSON by default. This works fine for simple data. It breaks when you're moving nested structures with circular references or non-serializable types. I hit this when trying to pass parsed schema objects between transformation steps. The workaround was straightforward—dump the objects to temporary Parquet files and pass file paths through the result bus instead. Serialization time went from 340ms per task to under 12ms, and I stopped getting timeout errors from the result bus. Memory management on workers is another area that deserves attention. Each worker runs a small internal buffer for task payloads. If you set max_concurrent_jobs too high on the coordinator while workers have limited RAM, you'll see out-of-memory kills during peak load. The sweet spot for most setups I've seen is somewhere between 8 and 16 concurrent jobs per worker, depending on your payload sizes. Profile your actual task memory usage before tuning this.

A Practical Example

Let's say you want to build a simple ETL pipeline with three stages: extract raw CSV files from an S3 bucket, transform the data by parsing dates and normalizing strings, and load the results into a PostgreSQL table. Here's how that looks in practice. Create a file called pipeline.yaml: job_name: daily_etl
tasks:
- name: extract
type: script
command: python3 scripts/extract.py
config:
source_bucket: s3://my-bucket/raw/
output_dir: /tmp/zizikis_stage/extracted
- name: transform
type: script
command: python3 scripts/transform.py
depends_on: [extract]
config:
input_dir: /tmp/zizikis_stage/extracted
output_dir: /tmp/zizikis_stage/transformed
- name: load
type: script
command: python3 scripts/load.py
depends_on: [transform]
config:
database: postgresql://dbhost:5432/mydb
table: daily_records
source_dir: /tmp/zizikis_stage/transformed

Submit the job: zizikis submit pipeline.yaml The coordinator resolves the dependency graph, schedules extract first, then transform when extract completes, then load when transform completes. Each stage runs on whichever worker picks it up. If any task fails, Zizikis retries it up to three times by default before marking the entire job as failed. You can adjust retry_count and retry_delay in the task config.

This isn't a perfect system. It doesn't handle micro-batch streaming, it has no native support for GPU-accelerated tasks, and the monitoring dashboard that ships with it is functional but bare-bones. For a team doing simple batch workflows with moderate complexity, it does the job without the overhead of more established tools. For anything larger or more complex, you'd likely outgrow it within six months.