Getting Valkyries to Actually Work in Production

I spent about three weeks last year trying to get Valkyries (Palantir's open-source orchestration library) to handle a pipeline that was barely processing 500 rows per second. The documentation makes it look dead simple. It's not, once your data starts getting messy or your downstream dependencies decide to timeout on you. Here's how it actually works when you're not following the tutorial walkthrough.

What Valkyries Actually Is

Valkyries is a Python-based workflow orchestration and data processing library. It sits somewhere between a lightweight Airflow alternative and a more opinionated version of Prefect. The core idea is that you define tasks, wire them into workflows, and let the engine handle execution, retries, and state management. It's not a full distributed computing platform like Spark, and pretending it is will burn you fast. People often confuse it with Palantir's commercial product (also called Valkyrie). They share DNA but are different things. The open-source Valkyries library is what you install via pip. The commercial one is a whole separate beast.

Setting It Up Without Wasting Two Days

The standard install is straightforward: pip install valkyries That gets you the core library. But here's where most people slip up — you also need the right backend configuration if you want anything beyond single-machine execution. By default, Valkyries runs everything in-process. That's fine for development. It's not fine if you need parallel task execution across multiple machines or even multiple CPU cores on the same box.

Get the Full Details

Valkyries survive Aces rally to take 2-0 lead in WNBA playoffs
Valkyries survive Aces rally to take 2-0 lead in WNBA playoffs

For anything production, you'll want to configure a result backend. SQLite works for small setups. Redis or PostgreSQL if you're doing anything real. I've seen people try to run multi-worker deployments against SQLite and watch their pipelines silently corrupt themselves because of write conflicts. Don't do that.

Defining Tasks the Way It Actually Works

The decorator-based approach looks clean in examples: @valkyries.task
def fetch_data(source: str) -> dict: But the reality is that task signatures need to be JSON-serializable if you're going beyond the default in-memory execution. Pass a pandas DataFrame around and you'll hit serialization errors when the engine tries to cache results or pass data between workers. I learned this the hard way on a project where I was moving 40GB of processed parquet files through the pipeline and kept wondering why tasks were failing with pickling errors that made no sense at first glance.

The workaround is simple but easy to miss: wrap your DataFrame operations in functions that return native Python types or file paths, then have downstream tasks read from those paths. It adds a layer of indirection but saves you from fighting the serialization layer constantly.

Valkyries once again named most valuable franchise in the WNBA as ...
Valkyries once again named most valuable franchise in the WNBA as ...

Common Pitfalls That Nobody Mentions

First, error handling is not automatic. Unlike some orchestration tools that give you decent retry logic out of the box, Valkyries expects you to define your own exception handling inside tasks or use its built-in retry decorators explicitly. If you skip this, one flaky API call and your entire workflow dies with no recovery. Second, dependency management is strict. If task B depends on task A, and task A returns None because it failed silently (no exception thrown), task B will still attempt to run with None as its input. This caused me to spend an afternoon debugging what I thought was a data quality issue before realizing the upstream task had returned None instead of raising an error. The fix was adding explicit validation at the start of each downstream task. Third, the logging is basic. You get stdout by default. If you want structured logs, distributed tracing, or integration with something like Datadog, you're writing that yourself. I ended up wrapping the core logger with a custom handler that pushed structured JSON to our monitoring stack. Took about half a day but was necessary for any meaningful debugging at scale.

When Valkyries Is the Wrong Tool

It handles moderate-scale workflows well — hundreds to low thousands of tasks per day, not massive batch jobs. If you're processing terabytes of data or need microsecond-level latency between task executions, look elsewhere. Prefect or Dagster will serve you better for complex dependency graphs with heavy parallelism. Airflow if your team already knows it and you need the massive ecosystem of prebuilt operators. Valkyries shines when you want something lightweight that you can embed directly in your Python application without spinning up a separate orchestrator service. It's good for microservices that need their own workflow logic, for ML pipelines that run inside a single deployment, and for prototyping before committing to a heavier orchestration layer.

A Real-World Edge Case

Here's a specific problem I ran into that the docs don't cover: time-sensitive task cascading. We had a workflow where task A fetched data from an API with a 30-second timeout, task B processed it, and task C uploaded results. The API would occasionally take 28 seconds, leaving barely any time for B and C. The workflow would complete but the upload would fail intermittently because the total wall-clock time exceeded our external service's timeout. The solution was to restructure the workflow so that task B ran in parallel with the tail end of task A's response handling, and to set individual task timeouts rather than relying on the overall workflow timeout. Valkyries supports per-task timeout configuration, but it's not obvious from the quickstart. You set it in the task decorator's timeout parameter. That changed our success rate from about 94% to 99.7%.

Valkyries reach $100M, first women's team — IJR News
Valkyries reach $100M, first women's team — IJR News

Where to Get It

The library is available on PyPI and GitHub under Palantir's organization. The documentation lives at the usual place and covers the basics well. The gaps are in the operational details — the things that only matter when something breaks at 2 AM.