A Practical Introduction to Using Colombia for Workflow Automation
Colombia is a Python package for parameterized scientific workflows. It sits between your code and the execution environment and handles configuration, job submission, and output collection in a way that stops you from writing fifty shell scripts that barely work on one machine but fail everywhere else. If you've ever spent an afternoon debugging why a simulation ran fine on your laptop but produced garbage on the cluster, this is probably worth your time. The package is available on PyPI, and you can install it with pip.
The basic workflow setup
Colombia works around the concept of a workflow definition. You write a YAML file that describes your tasks, the parameters you want to sweep, and how jobs should run. Then you call a small amount of Python to submit and collect results. That's really the whole thing. Here is a minimal example of what a config looks like: A typical config has a workflow section, a data section pointing to where inputs live, a parameters section for sweeps, and a jobs section for execution details. You don't need to understand every field before starting. Just get one task running end to end, then expand from there.
Installing and running your first job
Pip install Colombia brings in the command line tools and the Python API. After installation, you point it at a workflow directory and you're set. The simplest command is just to dry-run your workflow to check that everything resolves. This catches most configuration errors before they become headaches later. After that, you submit with the run command. Colombia handles creating the execution directories, writing parameter files, and launching jobs based on your configuration. One thing beginners often miss: Colombia separates the concept of a task from the concept of a job. A task is what you want to do. A job is a single instance of that task with specific parameters. This distinction matters when you start managing collections of parameter sweeps. You'll be glad you understood it early.
Get the Full Details

Parameter sweeps and collections
This is where Colombia actually saves you time. Instead of manually varying parameters across dozens of runs, you define collections in your config. Each collection expands into separate jobs automatically. For example, if you're sweeping two parameters across three values each, you get nine jobs. Colombia creates the directory structure, writes the configuration for each job, and submits them according to your backend settings. You don't touch individual jobs unless something goes wrong. I once spent two days debugging what I thought was a bug in my simulation code. Turns out I had accidentally defined overlapping parameter ranges in two different collections, so jobs were overwriting each other's output directories. Colombia does not prevent this by default. You have to structure your collections carefully or set collision rules in your config. Always run a dry run and check the job list before submitting anything to a cluster.
Backend configuration and execution
Colombia supports multiple execution backends. The default is local execution, which is fine for testing. For actual compute, you typically configure a cluster backend or a queueing system. The backend configuration lives in your workflow config and controls things like job queue names, resource requests, and submission commands. Getting this right depends heavily on your specific HPC environment. Documentation covers the common cases, but you will need to adapt it to your system. One counter-intuitive thing: Colombia does not manage dependencies between tasks the way some workflow engines do. Each task collection runs independently. If your analysis depends on simulation output, you handle that sequencing yourself, usually by running the simulation collection first and then the analysis collection. This is simpler than you might expect, but it catches people off guard when they assume automatic dependency resolution exists.
Common pitfalls and how to avoid them
Colombia is not a silver bullet. It has real limitations that matter in practice. The configuration format is YAML, and while that is convenient, complex workflows can produce configs that are hundreds of lines long and hard to debug. I've seen people maintain workflow configs as sprawling monoliths. Break your workflows into smaller pieces and compose them instead. Colombia supports this pattern and it makes debugging significantly easier. Another limitation: output management assumes you define clear output paths upfront. If your simulation produces output files with dynamic names, Colombia won't track them automatically. You either name outputs predictably or write custom collection logic. This is not a dealbreaker, but it is something to plan for.

For workflows that require tight coupling between steps, conditional execution, or complex branching logic, Colombia is not the right tool. You would be better served by a full workflow orchestration system. Colombia excels at parameter sweeps and batch processing, not at building directed acyclic graphs of dependent tasks.
When Colombia actually works well
The sweet spot is reproducing numerical studies where you need consistent parameter variation, clean output organization, and the ability to rerun everything from scratch. If your research involves running the same model across many parameter combinations and you want the results to be structured and reproducible, this is where it shines. It also helps with portability. Once your workflow is configured, you can move it between your laptop, a lab server, and a cluster with minimal changes. The backend abstraction does most of the heavy lifting there. I recommend starting small. Get one parameter sweep working on your local machine. Make sure the output structure matches your expectations. Then expand to more parameters and eventually a cluster backend. Don't try to configure everything at once. The package works fine, but the configuration complexity grows faster than most people anticipate, and you'll learn more by adding features incrementally than by trying to get it right in one shot.