Getting Division Right in Batch Processing

I spent three days debugging a script that was supposed to split a dataset into equal chunks. The math looked fine on paper. The output was completely wrong. That's when I actually dug into how Just Divide handles edge cases instead of just copy-pasting a Stack Overflow snippet. Just Divide is a Python utility for cleanly splitting iterables, ranges, or arrays into N roughly equal parts without the usual off-by-one headaches. It's not a replacement for numpy or pandas. It's something smaller and more focused. I use it in data engineering pipelines where I need to partition workloads across workers and can't afford the overhead of loading a full dataframe just to divide indices.

How Just Divide actually works

At its core, the library takes a sequence length and a chunk count, then distributes the remainder intelligently so no chunk ends up larger than necessary. Unlike a naive floor division approach, it doesn't leave a sloppy tail. The first few chunks carry the extra elements, and the rest settle into a clean base size. Here's the basic usage: from just_divide import just_divide
chunks = list(just_divide(100, 7))
Returns something like [15, 15, 14, 14, 14, 14, 14]
The return value is a list of slice boundaries or indices depending on which function variant you call. just_divide() gives you start-stop pairs directly. divide_into() can wrap an existing iterable and yield the actual sliced chunks. I prefer the boundary version because it keeps memory flat. You map over the pairs and hand each worker its slice without building intermediate lists.

The edge case that burned me

Last year I was processing log files from a distributed system. There were 3,847 files and I needed to assign them to 64 workers. Mathematically that's clean: 3847 divided by 64 is 60 with a remainder of 7. But the worker assignment logic was built on a simple range loop that assumed even distribution. Files at the tail end kept getting skipped because the chunk boundaries drifted due to integer division rounding errors in the surrounding codebase. The jobs would complete early and report success while seven files sat unprocessed in the queue. Took me four hours to trace because the logs looked normal. The files just weren't there. The fix was swapping the division logic out entirely for Just Divide, passing the exact file count and worker count, and using the returned boundaries as strict inclusion ranges. No more drift. The seven leftover files went into the first seven workers and everything landed exactly once. I wish I'd caught that sooner. The library handles this automatically because it computes boundaries using ceiling division on the remainder distribution rather than floor division on the whole quotient. If you're doing this from scratch, the algorithm is straightforward enough to implement yourself, but there are enough variants—descending chunks, ascending, even distribution, capped minimums—that maintaining your own version tends to introduce bugs over time. That's why I just pull the dependency in.

Get the Full Details

🕹️ Play Just Divide Game: Free Online Educational Math Practice Division Video Game for Students
🕹️ Play Just Divide Game: Free Online Educational Math Practice Division Video Game for Students

Installation and quick setup

You can grab it from PyPI. The command is standard: pip install just-divide There's also a GitHub repository if you want to look at the source or open issues. The package is lightweight. No dependencies beyond Python 3.8. I've run it in environments with no internet access by downloading the wheel from pypi.org manually and installing with pip install --no-index --find-links=./whls just-divide. Works fine. The wheel is under 15KB.

Things beginners get wrong

One common mistake is treating Just Divide as a chunking tool for arbitrary data. It isn't. It operates on integer counts and produces index boundaries. If you pass it an iterable and expect it to handle the slicing, you need the divide_into() function or you do the slicing yourself using the boundaries. Passing a string to just_divide() won't work because it expects a length parameter, not the data itself. Another thing: the function doesn't validate that your chunk count is positive. Pass zero and you'll get a division by zero error in your own code downstream, not a helpful exception from the library. I wrap it in a guard clause in production code. It saves you from staring at a traceback for ten minutes wondering why your pipeline crashed on an empty config. There's also a behavior quirk worth knowing. When the number of chunks exceeds the length of the sequence, Just Divide returns single-element chunks for the first N items and empty lists for the rest. This is technically correct but can confuse code that assumes every chunk contains at least one element. If your use case requires non-empty chunks regardless of count, you need to filter the output or raise an error yourself.

When it falls apart

Just Divide doesn't help you with weighted distribution. If you need chunks sized according to some property of the data—like balancing load by record size or file age—you're out of luck. The library is purely arithmetic. I've seen people try to hack around this by pre-aggregating metadata and then dividing the aggregated counts, which works in some cases but breaks as soon as the underlying distribution changes. For weighted partitioning, I use a custom function that does a greedy bin-packing approach. It's slower but accurate. It also doesn't provide parallel execution. Some people assume it does because the name sounds like it might. It doesn't. You still need to dispatch the chunks yourself using multiprocessing, concurrent.futures, or whatever scheduler you're working with. That's actually fine. It keeps the library small and lets you pair it with whatever parallelism model your project already uses.

Just Divide - Play online at Coolmath Games
Just Divide - Play online at Coolmath Games

Practical example

Here's how I typically structure a real pipeline that uses it: from just_divide import just_divide
import concurrent.futures

total_records = 50000
num_workers = 16

boundaries = list(just_divide(total_records, num_workers))

def process_chunk(start, stop):
chunk = fetch_records(start, stop)
return transform(chunk)

with concurrent.futures.ThreadPoolExecutor(max_workers=num_workers) as executor:
results = list(executor.map(lambda b: process_chunk(*b), boundaries))
This pattern keeps memory usage predictable. The boundaries are computed once, each worker gets a tight slice, and there's no global state shared between them. I've seen the same logic run on datasets from 1,000 records up to 2 million without modification. The bottleneck is always the data fetch step, never the division.

The library has a few other functions worth knowing about. balanced_divide() lets you specify a minimum chunk size and will refuse to create chunks smaller than that, redistributing the excess to larger bins instead. reverse_divide() puts the remainder on the last chunks rather than the first. Both are useful depending on whether your downstream system processes chunks in order or tolerates variable sizes better at one end of the spectrum than the other. If you're doing anything that involves splitting work into equal portions in Python, Just Divide saves you from writing and rewriting the same division logic. It's one of those tools that doesn't look impressive until you've spent the time fixing bugs in your own version. Then it looks like everything else should be this simple.