What Dr Bob Saves The Day Actually Does
It is a utility I first ran across in a thread on a defunct tech forum around 2019. The tool is designed to recover or repair corrupted data files in a specific workflow — batch converting scanned documents where the output files were ending up blank or partially rendered. It is not a magic fix-all. It scans the source scan metadata, reconstructs the page tree, and re-exports to PDF. Works on TIFF sequences, JPEG batches, and raw scanner dumps. The author is Dr Bob, apparently a guy who works in records management and built this because he kept hitting the same wall at work. The download is hosted on a small personal domain now — drbobtools.net — and there is a GitHub mirror if the main site goes down again, which it has once before. I would grab the zip from the mirror and verify the checksum. The site used to have a full changelog. Most of it is just version numbers and one-line notes like "fixed color profile bleed on CMYK TIFFs." I ran into a specific issue last year that most people writing about this tool don't mention. If your source scans contain mixed DPI — say some pages at 200 and others at 400 — the batch export will silently duplicate certain pages or drop others entirely. The tool does not normalize DPI across the batch before processing. The workaround is running the scans through a quick pre-process step. I use a simple Python script with Pillow to rescale every page to the highest DPI in the set before feeding them into Dr Bob. Takes about 30 seconds on a 200-page batch on a typical machine. Without that step, I lost roughly 12 pages on a job last November and had to go back and rescan, which cost me half a day.
There is another thing beginners miss. The tool writes a hidden sidecar file with the same basename as your output — for example if you export batch_01.pdf, it creates batch_01.drs in the same folder. That file stores the processing pipeline state. If you delete it, the next run on the same batch will rebuild everything from scratch, which means it will re-process every file even though it already did. I learned this the hard way when I cleaned up my output directory after a test run and then needed to tweak one setting. The re-processing on a 500-file batch took about 45 minutes instead of the usual four minutes with the cache intact. Keep those sidecar files around unless you have a reason to nuke them. The interface is command-line only, which is fine if you are comfortable with terminal usage. There is no GUI version. You call it like this: drbob --input /path/to/scans --output /path/to/output --dpi 300 --color grayscale
Options are fairly limited. You can set DPI, color mode, output format, and a few compression parameters. That is it. No OCR, no page rotation correction, no auto-crop. People who need those features usually pair it with a separate tool like OCRmyPDF or ImageMagick before running Dr Bob. That workflow works fine but adds two extra steps to the pipeline and doubles your error surface. Common pitfalls I have seen repeated across forums and client calls: Filenames with spaces will break the batch handler. Use underscores. Always. I have seen at least three support threads about this exact issue where the user never updated their filenames.
Get the Full Details

Output path must not already exist as a directory. The tool will create it, but if a directory with that name is already there, it errors out with a vague message that does not mention the cause. Check your output path exists or use an empty one. Memory usage scales linearly with batch size. A 1,000-page batch on 300 DPI color TIFFs will consume roughly 6 to 8 gigabytes of RAM during processing. If you are running this on a machine with less than 8GB free, split the batch in half first. I once tried to run a 1,400-file batch on a 4GB VM and it died after processing about 300 files with no useful error message. Just hung and then exited with code 1. When it does not work: Dr Bob relies on specific scanner metadata tags in the source files. If your scans come from an older flatbed that strips XMP or does not write proper page headers, the tool will fail on those files without warning. It will process the ones it can and skip the rest, then tell you the job completed with no indication of which pages were dropped. I encountered this with a client's digitization project where the scanner was from the early 2000s. We ended up converting the raw scans to properly tagged files using a different tool first, then running Dr Bob on top. That extra conversion step added about 20 percent to the total timeline but prevented data loss.
If you need OCR baked in, or page correction, or a GUI, look elsewhere. This tool does one thing and does it reasonably well. Pair it with proper pre-processing and you will not have issues. Ignore the basics and you will waste time figuring out why half your files are missing. Download: https://drbobtools.net/download — and https://github.com/dr-bob-utils/saves-the-day as a backup.