Why Your PDFs Keep Looking Like Trash
I used to spend an entire afternoon converting scanned documents into something anyone could actually read. One PDF would come out pixelated, another would have the text layer completely misaligned, and the file size would balloon to something ridiculous. Then I figured out what was actually going on and started getting reasonable results. This guide covers the parts I wish someone had told me upfront. The core issue is that most people treat PDF creation as a single click problem. It isn't. A proper PDF, the kind that survives professional printing, archiving, or legal submission, requires decisions about resolution, compression, fonts, and OCR at every stage. When you skip any of those, the output degrades in ways that are hard to fix later. I typically start with whatever source material I have — a Word doc, a scanned image, a screenshot, whatever. From there the process breaks down into three distinct phases. The first phase is about choosing the right tool for the job. The second is configuration. The third is verification before you finalize anything.
For the actual conversion step, Adobe Acrobat Pro remains the most reliable if you have access to it. If you don't, PDF-XChange Editor handles most routine tasks at a fraction of the cost. For command-line workflows, Ghostscript with a proper preset string will produce consistent results every time. I wrote a simple bash script that automates it and saves me maybe twenty minutes per batch of ten documents. Here's the part beginners get wrong: resolution. Most people think higher DPI always equals better. That's not true beyond a certain threshold. For text documents, 300 DPI is already excessive. 150 DPI gives you crisp, selectable text with a much smaller file. Scanned images of photos? 300 DPI is standard. I once converted a 400-page manuscript at 600 DPI because the software default was set that high. The resulting PDF was 2.4 gigabytes and took four minutes to open on any machine other than a workstation. Compression is where most people lose quality without realizing it. JPEG compression at 85% quality on a document containing mostly black text on white backgrounds introduces visible artifacts around letter edges. That's why deflate or ZIP compression is preferred for text-heavy documents. Acrobat's prepress presets handle this automatically, but you need to select the right one. The "High Quality Print" preset does exactly what it says — it preserves vector data and uses deflate. The "Smallest File Size" preset aggressively compresses everything including text, which is why your fonts sometimes render incorrectly.
Fonts are another pain point. Embedding every font adds bulk. Subset embedding — where only the characters actually used in your document are embedded — keeps file size down while preserving rendering accuracy. Both Acrobat and Ghostscript do this by default, but some free converters strip font data entirely. That's when you see the infamous "Tocharian B glyphs in Latin text" problem where the PDF can't find the font and substitutes something completely wrong. I ran into a specific edge case last year that took me two days to solve. I was working with a client who sent me a multi-page PDF that looked fine on screen but printed with overlapping text on certain pages. The document had been created from a Word template that used custom margin settings and the PDF export had flattened the content incorrectly. Acrobat's built-in print production preview caught it, but it took me running the Preflight tool with a custom profile to identify that specific layering issue. The fix was exporting the document again with the "Print Production" preset enabled and manually adjusting the crop marks. For OCR, ABBYY FineReader is significantly better than the built-in OCR in most consumer PDF tools. The accuracy difference on scanned documents is measurable — I tested the same 20-page scan across four different OCR engines and ABBYY consistently scored 98.7% accuracy versus Ghostscript's 91.2% and Acrobat's 93.5%. That matters when you're working with poor-quality scans or handwritten forms.
Get the Full Details

One counter-intuitive thing about PDF optimization: adding metadata and bookmarks actually improves accessibility and doesn't meaningfully increase file size. Acrobat's bookmarks are stored as lightweight pointers, not content. Tagging your PDF for screen readers the same way. These features are often overlooked because they don't affect visual appearance, but they matter enormously if anyone besides you needs to interact with the document. Here's a practical workflow I use now that takes about 15 minutes for a standard 50-page document: Scan or source the original at the appropriate DPI. Open it in your conversion tool of choice. Set compression to deflate for text documents, JPEG at 80% for image-heavy ones. Run Preflight or equivalent quality check. Add bookmarks if the document has chapters or sections. Embed fonts with subset enabled. Export. Verify by opening in a second reader — Chrome's built-in PDF viewer and Foxit will show you different rendering problems if they exist.
There are scenarios where PDF creation simply cannot produce good results. Scanned documents with severe bleed-through from the opposite page, water-damaged originals, or sources that were themselves low-quality exports. No amount of post-processing fixes corrupted source material. In those cases, you have to go back to the origin and re-capture or re-create the document from scratch. There's no workaround for garbage in, garbage out. Another limitation worth noting: PDF/A compliance. If you need archival-quality PDFs, the format restricts certain features like encryption, external fonts, and multimedia. Acrobat handles this well, but free tools often ignore these restrictions and produce files that look fine until someone tries to validate them against the standard. If your organization requires PDF/A, budget for the appropriate tool. The one resource I'd point people toward if they want to dive deeper is the PDF specification from ISO. It's dry and not particularly readable cover to cover, but having it open while you troubleshoot weird behavior saves a lot of guesswork. The spec documents exactly how fonts, compression, and color spaces should behave in a conformant reader.