The HTML-to-PDF Pipeline I Actually Use

Most people try to write content and generate PDFs at the same time. This breaks flow. I separate the two. I write in whatever editor I prefer. Then I export to HTML. Then I run a script that converts that HTML to a clean PDF. That's the pipeline. It took me about three days to set up properly. After that, generating a formatted PDF for any piece of content takes roughly 90 seconds. The upfront cost is real. The daily time savings compound quickly. I wrote my first batch of PDF content in 2019 using a combination of pandoc, a custom HTML template, and wkhtmltopdf. That was before most of the modern tooling became stable. I learned the hard way that CSS print media queries are not a suggestion. They are a requirement. Without them, your PDF looks like a screenshot of a webpage with weird padding. I spent six hours debugging header overflow on one document. The fix was adding @page { margin: 2cm; } to my stylesheet. That single block solved half the layout problems I was having.

Pdf For Content Creation Daily

The core idea behind using PDFs for daily content creation is that a PDF is a finished artifact. Word documents are drafts. A PDF says this is done. When you distribute a PDF, your reader cannot accidentally reformat your headings or break your layout. That matters more than people admit. I once sent a client a nicely styled guide in DOCX format. They opened it on a different machine. The tables collapsed. They complained. I never made that mistake again. Here is what my actual workflow looks like now. I keep a master HTML template in Google Drive. The template has the footer, the branding, the table styles, the font imports, everything. When I finish writing a piece, I duplicate the template folder. I replace the body content. I run my conversion script. The script uses Puppeteer in headless mode with a custom CSS file injected. The output is a PDF named exactly like the source file plus -v1.pdf. I push it to my shared drive. Done. I do not open any PDF editor. I do not click around menus. The script itself is about forty lines of JavaScript. It loads the HTML file, applies a print stylesheet, sets the paper size to A4 with 2cm margins, and renders. There are no complex dependencies. Just Puppeteer and Node. I run it from the terminal. The conversion time is usually between two and four seconds per document. If your HTML has heavy images, it can take longer. I optimize images before they enter the template. That is step zero.

One edge case that nearly cost me a client deadline. I was generating a 200-page PDF from a long HTML file. The first fifteen pages looked fine. Page sixteen had a massive image that pushed the rest of the content onto a new page. The pagination was broken. Every page after that was shifted by one. I spent two hours figuring out why. The problem was page-break-inside: avoid on the image container. I removed that rule. The image split across two pages. It was ugly but consistent. I accepted the tradeoff. A consistent bad layout is better than a random one. Another thing beginners miss. Font embedding. If you use a web font like Inter or Poppins, Puppeteer will download it during headless rendering. Sometimes. It depends on your internet connection and whether Chrome's font cache is populated. If the font fails to load mid-render, your PDF looks wrong and you will not know until the file is generated. I solve this by downloading the font files locally and referencing them with @font-face pointing to a local path in my template. Now the conversion is offline and reproducible. Every single time. The biggest limitation of this approach is that HTML-to-PDF conversion struggles with complex layouts. Tables with nested cells. Floating elements. Anything that relies on CSS Grid or Flexbox in a way that the print renderer does not fully support. If your content is mostly text with simple headings and images, this pipeline works great. If you need magazine-quality layouts with anchored text boxes and overlapping elements, you should look at LaTeX or a dedicated design tool like Affinity Publisher. I tried using this same HTML approach for a poster layout. It took me four hours to get it to look acceptable. I switched to a vector tool. Took twenty minutes.

Get the Full Details

Content Creation Roadmap Guide | PDF
Content Creation Roadmap Guide | PDF

If you want to start with something simpler than a custom script, there are online converters. You upload an HTML file. They give you a PDF back. These are fine for occasional use. They are not fine for daily content creation because you hit rate limits and lose control over styling. My script runs locally. I control every variable. The tradeoff is that I maintain the script. But maintenance is low. The script has not changed in eighteen months. Here is a minimal version of what my conversion script does, so you can see the actual structure before you build your own. It is simplified. I removed authentication and logging for readability. const puppeteer = require('puppeteer');
const fs = require('fs');
const path = require('path');
const inputPath = 'content/blog-post-047.html';
const outputPath = 'pdfs/blog-post-047.pdf';
const cssPath = 'styles/print.css';
const htmlContent = fs.readFileSync(inputPath, 'utf-8');
const cssContent = fs.readFileSync(cssPath, 'utf-8');
const fullHtml = htmlContent + '';
(async () => {
  const browser = await puppeteer.launch({ headless: 'new' });
  const page = await browser.newPage();
  await page.setContent(fullHtml, { waitUntil: 'networkidle0' });
  await page.pdf({
    path: outputPath,
    format: 'A4',
    margin: { top: '2cm', bottom: '2cm', left: '2cm', right: '2cm' },
    printBackground: true,
    preferCSSPageSize: true
  });
  await browser.close();
})();

The CSS file print.css is where most of the magic happens. I keep it in a separate directory. The contents are mostly reset rules and print-specific overrides. Things like body { font-family: 'Inter', sans-serif; font-size: 11pt; line-height: 1.6; color: #1a1a1a; }. Simple. Predictable. This is the file I tweak when the PDF does not look right, not the JavaScript. The JS is boilerplate. The CSS is craft. I also include a pagination script that runs after the PDF is generated. It checks the page count. If the document is under ten pages, it adds a note in the console. Long documents get flagged differently. I use this for quality control. Two percent of my conversions have layout issues that slip through. The pagination check catches about half of them before I send anything out. For download purposes, you do not need to build this from scratch. There are open-source projects on GitHub that replicate this exact workflow. Search for puppeteer-pdf-template or html-to-pdf-node. Clone one. Swap the template. Adjust the CSS. You will have a working pipeline in an afternoon. The risk is that these projects get stale. Dependencies break. I recommend pinning your Puppeteer version and testing on your own machine before committing to a project. If it works once, it will work every time. If it works sometimes, it is not ready for daily use.

The total time investment to get from zero to a reliable daily PDF generator is about twelve hours. The first hour is setting up Node and Puppeteer. Three hours are spent building the template and CSS. Four hours are debugging layout issues, especially with images and page breaks. The remaining five hours are spread over a week of actual use as you refine the workflow. After that week, you are spending about three minutes per PDF. That is the number that matters. If you create content daily, the question is not whether you need a PDF pipeline. It is whether you can afford to keep doing it manually. The answer is usually no.

Ultimate Guide To Content Creation | PDF
Ultimate Guide To Content Creation | PDF