The Problem With Math Extraction
Math equations in scanned PDFs, photos of whiteboards, or legacy academic papers don't just convert cleanly to editable text. OCR software treats them like images or breaks fractions into separate characters. I spent two years dealing with this exact issue while migrating decades of engineering documentation, and the standard solutions either corrupted the equations entirely or missed half the content. There are multiple approaches depending on where your source material lives. I'll walk through the methods that actually work, the ones that look good in demos but fail in practice, and what to do when you hit the inevitable edge cases.
Extract Math Equations From Documents and Images
The most reliable method I've found involves a pipeline: first identify where the equations are, then process them separately using LaTeX-aware tools, then reassemble. Don't try to extract everything in one pass. It rarely works well. For PDFs, the first tool to check is whether the PDF already contains embedded math. Many technically-sourced PDFs use mathematical typesetting engines at the source, meaning the data is already there in LaTeX or MathML form. Run a plain text extraction and search for backslashes or angle brackets. If you find them, you're done. You can convert directly to LaTeX using specialized libraries. If that returns nothing, you're dealing with rendered output, which means optical recognition is your only path. The current state of the art here is Mathpix. Their API converts images and PDFs directly to LaTeX with decent accuracy on standard notation. It costs about $0.01 per equation on their pay-as-you-go tier, and processing a 200-page textbook with heavy equations runs roughly $15 to $30. Not cheap for large volumes, but dramatically faster than manual transcription.
For people who need a free option, LaTeX-OCR from GitHub handles images reasonably well for simpler equations. It uses a transformer model trained on math notation. The quality drops noticeably on complex multi-line derivations or handwritten content, but it handles printed textbooks adequately. Run it locally if you can. The model is around 400MB and works on CPU, though a GPU cuts processing time significantly. I ran into a specific problem last year with a collection of Soviet-era physics textbooks scanned at 300 DPI. The pages were yellowed, the ink had bled slightly, and the equations included Cyrillic variable names alongside standard Latin notation. Mathpix misidentified about 15 percent of the variables. The workaround was to preprocess the images with adaptive thresholding using OpenCV before feeding them to the OCR. Specifically, I applied a Gaussian blur, then Otsu's thresholding, which cleaned up the noise without destroying the thin stroke weights of the fractional bars. This pushed accuracy from roughly 85 percent to about 94 percent.
Get the Full Details

Tools and Libraries
Beyond Mathpix and LaTeX-OCR, a few other options exist depending on your workflow. Tesseract with the equation module can handle basic cases if you configure it properly. The default training data doesn't include mathematical symbols, so you need to load the model or train a custom one. The setup is tedious and the results are inconsistent, which is why most people move past it quickly. For structured documents like Word files or Jupyter notebooks, you can often extract equations directly from the XML. Word stores MathML in its internal format. A simple script parsing the document XML and extracting nodes containing m:
namespaces gives you edit