A Practical Look at The Gravedigger

The Gravedigger is a bulk data recovery and archival utility that pulls dead links, orphaned assets, and buried reference material out of WordPress installations and static site builds. The name comes from how it works: it crawls a site, finds the content that no longer has an active inbound link, and flags it for review rather than just deleting it outright. That distinction matters because a lot of tools labeled as "broken link finders" will simply report 404s. The Gravedigger goes one step further and checks attachment metadata, sidebar widgets, theme template references, and even database post meta fields to locate content that technically exists but has been functionally removed from navigation. It runs a depth-first crawl starting from the homepage or any seed URL you provide. During the crawl it maintains a live set of discovered internal references and cross-references those against the site's published content table. Anything that lacks a valid internal anchor gets logged. The output is a CSV or JSON report with severity tags, a confidence score, and a suggested action: redirect, merge, or archive. The confidence score is the part most people skip, and it is worth looking at because it correlates with whether the dead reference is a genuine problem or just a tracking parameter that looks broken on paper. I downloaded the latest release from the project page about nine months ago and have been running it against client WordPress installs ever since. The installer bundles a small daemon process, so make sure your firewall allows outbound connections on the default port. If you run it behind a corporate proxy, the config file has a proxy block near the top. Leaving that empty will cause the crawl to stall after about two hundred pages, which is a common failure mode that looks like a bug but is just a missing environment setting.

Here is the basic setup sequence: Extract the archive to a working directory. Open config.json and set your base_url, output_format, and max_depth values. The default max_depth is three, which covers most brochure sites in under twenty minutes. If you are scanning an e-commerce site with layered category hierarchies, bump it to five and expect the run to take closer to two hours depending on page count. Add any user-agent strings your hosting provider requires, set the rate_limit_seconds value to something between one and three to avoid triggering WAF rules, then run the binary from the command line. The first scan I ever ran produced a report with nearly fourteen hundred flagged items. Most of those were false positives caused by URL parameters that the parser treated as separate pages. The workaround was to add a canonical_rules array to the config file with the parameter patterns I wanted to ignore. I added a few common tracking segments, session IDs, and UTM permutations. After that, the true positive count dropped to about eighty, which is a manageable number for a content team to review in a single afternoon.

Using The Gravedigger on Real Sites

The tool does not fix anything by itself. It generates a report and stops. That is intentional, and it keeps the execution fast. You import the CSV into a spreadsheet, filter by severity, and decide which rows need redirects, which need content merges, and which can be ignored. I usually split the work into two passes. The first pass handles the high-confidence broken internal links that point to deleted product pages or removed blog posts. The second pass tackles medium-confidence orphaned assets like images referenced in post meta but not embedded in any visible template. One edge case I ran into recently involves multilingual WordPress setups using a translation plugin that stores language slugs in a custom taxonomy. The Gravedigger saw those taxonomy terms as orphaned because they had no visible internal link from the primary language archive. The fix was to add an exclusion rule for that specific taxonomy namespace rather than marking the terms as dead content. Without that exclusion, the report was inflated by about twelve percent and the team wasted time reviewing taxonomy entries that were functioning correctly. I learned to add the exclusion before running any scan on a multilingual install now instead of cleaning up the report afterward.

Common Pitfalls and Where the Tool Falls Short

The biggest limitation is that it only understands internal references. If your site loads content through an API endpoint and the API returns HTML fragments that contain links back to the same domain, The Gravedigger will not see those unless you configure the API URLs as additional seed pages. That is not a flaw in the logic, it is just the scope of the tool. Another issue is caching. If your CDN serves a cached page that includes a link to a resource you deleted yesterday, the crawler will read the cache and report the link as healthy. The solution is to disable cache for the crawler user-agent or to run the scan during a low-traffic window when the cache is cold. Neither option is ideal, but both prevent the report from lying to you. Performance-wise, the tool uses a single-threaded crawler by default. You can enable concurrent fetching in the config, but it increases memory usage and can trigger rate limits on shared hosting environments. I typically run it on a staging copy of the production site first, verify the output, and then run it live once I know the concurrency settings will not overwhelm the server. That extra step usually saves about an hour of debugging compared to running it blindly on production.

Why The Gravedigger Is Worth Using Anyway

Most sites accumulate dead references over time. Content gets updated, products get discontinued, and internal links become stale. Left unchecked, that debris shows up in search console reports, slows down page load through unnecessary requests, and creates a poor experience for users who click a menu item and land on a missing page. The Gravedigger catches that debris before it becomes a support ticket. It is not a perfect scanner and it will misreport on complex SPAs, but for traditional WordPress and static site architectures it produces a clean, actionable report in a reasonable amount of time. If you maintain a content-heavy site, it is worth adding to your quarterly maintenance routine.

Get the Full Details

Who will be in the Detroit Lions’ starting secondary in Week 1? | Pride ...
Who will be in the Detroit Lions’ starting secondary in Week 1? | Pride ...