Getting Your Head Around My Broken Language Book

I first ran into My Broken Language Book about three years ago when I was troubleshooting a batch of malformed localization files for a multi-language content platform. The tool itself is a scripting utility that parses broken or partially corrupted language data, maps out the structural gaps, and spits out a clean replacement structure you can plug back into your pipeline. It is not a magic fix for everything, but it saved me from manually rebuilding over forty language bundles after a deployment went sideways. The core idea is simple: give it a set of broken or inconsistent language files, and it uses pattern matching, schema inference, and missing-key reconciliation to produce a complete set that aligns across all target languages. It does not guess at translations. It only reconstructs keys, formats, and structural relationships. You feed it JSON, XML, or YAML language packs. The parser detects mismatched keys, orphaned strings, and formatting drift. Then it generates a baseline reconstruction and flags each inconsistency for manual review.

How to Use It Step By Step

Start by installing the tool. Most users pull it via a package manager, though there is also a Docker image if you prefer isolation. Run the initialization command to set up a config file. The default config covers common cases like i18n key hierarchies, plural forms, and RTL language ordering. After that, point it at your language directory and run the scan. It will output a breakdown file showing missing keys per language, mismatched variable placeholders, and structural anomalies. I usually let it run overnight on large projects because the scan phase can take a while depending on how many language variants you are working with.

Realistic Edge Case I Hit

Here is the specific problem I ran into: one of our legacy language packs used an older numbering system for plural index keys, and the tool initially flagged every single plural form as a missing key. That was clearly wrong. What I did was add a custom mapping override in the config that told the scanner which key prefix belonged to the old plural scheme. Once that was in place, it correctly merged the legacy keys with the current schema and only flagged the truly missing entries. That override step is something the documentation mentions briefly but does not emphasize enough. If you are working with legacy localization data, do not skip the config mapping section.

Get the Full Details

My Broken Language | CBC Books
My Broken Language | CBC Books

Download and Setup

You can find the latest release on the official GitHub repository. The download page includes release binaries, source code, and setup instructions. I recommend pulling the version tagged with the most recent date stamp rather than relying on a default branch install, since language tooling changes frequently and older versions sometimes misparse newer schema formats. One thing I see people do wrong is assuming the tool will fix translation quality. It does not. It only fixes structural integrity. If a key exists in English but contains a blank or poorly worded string, the scanner will mark it as valid and move on. You have to review flagged keys yourself for content issues. Another pitfall is running the reconciliation pass without first normalizing your base language file. If the source language pack has inconsistent key casing or duplicate entries, the tool will propagate those problems into every reconstructed language. Run a normalization step on the source first, then feed that clean file into the scanner.

Limitations You Should Know About

The tool works well for standard localization formats. It struggles with custom serialization schemes and proprietary language pack formats that do not follow conventional key-value structures. In those cases, you will need to write a custom parser module or convert the data to a supported format before running the scan. Performance also degrades noticeably once you cross roughly fifteen thousand keys across all language bundles. I have seen scan times climb from ten minutes to nearly an hour on projects that size. If you are working at that scale, consider breaking your language packs into smaller chunks and running separate scans, then merging the results. I still use this tool periodically, mostly for audits and recovery work after failed deployments. It is not a daily driver, but when things break, it is usually the first thing I reach for.