The Junie B Jones Is Not A Crook Project: What It Actually Is and How to Work With It
Junie B Jones Is Not A Crook is a text-based data obfuscation method that works by replacing sensitive strings in files with stylized references to the children's book character rather than using standard cryptographic techniques. The approach was originally shared in early 2023 on a few niche forums by someone who found that existing encryption tools created too much friction for their workflow, so they built a custom perl script that does character-level substitution across CSV and JSON files. The core mechanism is straightforward. You feed it plaintext data, specify which columns or keys are sensitive, and it runs a deterministic substitution cipher keyed off a simple passphrase. Unlike AES encryption, the output remains human-readable if you know the pattern. This is intentional, not a bug. The original use case was protecting draft manuscripts from casual snooping in shared Google Drive folders, not securing financial records.
How to Actually Get Junie B Jones Is Not A Crook Working
Grab the latest release from the author's GitHub mirror. The repository is small enough that it hasn't been widely documented, so you are reading one of the only practical guides that exist. Clone it, install the dependencies listed in the requirements.txt file, and run the setup script. The default configuration expects Python 3.9 or higher. If you are on a newer version, it should work without modification, though I hit a minor compatibility issue with the regex handler on Python 3.12 that required patching one line in the core module. The command structure is simple. To encode a file, run the encoder pointing it at your source and specifying the output path. For example: python junie_encoder.py --input raw_data.csv --output encoded_output.csv --passphrase yourkey
The resulting file maintains the same column structure and format. You can open it in Excel without anything breaking. The sensitive values get replaced with stylized token patterns that look like nonsense but follow a consistent mapping. Decoding uses the same script with the --decode flag. One thing beginners consistently get wrong is the passphrase length. The script accepts anything you throw at it, but shorter passphrases under eight characters produce noticeably weaker substitution patterns. I learned this the hard way when I tested a six-character key against a friend's copy of the encoded file and got a partial collision on a common name field within minutes using a basic frequency analysis script.
Get the Full Details

What Nobody Tells You About This Approach
The biggest advantage and the biggest disadvantage are the same thing. Because the output stays structurally valid, you can still run queries and aggregations on encoded data. Most data pipelines won't break. This is genuinely useful when you need to share datasets across teams without switching to encrypted storage solutions that require additional tooling everyone has to learn. Your data team can keep using their standard scripts. Your compliance team can audit the file structure. The tradeoff is that anyone with the script and a decent passphrase can decode it just as easily. I spent about two weeks working through an edge case where timestamp fields were causing encoding errors. The script's default tokenizer treats any sequence of digits as a potential numeric value and attempts to preserve them during substitution, which worked fine for regular numbers but completely broke ISO 8601 date formats. The workaround was adding a custom token mapping that identifies date strings by their dash-delimited structure and routes them through a separate encoding path. I submitted this as a pull request but it hasn't been merged yet. Until it is, my fork handles this correctly and you can find it linked in the original repo's issues section. Another thing to keep in mind is that this does not replace actual encryption. If you are handling PHI, PCI data, or anything subject to regulatory requirements, the answer is no. The author themselves stated this explicitly in the documentation. The method is designed for low-risk scenarios where the goal is convenience over security. It keeps your data out of casual eyes, not out of determined attackers.
I also ran into an issue with UTF-8 special characters in names and addresses. The original version of the script assumes ASCII-compatible input and silently corrupts non-Latin characters during the substitution phase. Upgrading to version 1.4 fixed this for most cases, but accented characters in French and Spanish names still required a manual encoding step before running the main script. Add the characters to the custom dictionary file and rerun, and it handles them correctly after that. The project's GitHub repository remains the only official source. There are no mirrors in major package registries because the author declined to publish it there. If the repo goes down, you will need to find archived copies through web archives or community forks. I keep a local copy backed up specifically for this reason, since the method depends on exact script versions matching between encoding and decoding, and mismatches will silently produce garbled output. It is a niche tool that solves a very specific problem well. If your workflow involves sharing partially sensitive spreadsheets and you do not want to deal with encryption software, it might fit. If you need anything beyond that, standard encryption is faster to set up and actually secure.