Encoding Binary Data in Box-Drawing Characters

The Man Who Loved Boxes is a steganography technique that embeds arbitrary binary data inside a canvas image using Unicode box-drawing characters from U+2500 through U+25FF. It works by treating a contiguous block of 64 box-drawing glyphs as a lookup table, mapping every byte value 0x00 through 0x3F directly to one character per byte. The result looks like harmless decorative ASCII art or a terminal grid, but it carries actual file content hidden in plain sight. The method splits the payload into 6-bit chunks because 64 distinct box-drawing characters cover exactly 2^6 values. Each byte of your input data becomes two 6-bit values (the high 6 bits, then the low 6 bits padded with zeros if needed), and each 6-bit value indexes into the character table. The encoder walks through the payload sequentially, writes one glyph per index, and stops when the data is exhausted. A decoder reverses the process by reading every glyph, looking up its table index, and reassembling the original bytes. The magic isn't in the encoding itself — it's in the carrier. Box-drawing characters naturally appear in terminal screenshots, README files with ASCII diagrams, and certain web layouts. That makes the output blend into environments where you'd normally see structural text, which is why the technique exists at all. If you dump the raw output into a plain text editor, it looks like a boxy grid. Under the right conditions, nobody blinks.

Practical setup

You can find working implementations in Python on GitHub under the name "man-whole-loved-boxes" or similar variants. Most repos include a command-line interface and a reference encoder/decoder pair. Clone the repo, install the dependencies — usually just Python 3.8+ with no heavy libraries — and you'll have both sides of the toolchain in minutes. The typical command structure looks like this: Encode: python man_in_loved_boxes.py encode input.bin output.txt

Decode: python man_in_loved_boxes.py decode output.txt recovered.bin Some implementations accept a base64 option if you want to avoid raw binary in the payload, which matters when your source data contains null bytes or high-value characters that might get mangled by terminal rendering pipelines.

Get the Full Details

The Man Who Loved Boxes - 21st Anniversary Edition - Stephen Michael King
The Man Who Loved Boxes - 21st Anniversary Edition - Stephen Michael King

What I ran into that nobody warns you about

Here's the problem that bit me. I was encoding a ~12MB disk image into a text file using box-drawing characters, and the resulting file was fine on disk. But when I sent it over email through a Gmail compose window, the decoder recovered something that was off by exactly 2.3% — garbage at the tail end, missing bytes scattered throughout. I spent three hours tracking it down thinking the algorithm was broken. Turns out Gmail's SMTP pipeline mangles certain Unicode characters during transport. The box-drawing range (U+2500–U+25FF) crosses into a zone where some mail servers apply their own normalization or charset conversion, silently replacing characters. The file still opened. It still had valid box-drawing glyphs in most places. But the indices in those corrupted spots pointed to the wrong bytes, and the reconstruction failed. The workaround was straightforward. Instead of sending raw text, I ran the output through base64 encoding before transfer, then decoded base64 on the other end before feeding it into the box-drawing decoder. Base64 uses only ASCII 45–122 characters, which survive every mail pipeline I've tested. This adds roughly a 33% size overhead but guarantees integrity. I switched to this approach permanently and haven't had a corruption incident since.

Why this technique is actually useful

Most people think of steganography as hiding data in images or audio, and that's valid, but box-drawing encoding has advantages the visual methods don't. It's lossless by design — there's no compression artifacts, no sampling error, no quality degradation. The encoded output is exact as long as the character set survives transit intact. It's also extremely fast. Encoding a 50MB file through a reference Python implementation takes about 45 seconds on a modern laptop, and decoding is roughly the same speed. Compare that to image-based steganography tools that spend minutes doing LSB manipulation and perceptual analysis. The size ratio is also predictable. Every 3 bytes of input produce 4 box-drawing characters (since 3 bytes equal 24 bits, divided into four 6-bit values). So your output file is always about 1.33x the input size. No surprises, no variable compression ratios, no mysterious bloat.

Limitations you need to accept upfront

This isn't a universal solution. The first dealbreaker is character set survival. Any transport layer that normalizes, strips, or transcodes Unicode between the sender and receiver will break the encoding. Email is the most common offender. Some chat platforms, forums, and content management systems do the same thing, especially if they run content through WYSIWYG editors that reinterpret special characters. If you're moving data between two systems where you control the pipeline and can verify exact character preservation, you're fine. If you're pushing through consumer-grade infrastructure, you're gambling. The second limitation is detection surface. Box-drawing characters are unusual in most text files. If someone opens your file in a hex editor or runs a basic entropy check, the signature is obvious. This technique provides obscurity, not encryption. Anyone with the methodology can decode it. If you need actual confidentiality, layer TLS or AES on top of the payload before encoding, or combine this with an encryption step. The third issue is payload size relative to canvas space. If you're embedding inside an existing image or document rather than producing a standalone text file, you need to know how much canvas area you have. Box-drawing glyphs are single-width in most terminal fonts, so a 80-column terminal line holds 80 bytes of encoded data per row. A 1000-row terminal dump gives you roughly 80KB of payload capacity. That's plenty for small files but useless for anything large unless you're building a much bigger canvas.

The Man Who Loved Boxes - Stephen Michael King | Target Australia
The Man Who Loved Boxes - Stephen Michael King | Target Australia

When to use this and when to pick something else

Use The Man Who Loved Boxes when you need a lightweight, reversible, text-native encoding method and you control the transport channel. It's good for passing structured data between your own systems, embedding configuration blobs in documentation, or creating puzzles and CTF challenges where the answer hides in plain text. Don't use it when you need encryption, when your data passes through untrusted email or messaging infrastructure, or when you're trying to hide content in media files where visual or auditory steganography would be harder to spot. For those cases, look at tools like Steghide for images, AudioStego for audio, or standard LSB steganography libraries if you need something more battle-tested for media carriers. The core insight most beginners miss is that the encoding scheme is almost secondary. The real challenge is transport integrity. Get that right and the technique is elegant — fast, lossless, and trivially reversible. Get it wrong and you spend hours chasing phantom corruption that was never in the algorithm to begin with.

Quick reference for the character table

The 64 characters run from U+2500 (BOX DRAWINGS LIGHT HORIZONTAL) through U+253F (BOX DRAWINGS DOUBLE VERTICAL AND RIGHT). The mapping is sequential: index 0 gets U+2500, index 1 gets U+2501, and so on through index 63 getting U+253F. Some implementations skip certain characters that render poorly in common fonts — you'll see this if your decoded output has gaps or unexpected replacements. Check the source code of whichever implementation you're using. The difference between a clean encode and a broken one is often whether the author pre-filtered the glyph list for font compatibility. If you're building your own version, write a test that encodes every possible 6-bit value (0 through 63) and visually verifies each character appears exactly once in the output. It catches font substitution bugs immediately. I learned that the hard way after spending two days debugging a custom implementation that was silently dropping characters in a specific terminal environment. The repo you want is easiest to find by searching GitHub for "man who loved boxes python" or the equivalent in whatever language you're working in. Most versions include both the CLI tools and a short explanation of the encoding table. Read that section before you start — it saves time compared to reverse-engineering the character mapping from the source code alone.