Those Old Hacks You Remember From the 90s
I spent way too many hours porting legacy C code from a DEC Alpha to x86 in the late 90s. What came with it was a whole catalog of Vintage Coding Tricks that made perfect sense on machines with 4MB of RAM and no memory protection. Some of them still show up in codebases I touch today, usually hiding behind a "don't fix what ain't broken" comment. Here is how I actually use these tricks, what breaks when you try them on modern hardware, and the specific edge case I ran into last month that proved some of them still matter.
Vintage Coding Tricks That Still Work (With Caveats)
The n & (n-1) bit count trick. Brian Kernighan wrote this one and it shows up everywhere. It clears the lowest set bit in a single operation instead of looping through all 32 or 64 bits. The code looks like this: int count = 0; while (n) { n &= n - 1; count++; } This used to be genuinely faster than a shift-and-check loop on machines without a population count instruction. On modern x64 with __builtin_popcount it compiles down to a single CPU instruction anyway, so the manual version is just there for tradition at this point. I still write it out in interviews because it filters out people who have never touched bit manipulation.
Integer division by magic numbers. This is the trick where you replace a division by a constant with a multiplication and a shift.gcc does it automatically with -O2, but writing it out by hand was a whole discipline back when compilers were stupid. The formula involves computing the multiplicative inverse of the divisor modulo 2^32 or 2^64, then right-shifting by a calculated amount. It sounds insane until you remember that integer division on a Pentium III took around 30-50 cycles and a multiply took 3. The difference was the difference between a real-time system meeting its deadline or not. The Duff device. Tom Duff invented this at Lucasfilm when he needed to optimize a block copy loop for the NeXT workstation. You unroll a switch statement into a while loop in a way that should not compile but does in C. It looks like garbage: register int n = (count + 7) / 8; switch (count % 8) { case 0: do { *to = *from++; case 7: *to = *from++; case 6: *to = *from++; case 5: *to = *from++; case 4: *to = *from++; case 3: *to = *from++; case 2: *to = *from++; case 1: *to = *from++; } while (--n > 0); }
Get the Full Details

It was fast on old RISC machines where loop overhead mattered. On anything with a decent auto-vectorizer it makes zero difference. I found it in a production graphics driver last year and spent an hour convincing my team that yes, this is a real thing and no, we should not remove it because someone might still be targeting a machine where it helps.
What Breaks When You Use These Today
Bit tricks assume two's complement representation. They work everywhere now because the C standard effectively mandated it in C99, but if you ever end up on a DSP with sign-magnitude arithmetic the whole category collapses. I ran into exactly this problem six months ago when someone tried to port a bitmap compression library to a TI C55x DSP. The right-shift behavior on signed integers was arithmetic shift on x86 but logical shift on that architecture. Values that should have been negative came out positive and the decompression produced garbage output that crashed the audio pipeline. The fix was casting to unsigned before every shift operation, which added maybe 8% overhead but stopped the segfaults. Another thing nobody warns you about: these tricks assume the machine does not have undefined behavior protections. Stack canaries, ASLR, address sanitizers, and hardened compilers will happily break code that depends on reading uninitialized memory or overflowing an integer. A lot of the classic tricks from the K&R days rely on properties of memory layout that modern security models explicitly destroy. If you are trying to use vintage tricks in code that runs under ASAN or UBSAN you will need to isolate that code behind a flag or a separate compilation unit. Magic number division is a false economy on modern CPUs. People still reach for it because they read about it in old textbooks. A Skylake core does integer division in roughly 20-80 cycles depending on the operand, but it also has a hardware multiply that does 64-bit in about 3-5 cycles. The magic number trick saves maybe 10 cycles in the best case and costs you code readability. Just let the compiler do it.
Manual loop unrolling fights your CPU. Out-of-order execution and branch prediction make hand-unrolled loops mostly harmful on anything newer than a Pentium 4. The CPU unrolls loops itself based on its trace cache. Hand-written unrolling just adds executable bloat and can actually hurt performance by filling up the instruction cache. I benchmarked a manually unrolled AES key expansion routine against the compiler's auto-unrolling and the compiler won by about 12% on a Coffee Lake chip.

When They Actually Help
Embedded systems without an FPU still benefit from integer math tricks. I was doing work on a Bluetooth audio codec for a chip that had no floating point unit and every cycle counted. The pitch detection algorithm used a float-based autocorrelation that ran at about 80% CPU load. Converting the critical inner loop to fixed-point integer math using bit shifts for multiplication by powers of two dropped it to 35%. The code was uglier but it fit. Game engines and real-time rendering code still use bit twiddling because those paths are hot enough to justify the maintenance cost. The alpha-to-coverage calculation in our renderer uses a table-lookup trick from the PlayStation 2 era that avoids a branch entirely. Branch misprediction costs about 15 cycles on that architecture and we call that code path thousands of times per frame. Code golf and competitive programming are legitimate use cases. Some of the classic bitwise hacks for finding the next higher number with the same population count, or for isolating the rightmost set bit with n & -n, are genuinely useful when you need to generate subsets or permutations in a tight loop.
My Actual Workflow for Deciding What to Keep
I look at the target platform first. If it is a modern desktop or server, the answer is almost always "use the standard library." If it is embedded, bare metal, or has hard real-time constraints, I evaluate each trick individually. I benchmark it. If the trick does not measurably improve performance on the target hardware, I drop it. Readability matters more than you think when you have to debug something at 2 AM. For the projects where I do keep vintage tricks, I wrap them in clearly labeled utility functions with a comment pointing to the original source. /* Bits from Kernighan, C Programming Language, 2nd Ed, p. 36 */ This way whoever maintains the code knows it was intentional and not just some random optimization from three years ago. There is also a preservation angle. Some of these tricks are part of computing history and understanding them helps you read older codebases that still run critical infrastructure. I have seen power plant control software and medical device firmware that still run on code written in the 80s and 90s. The people maintaining it need to understand why the code does what it does, not just what it does.
Where to Actually Learn This Stuff
Hacker's Delight by Henry Warren is the canonical reference. It covers bit manipulation tricks with explanations of why they work and when they are appropriate. Bitmap Tricks and Optimization Techniques by Sean Anderson is a much longer online resource that goes deeper into stuff like integer square root approximations and fast inverse square root variants. David Stafford's bit Twiddling Hacks page is another good starting point for people who want to see a large collection organized by operation type. For the assembly-level vintage tricks, looking at old GCC optimization passes and the Godbolt compiler explorer will show you what the compiler actually emits. You learn a lot faster when you can see that your fancy manual trick compiles to worse code than the straightforward version. The old Usenet groups like comp.lang.c and comp.arch are full of threads about these techniques. They are not always accurate but they are primary sources. I still go back to them when I am trying to remember whether a particular trick was controversial or universally accepted.
