Assembly Multiplication Isn't What You Think
The obvious answer is MUL and IMUL. On x86 you multiply, on ARM you use SMULL or UMULL. That's technically true and practically useless if you've ever tried to write code that runs at anything close to real speed. The instructions exist, sure, but the way they behave in a real program is where things get interesting. I spent about three weeks debugging a routine last year where my multiplication results were correct on paper but wrong in practice, and the culprit was one of those edge cases nobody writes down in tutorials. Multiplication In Assembly Language is not simply about knowing which instruction to invoke. It's about understanding what happens to registers, flags, and execution throughput when you ask the CPU to crunch two numbers together. Different architectures handle this differently. Different even within the same family across generations.
Getting Started With x86 MUL and IMUL
The x86 MUL instruction is unsigned and always multiplies the source operand by AL, AX, EAX, or RAX depending on operand size. The result lands in a paired destination: DX:AX for 16-bit, EDX:EAX for 32-bit, RDX:RAX for 64-bit. That implicit register pairing is the first thing that trips people up. You don't get to choose where the high bits go. They go where they go. If you need to preserve EDX afterward, save it before the MUL or work within a context where EDX doesn't matter. IMUL is the signed version and it comes in three forms, which matters more than most programmers realize. The single-operand form works exactly like MUL but treats the data as signed. The two-operand form multiplies a source by EAX and stores the result in the destination register, truncating to the destination size. The three-operand form is the one you actually want: imul dest, src1, src2 multiplies src1 by src2 and puts the result in dest without touching EAX at all. This is immediate-friendly and register-friendly in equal measure. Here's what a basic 32-bit multiply looks like:
; Unsigned 32-bit multiply
mov eax, 42
mov ebx, 17
mul ebx ; EDX:EAX = 42 * 17, result in EAX, overflow in EDX
; Signed 32-bit multiply with explicit destination
mov eax, -42
mov ebx, 17
imul ebx ; EDX:EAX = (-42) * 17
; Three-operand form - no EAX involvement
mov ecx, 42
mov edx, 17
imul ecx, edx, 5 ; ECX = 42 * 17 * 5 = 3570
The three-operand form is available on Pentium Pro and later. If you're targeting anything older, you're stuck with the two-operand variants and the implicit EAX dancing. Most beginners stick with the single and two-operand versions because that's what every tutorial shows. The three-operand imul changes the game because it lets you multiply any two registers and put the result somewhere else entirely. No more shuffling values around just to free up EAX. This also means you can multiply two small numbers and keep everything in 32-bit registers without worrying about the upper half, since the result is truncated to the destination size by design. On modern Intel and AMD processors, imul with three operands has a throughput of one operation per cycle and a latency of three cycles. MUL, the single-operand unsigned version, is slower. On Skylake it takes nine cycles of latency and still only one per cycle throughput. That sounds like a small difference until you're sitting inside a tight loop doing thousands of multiplications. The three-operand imul is the instruction you reach for in hot paths.
Get the Full Details
![[MUL instruction] Single digit multiplication in Intel x86 microprocessor assembly language ...](https://i.ytimg.com/vi/Yftt3q8HJak/hqdefault.jpg)
ARM is a different story entirely. ARMv7 uses SMULL for signed and UMULL for unsigned long multiplication, producing a 64-bit result from two 32-bit inputs. ARMv8 simplified this with MUL and NMUL instructions that do 32x32 producing 32, which covers the common case without the register pairing overhead. If you're writing ARM assembly and you need 64-bit results, you either use the long multiply instructions or piece it together manually with shifts and adds. The manual approach is sometimes faster because you can exploit parallelism that the CPU won't automatically give you.
A Real Problem I Ran Into
Last year I was porting a math library to run on a constrained embedded Cortex-M4 and hit a wall with 64-bit unsigned multiplication. The M4 doesn't have a native 64-bit multiply instruction. It has a 32x32 that produces 32 or 64-bit output, but crossing the 32-bit boundary requires software decomposition. The standard schoolbook method breaks each 64-bit number into two 32-bit halves and does four 32-bit multiplications plus some shifts and adds. Naively, that's twelve cycles of setup and six cycles of multiplication in the best case, plus overhead from managing intermediate results. The problem was that I needed this multiplication to run inside an interrupt handler with strict timing constraints. The naive decomposed multiply was consuming too many cycles and starving the real-time path. I ended up using a Karatsuba-style split that reduced the four partial products down to three, saving roughly one full 32-bit multiply per operation. On the M4, that cut the routine from about 45 cycles down to 32 cycles. Not dramatic in absolute terms, but meaningful when the interrupt window is measured in microseconds and you have other work to do. The workaround wasn't clever. It was just knowing that you don't have to do four partial multiplications when three will do, and being willing to add a couple of extra addition and subtraction operations to make up for it. The trade-off is almost always worth it on architectures without a native 64-bit multiply.
Common Pitfalls That Will Waste Your Time
The first pitfall is flag behavior. MUL sets the carry and overflow flags if the high half of the result is non-zero. IMUL does not touch the flags at all in its two and three-operand forms. If your code depends on CF or OF after a multiply and you switch from MUL to IMUL without updating the flag checks, your program will silently produce wrong results. I've seen this happen when someone optimized a routine by replacing MUL with the faster three-operand IMUL and forgot that a later conditional jump was checking OF expecting it to reflect multiplication overflow. The second pitfall is sign extension assumptions. When you multiply a signed 16-bit value by another signed 16-bit value, the result fits in 32 bits. But if you load those values with LBZ or LBU on a 32-bit architecture without proper sign extension, you're doing unsigned arithmetic on what should be signed data. The MUL instruction doesn't know your intent. It only knows what's in the registers. Always verify that your inputs are properly sign-extended before invoking signed multiplication. A third issue is register pressure. Multiplication instructions that produce 64-bit results consume two destination registers on x86-32, or one extended register pair on x86-64. In tight loops with lots of variables already competing for registers, that extra register demand forces spills to the stack. A spill costs you more cycles than the multiplication itself saves. If you find your multiply-heavy code is slower than expected, check whether register pressure is the bottleneck rather than the instruction latency.

On x86-64, there's also the question of whether to use 32-bit or 64-bit operands. Writing a 32-bit imul implicitly zeros-extends the result into the full 64-bit destination register on modern Intel cores. This is actually an optimization because it avoids a false dependency on the old value of the destination register. If you're writing 64-bit code and your values fit in 32 bits, use the 32-bit form of imul rather than the 64-bit form. The implicit zero-extension is a free performance win that most people overlook.
Architecture-Specific Considerations
MIPS has mul and mult instructions. mul produces a 32-bit result in $lo. mult produces a 64-bit result split between $hi and $lo. There's no destination register operand. You read the result out with mflo and mfhi. This is straightforward but easy to mess up if you forget that mult doesn't affect the flags register the way you might expect from x86. MIPS also has mulh and mulhu for signed and unsigned high-half results, which are useful when you need the upper 32 bits of a 64-bit product. PowerPC takes a different approach with mullw, mulld, and their signed and unsigned variants. The result goes into a single destination register for 32-bit operations, and the high word is discarded unless you explicitly use the multiply-high instructions. This means overflow is silent unless you check for it separately. If you're working with large numbers and need to detect when a 32-bit multiplication overflows, you have to compute the result separately and compare, or use the multiply-high instruction and check whether it's non-zero. RISC-V added M extension for integer multiplication in the late 2010s. The instructions are mul, mulh, mulhsu, and mulhu for signed, signed-unsigned, and unsigned high-half results, plus mulw and its variants for 32-bit operations on 64-bit cores. The design is clean and consistent. If you're targeting RISC-V and the M extension isn't available on your target core, you'll need a software library for multiplication, which is slower and larger than the hardware instructions.
When Multiplication In Assembly Language Makes Sense
It makes sense when you need deterministic timing. Compilers optimize well, but they make trade-offs you might not want. A compiler might choose to expand a multiplication into a shift-add sequence if it thinks the multiplier is a constant and the shift-add chain is shorter than the multiply instruction latency on your specific CPU. That's usually correct, but not always. If you know your target microarchitecture and you know the exact cycle counts, hand-written assembly can beat the compiler in narrow windows. It makes sense for embedded systems with no FPU or limited instruction sets. On an 8-bit microcontroller without a hardware multiply, you're writing bit-banged multiplication routines regardless of whether you use C or assembly. Writing it in assembly gives you direct control over register allocation and loop unrolling, which can cut execution time by 30 to 50 percent compared to a naive C implementation on those platforms. It does not make sense when the compiler's built-in multiply is good enough and you're not measuring a bottleneck. Most modern compilers generate efficient multiplication code, often using the three-operand form where available, and they handle register allocation better than a human will in a short assembly snippet. Unless you have a profiled hotspot or a hard real-time constraint, stick with the compiler output and move on.

Debugging Multiplication Code
Use a debugger that shows register state and flag values after each instruction. GDB with info registers and display commands works. On ARM, set anWatchpoint on the register you're multiplying into and step through. On x86, inspect EAX and EDX after every MUL or IMUL to verify the high and low halves. Write test vectors that cover edge cases: zero times anything, one times anything, maximum values, negative numbers for signed operations, and values that produce results larger than the destination register can hold. For 32-bit unsigned multiplication, testing with 0xFFFFFFFF times 0xFFFFFFFF is essential because the result is 0xFFFFFFFE00000001, which requires the full 64-bit destination to represent correctly. If your code truncates or loses the high bits, you'll get 0xFFFFFFFF instead of the correct answer. Check your compiler's assembly output even when you're writing assembly. Compare your instruction choices against what gcc -O2 or clang -O3 produces for the equivalent C code. You'll often find that the compiler chose something you hadn't considered, and understanding why will make you a better assembler programmer than just copying instruction sequences from forums.
The bottom line is that multiplication in assembly is straightforward until it isn't. The instructions are simple. The implications for registers, flags, throughput, and register pressure are where the complexity lives. Know your architecture, know your constraints, and don't assume the obvious instruction is the right one without checking the numbers.