Working with Cortex M chips in both assembly and C is less glamorous than people make it seem

I spend most of my time reading register manuals and debugging hard faults that make absolutely no sense until 2 AM. That is just the life. People think mixing assembly with C on ARM Cortex M cores sounds like some elite low-level hacking endeavor. It is not. It is mostly reading reference manuals and accepting that your compiler will do things you did not expect. The Cortex M family covers everything from the ultra-cheap M0+ up through the M33. The architecture is Thumb-2 based, which means every instruction is either 16 or 32 bits. The compiler handles this transparently. When you write inline assembly, you are working inside the ARM GNU toolchain, which uses GAS syntax. That is the GNU Assembler. It does not look like the Keil or IAR assembly you might have seen in older tutorials. Let me give you a concrete example. Say you need to read the current program status register. In C you would write something like this.

uint32_t psr;
__asm volatile("mrs %0, ipsr" : "=r"(psr));

The compiler generates the right Thumb instruction. You are using the volatile keyword because without it, the optimizer might remove the whole thing if it thinks the result is unused. This is one of those things that trips up everyone at least once. The real friction comes when you start trying to interface assembly functions with C functions. The calling convention on Cortex M is relatively straightforward. Arguments go into R0 through R3. Any extra arguments spill onto the stack. Return values sit in R0. The compiler expects the callee to preserve R4 through R11. If you clobber those registers in an assembly routine and do not tell the compiler, you are going to get strange behavior that takes hours to track down. I spent an afternoon once debugging a peripheral driver where an assembly routine I wrote was silently corrupting a pointer. I had forgotten to preserve R7. The function compiled fine. It ran fine for a while. Then some random variable halfway across memory would get overwritten and the MCU would hard fault. The fix was adding a simple push {r4-r7} at the top and pop {r4-r7} before the bx lr. Basic stuff. It should have been obvious.

Here is something most beginner guides do not mention. The Cortex M core has a feature called the unaligned access exception, which is disabled by default on M3 and M4. If you are writing assembly that moves data around and your pointers are not properly aligned, the hardware throws a usage fault. You will see the fault address point somewhere random. The actual problem is that your assembly loaded a half-word from an odd address. Most C code avoids this because the compiler aligns variables automatically. Assembly does not get that courtesy. When it comes to choosing between writing things in C versus assembly, the rule of thumb is that assembly is worth it for a handful of scenarios. Interrupt service routines where you need precise timing. Bit-banging a protocol that the hardware peripheral cannot handle. Reading device-specific registers that the vendor does not expose through C headers. Everything else is almost certainly better written in C. The compiler is better at register allocation than you are, and it will optimize loops in ways you will not bother to think about. I once optimized a software UART routine by rewriting the critical section in assembly. The C version ran at about 9600 baud with interrupts enabled. The assembly version pushed it to 115200 baud reliably. The difference came down to knowing exactly how many cycles each instruction took. In C, the compiler inserts stack frame setup code and padding that you cannot control. In assembly, you count every cycle and line up the bit timings precisely. That was worth the maintenance headache.

Get the Full Details

A #lady returning #home after a #hard #days #work in the #… | Flickr
A #lady returning #home after a #hard #days #work in the #… | Flickr

But here is the honest part. That same software UART written in assembly took me about six hours to get right. The C version took twenty minutes and worked fine at lower baud rates. If you are building a product and the timing budget is not tight, use C. Assembly is not a shortcut. It is a trade-off. The toolchain matters more than people admit. I use the Arm GNU Toolchain, specifically gcc-arm-none-eabi. Version 12 and 13 have good Thumb-2 support. Older versions had quirks with certain inline assembly constraints. If you are working with an M0+ and need to squeeze performance, check the actual generated assembly. Sometimes the compiler expands a simple loop into something inefficient. A hand-written assembly version can cut execution time significantly on these low-clock devices. One counter-intuitive thing about inline assembly on Cortex M: the compiler does not always generate the exact instruction sequence you might expect. If you write a constraint like "r" for a register operand, the compiler picks the register. It might pick R0 even when R0 already holds something important. You can work around this by using the earlyclobber modifier. Writing "=r" instead of "="r" tells the compiler that the output register cannot overlap with any input registers. This prevents subtle bugs in multi-step assembly blocks.

Debugging assembly is painful. The STM32CubeIDE debugger shows the assembly alongside the C source, which helps. But when you step through inline assembly, GDB sometimes jumps around in ways that feel broken. The issue is usually that the compiler reorders instructions or unrolls loops behind your back. You can force single-stepping through the assembly by disabling optimization for that specific function. Add __attribute__((optimize("O0"))) to the function signature. It is a dirty workaround but it makes the debugger behave. Memory mapping is another area where assembly and C intersect in annoying ways. The Cortex M memory map is fixed. Boot address, vector table, flash, SRAM, peripherals. If you are writing a bootloader in assembly, you need to understand this map cold. I once wrote a minimal boot sequence that jumped to an application starting at address 0x8002000. The application assumed the vector table was at 0x08000000. The NVIC does not move the vector table automatically. I had to copy the table in assembly before jumping, or configure the VTOR register in C before the jump. Both approaches work. The second is simpler. The downsides of writing assembly are real. You lose portability. Code written for an M4 does not run on an M0. The instruction sets overlap but not completely. You lose compiler optimizations. You lose readability. Your code becomes unmaintainable after six months unless you document every decision. And let me be clear: this is not something you should do just to learn. Do it when you have a specific performance or timing requirement that C cannot satisfy.

For anyone starting out, I would recommend writing the C version first. Profile it. Measure the actual execution time with a toggle pin and an oscilloscope. If the C code meets your requirements, ship it. If it does not, identify the bottleneck. Is it a tight loop? An interrupt handler? A bit manipulation sequence? Then rewrite only that part in assembly. You will save time and your code will be easier to maintain. The download resources for toolchains and debuggers are straightforward to find. The Arm website hosts the GNU toolchain. STMicroelectronics provides CubeIDE for their chips. Nordic has its own setup for nRF devices. You do not need to chase down obscure compilers. The standard tools work fine for mixing C and assembly on Cortex M platforms. One final practical note. The vendor header files for Cortex M microcontrollers define all the peripheral registers as C structs with memory-mapped addresses. You do not need to write assembly to access them. Reading a register is just a C assignment. Writing is just a C assignment. The assembly is only necessary when the C abstraction layer is too slow or too imprecise for your needs. And those cases are rarer than most tutorials imply.

Rice Terraces | There are no flat patches in sapa, so the vi… | Flickr
Rice Terraces | There are no flat patches in sapa, so the vi… | Flickr