Getting Started With Real-Time Interfacing On Cortex-M

Real-time interfacing on Arm Cortex-M microcontrollers isn't about fancy middleware or a commercial RTOS license. It's about understanding the hardware interrupt system, getting the NVIC configuration right, and writing peripheral drivers that don't accidentally starve your timing-critical code. The whole thing breaks down into a few practical layers, and most of the headaches come from people skipping the bottom two. I spent about three months debugging a UART RX interrupt that was supposed to grab data at 115200 baud with zero missed characters. The MCU was a STM32F407 running at 168 MHz. The problem wasn't the baud rate, wasn't the buffer size, and wasn't the CPU speed. It was a DMA channel that had been configured but never actually started, and a half-written ISR that fell through to a default handler which re-enabled interrupts before clearing the flag. Misfiring interrupts at that frequency created a cascade that starved the main loop and caused timestamp drift that looked like a software bug for two weeks. The fix was writing a single clean ISR that checked the flag, read the DR register, incremented the ring buffer head, and cleared the interrupt pending bit—all in under 40 instructions. It ran in roughly 1.2 microseconds on that core. The NVIC is where real-time behavior lives or dies. Each Cortex-M core has a nested vectored interrupt controller that handles priority grouping, interrupt latency, and context saving automatically. When an interrupt fires, the hardware pushes eight registers onto the stack in about 12 cycles. If you have multiple interrupts with the same priority level, they're handled in fixed hardware order, not in the order you registered them. That matters more than people expect when you're doing something like simultaneous ADC conversions and PWM updates.

Priority levels on Cortex-M use only the top bits of the 8-bit priority field. By default, most vendors configure 4 bits for preemption priority and 0 bits for sub-priority, giving you 16 distinct preemption levels. If you call NVIC_SetPriorityGrouping() with a different split, you change how many levels can preempt each other. I once saw a project where someone set four sub-priority bits and zero preemption bits, which meant every interrupt could be preempted by every other interrupt. That sounds flexible until you realize your highest-priority timer interrupt gets delayed by a low-priority debug UART interrupt that fires constantly. The solution is usually keeping preemption bits at 4 or higher and only allowing true timing-critical interrupts—usually one or two—to have higher preemption priority. The clock tree is the next thing that trips people up. Cortex-M peripherals don't run at core clock speed by default. On an STM32, the APB1 peripherals max out at 42 MHz even if your core is running at 168 or 180 MHz. APB2 peripherals can go up to 84 MHz. If you need a timer to generate a precise interrupt at a specific frequency, you have to account for the actual peripheral clock, not the SysTick or core clock. A common mistake is dividing the core frequency by your desired timer frequency and getting a value that's exactly half what it should be because the peripheral clock is running at APB divider speed. Check your reference manual's clock tree diagram before writing any timer initialization code. It saves about an hour of head-scratching per peripheral.

Setting Up The Interrupt Infrastructure

Start by defining your interrupt priorities in a single header file. Don't scatter them across multiple source files. I've seen projects where two different engineers assigned the same numeric priority to two different high-frequency peripherals, which meant whichever one got wired first in the linker would silently win preemption. Put it all in one enum or constant block and commit it early. ISR design principle: keep them short. A real-time ISR on Cortex-M should ideally execute in under 100 instructions. Anything longer, and you're risking missing subsequent interrupts or inflating your worst-case latency. Do the time-sensitive work inside the ISR—read the peripheral register, update a flag or buffer pointer, clear the interrupt—and push the heavier processing into the main loop or a lower-priority deferred task. Ring buffers are the standard way to handle data flowing from ISRs to the main loop. They work because the ISR writes to a head pointer and the main loop reads from a tail pointer, and as long as you use atomic operations or disable interrupts briefly around the pointer update, you avoid race conditions. A 256-byte ring buffer on a typical Cortex-M can handle several kilobytes per second of serial data without ever dropping a byte, assuming your ISR is clean.

Get the Full Details

Embedded Systems: Real-Time Interfacing to Arm Cortex-M Microcontrollers: Valvano, Jonathan W ...
Embedded Systems: Real-Time Interfacing to Arm Cortex-M Microcontrollers: Valvano, Jonathan W ...

SysTick is the built-in system timer that runs independently of your peripherals. It's ideal for generating a stable system tick, but it conflicts with FreeRTOS and other RTOS kernels if you try to use both. I've worked on projects where we ran FreeRTOS on one core of a dual-core STM32H7 and used SysTick on the other core purely for timestamping sensor data. That required disabling SysTick in the RTOS configuration and writing our own microsecond counter based on a hardware timer instead. It added maybe 20 lines of code and eliminated the tick conflict entirely.

Peripheral-Level Timing Techniques

GPIOs on Cortex-M can be toggled in a single cycle if you use the ODR register or the BSRR register correctly. The BSRR is especially useful because you can set and clear pins in the same instruction without a read-modify-write sequence. That matters when you're doing software PWM or generating precise pulse trains. A single BSRR write takes one cycle, so toggling a pin at 10 MHz on a 168 MHz core is feasible if you're careful about what else is happening in the loop. Timer interrupts are the backbone of real-time systems. Configure your timer prescaler and auto-reload value to hit your desired interrupt period exactly. The formula is straightforward: interrupt frequency equals the peripheral clock divided by (prescaler plus one) divided by (auto-reload plus one). Use a calculator or write a small Python script rather than doing this by hand. Getting the prescaler wrong by one value changes your timing by a percentage that compounds across multiple timers in the same system. DMA transfers paired with timer events are how you handle high-throughput data acquisition without CPU intervention. A typical pattern is configuring a timer to trigger a DMA request on each conversion complete event, and having DMA write directly into a circular buffer in RAM. The CPU only gets involved when the buffer wraps around and an interrupt tells it to process the accumulated data. This pattern can handle ADC sampling at several hundred kilohertz on a mid-range Cortex-M with virtually zero CPU overhead after the initial setup.

Watchdog timers are often treated as an afterthought, but they're critical for real-time systems that run for extended periods without human intervention. The independent watchdog (IWDG) on STM32 runs from its own RC oscillator, so it keeps ticking even if the main clock fails. That's useful for detecting software hangs, but it also means you need to service it regularly or your system resets unpredictably. I've seen production boards reset every 72 hours because a developer added a new peripheral initialization path that forgot to call the watchdog refresh function. The fix was adding the refresh call to every code path that could potentially run for extended periods without hitting the watchdog routine.

Embedded Systems: Real-Time Interfacing to ARM Cortex-M Microcontrollers (Introduction to Arm ...
Embedded Systems: Real-Time Interfacing to ARM Cortex-M Microcontrollers (Introduction to Arm ...

Debugging Real-Time Behavior

Logic analyzers and oscilloscopes are non-negotiable for real-time interfacing work. Software-only debugging—printf statements, breakpoint stepping, variable inspection—changes the timing of your system enough to make real-time bugs disappear or appear inconsistently. I once spent two days chasing an interrupt latency issue that turned out to be caused by the debugger halting the core during a specific timer interrupt, which made the ISR appear to fire late when it was actually firing on time and the measurement tool was the problem. ITM (Instrumented Trace Macrocell) is available on Cortex-M3 and later cores through the SWO pin. It lets you send formatted strings from your code to a host debugger without using UART or blocking. The overhead is roughly 4 microseconds per character at 1 MHz SWO clock, which is negligible compared to what a blocking UART print would do. It's useful for logging timing data and interrupt entry/exit points without corrupting your real-time measurements. Swoole and OpenOCD can capture trace data and decode ITM packets on the host side. The setup is a bit fiddly—the SWO clock configuration has to match between the MCU and the debugger, and not all evaluation boards break out the SWO pin—but once it's working, you get real-time streaming of diagnostic data without stalling your application.

Common Pitfalls And What To Do Instead

One of the most common mistakes is enabling interrupts globally inside an ISR before clearing the pending flag. Some NVIC implementations allow nested interrupts even without explicit nesting configuration, and if you enable global interrupts prematurely, you can get re-entered ISRs that corrupt your buffer state or cause double-processing of the same event. Always clear the interrupt flag first, then do your work, and only re-enable if your architecture requires it. Another issue is interrupt priority inversion in systems that use a bare-metal approach with no scheduler. If your main loop takes longer to execute than the period of a high-priority interrupt, the interrupt will starve the main loop of CPU time, and your system becomes effectively interrupt-driven without any of the guarantees you'd get from a proper RTOS. Monitor your main loop execution time with a GPIO toggle and measure it on an oscilloscope. If it's more than 80 percent of your timing budget, you need to restructure your code or move processing to lower-priority deferred tasks. Cortex-M cores have a feature called Tail-Chaining. When two interrupts of the same priority are pending and the first ISR finishes, the core jumps directly to the second ISR without doing a full context save and restore. This saves about 12 cycles compared to a normal interrupt return and re-entry. It's automatic and you don't configure it, but understanding that it exists helps explain why your measured interrupt latency sometimes looks better than the theoretical worst case.

The trade-off with aggressive interrupt handling is power consumption. A Cortex-M4 running all its peripherals and interrupts enabled at full clock speed can draw 80 to 120 milliamps depending on the specific device. If your real-time system has idle periods, gating clocks to unused peripherals and putting the core into sleep mode between events can cut that down to single-digit milliamps. The wakeup latency from sleep modes ranges from 3 cycles for Sleep mode to around 5 microseconds for Stop mode on newer parts, so choose the right sleep level for your timing requirements. There's no free lunch with real-time interfacing. You gain determinism by accepting complexity in your interrupt management, you gain speed by accepting that your code has to be careful about shared resources, and you gain precision by accepting that debugging requires hardware tools, not just a debugger and some print statements. The Cortex-M architecture gives you the primitives—fast NVIC, DMA, hardware timers, low-latency GPIO operations—and the rest is about knowing how to combine them without tripping over the edges.

Embedded Systems : Real-Time Interfacing to the Arm Cortex-M Microcontrollers by Jonathan ...
Embedded Systems : Real-Time Interfacing to the Arm Cortex-M Microcontrollers by Jonathan ...