The Reality of Embedded Software
Most people entering this field come from a web or mobile background and assume the code is the same thing just with smaller RAM. That assumption will break your board within the first week. The differences are structural, not cosmetic. Memory is scarce. Timing is often hard. The hardware does not forgive assumptions. Start by picking a development board you can actually buy without waiting three months, and pick a toolchain that runs on whatever machine you own. STM32 with ST-LINK or DAPLink is fine. NXP's i.MX RT series works if you need more compute headroom. For toolchains, ARM GCC with OpenOCD covers the majority of MCU work. Keep your toolchain pinned to a specific version and stick with it. Version drift across compiler releases causes bugs that look like hardware problems. Before writing any application logic, verify your bringup path. Flash a known-good binary. Confirm you can halt the debugger. Check clock configuration. Verify UART output at the baud rate you expect. Most projects fail here because someone changed a prescaler value and forgot to update the console logging line. Document the exact clock tree state you expect, then write a short function that reads back and reports every peripheral clock frequency at startup. This takes five minutes and saves days of debugging later.
How Memory Actually Works in Practice
Static allocation is not a preference. It is usually a requirement for anything running in a safety context or on resource-constrained silicon. Dynamic allocation introduces fragmentation, which is invisible at first and then manifests as a random crash six months after deployment when the heap has been stressed by a specific sequence of events that nobody anticipated. I use a bump allocator for temporary structures and fixed pools for recurring objects. Here is what that looks like in a real project:
typedef struct {
uint8_t buffer[CONFIG_POOL_SIZE];
size_t offset;
} pool_t;
static pool_t g_sensor_pool;
void* pool_alloc(pool_t* p, size_t size) {
if (p->offset + size > sizeof(p->buffer)) return NULL;
void* ptr = &p->buffer[p->offset];
p->offset += size;
return ptr;
}
void pool_reset(pool_t* p) {
p->offset = 0;
}
This is trivial. It is also more reliable than calling malloc in a real-time loop. The cost is that you cannot free individual allocations, but in embedded work you rarely need to. You reset the pool at a natural boundary, like the start of each control cycle or at init time. Microcontrollers do not run at a single clock speed across all peripherals. The CPU, USB, ADC, timers, and communication interfaces often run off different clocks. Getting these wrong does not always cause an immediate failure. It causes timing that is off by a factor you did not expect, and the symptom is intermittent. A UART error flag that only appears under certain conditions. An ADC sampling rate that drifts when the regulator enters light-load mode. A timer interrupt that fires at half the expected frequency because the APB prescaler was not accounted for in the peripheral clock calculation. Always calculate peripheral clock frequencies from the clock tree diagram, not from the system clock value alone. Write a helper function that prints the actual configured frequency for every peripheral your software touches. Do this at boot. Compare against the datasheet. If they do not match, the rest of your debugging is already compromised.
Get the Full Details

Interrupts and Critical Sections
Shared variables between an interrupt service routine and the main loop require careful handling. A volatile keyword alone does not make a multi-byte read atomic on a Cortex-M0. It prevents the compiler from caching the value, but it does nothing for the hardware behavior of reading a 32-bit variable across two 16-bit accesses on a 16-bit bus. The workaround is alignment and access size control. Declare shared variables with the correct alignment for their size, and use the compiler's built-in atomic operations when available. On Cortex-M3 and later, LDREX/STREX instructions exist. On M0, you need to disable interrupts around the access or use a lock-free ring buffer designed for the target architecture. I learned this the hard way on a motor controller project. The current-sense ADC result was a 12-bit value stored in a uint16_t. The ISR wrote it. The control loop read it. The variable was declared volatile. Everything looked correct. The motor would jerk randomly under load. The root cause was that the control loop was reading the uint16_t in two separate 8-bit accesses, and between those two accesses, the ISR updated the value. The result was a blended number that never actually existed. Switching to a uint32_t with proper alignment and using a disable-enable critical section around the read eliminated the issue entirely. The jerk disappeared. The motor ran smoothly.
Debugging Without a Debugger
Sometimes you do not have debug access. The device is in the field. You need to understand what happened. Hardware trace ports like ETM are expensive and overkill for most projects. A lightweight logging strategy using a circular UART buffer gives you enough visibility without impacting performance significantly. Key design points for this approach: Use a fixed-size UART TX buffer. Write logs to the buffer in the interrupt or main context. Have a single UART IRQ drain the buffer. Do not block the application while transmitting. Timestamps from a hardware timer are useful, but keep the timestamp format small. A 32-bit tick count is enough. Do not format dates and times in an embedded log. Parse the ticks afterward on the host side if you need readability.
I once debugged a thermal shutdown issue in a field-deployed HVAC controller using nothing but UART logs and a logic analyzer. The device would shut down after approximately 40 minutes of operation. The fix was that a watchdog timer was being fed in the main loop but a secondary timeout detection in the application code was not accounting for a scheduler latency spike during a specific state transition. The secondary watchdog expired, resetting the processor, and the temperature sensor reading was stale. The heating element stayed on. By the time the system recovered, the thermal cutoff had triggered. Logs showing the scheduler tick count at each state entry point made the gap obvious.

Build Systems and Reproducibility
Makefiles work. CMake works. The important part is that the build is reproducible. Pin every dependency. Record the compiler version. Store the linker script and startup code in version control alongside your application. Do not assume that rebuilding the project next month with a slightly different toolchain version will produce the same binary. It will not. Static analysis tools catch more issues than you think they will. Coverity, cppcheck, or even the built-in analysis in newer IDEs will flag unreachable code, potential null dereferences, and integer overflow conditions before they reach the hardware. Run them on every build, not just before release.
Pitfalls That Beginners Miss
The biggest gap between hobbyist and professional embedded development is not knowing what you do not know. It shows up in assumptions about timing, stack usage, and failure modes. Here are specific areas where projects tend to fail: Stack overflow is almost never detected at runtime unless you deliberately implement guard patterns. Fill unused stack space with a known pattern like 0xDEADBEEF at initialization. In the idle task or a periodic check, scan the stack region and verify the pattern has not been overwritten. This takes negligible CPU time and catches stack exhaustion early. Power supply noise coupling into analog readings is a real issue. Digital switches create ground bounces that shift your ADC reference. Separate analog and digital grounds at the PCB level if you can. If you cannot, place the ADC conversion during a quiet period when no high-current digital peripheral is switching. Check the datasheet for recommended timing relationships between digital activity and ADC conversion.
Firmware updates over a serial interface require a robust mechanism. A corrupted flash write during an OTA update bricks the device. Implement a dual-bank bootloader with a rollback strategy. The active and standby firmware images should coexist. Validate the new image checksum before switching. If validation fails, remain on the known-good image. This adds complexity to your bootloader but eliminates a class of field failures that is difficult to recover from.
Testing Strategies That Actually Work
Unit tests for embedded code are possible but require a different approach than application code. The target code should be structured so that hardware-dependent layers can be swapped for mocks. Use build-time conditional compilation or compile-time injection to separate driver code from business logic. Test the business logic on a host machine. Test the driver code on the target hardware. Hardware-in-the-loop testing is where most production-quality work happens. Even a simple setup with a microcontroller connected to a test rig that simulates sensor inputs and monitors outputs catches more issues than pure software simulation. Automate the test sequences. Record the results. A failing test on the bench is preferable to a failing test in the customer's installation. I worked on a medical infusion pump where the software had to meet IEC 62304 Class B requirements. We spent more time on the verification plan than on writing the application code. Every requirement was traced to a test case. Every test case was executable. Code coverage analysis was not optional. The regulatory audit reviewed the test evidence, not the code itself. This meant the testing infrastructure had to be as reliable as the application. It was slower development initially, but it prevented the kind of retrospective panic that happens when a device needs recall.
Tooling Recommendations
For most MCU development, this stack is sufficient: ARM GCC for compilation, OpenOCD for debug and programming, ST-LINK or J-Link for hardware connection, GDB for debugging, and a serial terminal program for UART communication. GDB plugins like SEGGER J-Link GDB Server make debugging faster than the default GDB protocol over SWD. The difference is noticeable during iterative breakpoint testing. For more complex systems with multiple cores or external memories, consider using a vendor-supplied IDE for initial bringup and then migrating to a command-line workflow for production builds. IDEs are convenient for exploring registers and peripherals, but they obscure the build process. Command-line builds make the build explicit and repeatable.
What This Approach Does Not Solve
Software engineering practices alone cannot fix a poorly designed PCB. If the power integrity is bad, no amount of careful coding will stabilize the analog readings. If the reset circuit is underspecified, brown-out events will cause unreproducible failures that software workarounds can only partially mitigate. Software can handle some hardware imperfections through filtering, retry logic, and graceful degradation. It cannot handle hardware that is fundamentally out of specification. Invest in good hardware design first. Good software design is a multiplier, not a replacement. Similarly, choosing a microcontroller with insufficient resources forces compromises that no amount of clever coding can overcome. A project that needs real-time image processing on a Cortex-M0 is going to struggle regardless of how well the code is written. Select the right silicon for the workload. Under-provisioning is the most common architectural mistake I see. Documentation is another area where embedded projects routinely fall short. Commenting code is not documentation. Architecture diagrams, interface specifications, and design rationale notes are what matter when the original developer is no longer available. Write these documents while the design is fresh. Update them when the design changes. The version control history will show what changed, but it will not explain why.
