Embedded C Interview Questions: Startup Code, Memory Layout, and Interrupt-Safe Reasoning
Quick Overview
A practical Embedded C interview article connecting an actual Cortex-M3 ELF and linker map to pre-main data copying, zero initialization, memory accounting, and a modeled interrupt lost update. Distinguishes official toolchain facts, original exercises, and untested hardware behavior.
Embedded C interview questions become easier to answer when you can follow a value from the firmware image into RAM, then explain who may change it. Start with three checks: where the initial bytes live, which startup code establishes the application’s state, and whether an interrupt can break the next operation. If a counter changes only when the debugger is attached, a definition of volatile alone will not explain the observation. Trace initialization and the competing writers instead.
This article uses an original Cortex-M3 linking exercise and two small executable models. For additional pointer practice, try PracHub’s C memory-layout question. Its hosted-system assumptions differ from the bare-metal example below; identifying that difference is part of the exercise.
Evidence boundary: Official toolchain and architecture documentation supports the technical rules. The examples and answer recommendations are original preparation material, not candidate-reported questions or evidence of a particular employer’s assessment.

What should happen before application code runs?
Startup code establishes the conditions that application code relies on: usable execution context, required memory initialization, and the platform setup needed to reach the application entry point. The exact sequence belongs to the target, runtime and boot chain. A bootloader, vendor startup file or operating-system loader may do some of this work.
In our deliberately small example, the application expects boot_mode == 3, retry_limit == 5, and two zero-initialized counters. The linker assigns their storage; the reset handler copies the initialized bytes and clears the counter region. GNU ld’s official example demonstrates this separation between image placement and runtime initialization. GNU ld load-address documentation
Do not claim that every device starts directly at your C function. Our vector fixture contains only a stack value and reset entry for inspection. It omits a complete exception table, device clocks, boot configuration, C++ initialization and production fault handling. Explain what another boot stage supplies before assuming the application owns it.
Read the load address and run address separately
VMA, the virtual memory address, identifies where a linked section is addressed during execution. LMA, the load memory address, identifies where its initial image is placed. In a flash-based bare-metal layout, an initialized writable object can therefore consume flash bytes for its initial value and RAM bytes for its running value.
Here is the important part of the original linker script:
.data 0x20000000 : AT(0x08000400) {
_data_start = .;
*(.data*)
_data_end = .;
} > RAM
_data_load = LOADADDR(.data);
.bss (NOLOAD) : {
_bss_start = .;
*(.bss*) *(COMMON)
_bss_end = .;
} > RAM
Official linker fact: LOADADDR returns a section’s LMA. Consequently _data_load is the copy source, while the start and end symbols inside .data describe its runtime destination range. A linker symbol can identify an address without allocating a C object that stores the address. Treat _data_load as the copy-source address, not as a word to dereference to discover that address. GNU ld built-in functions
Actual build output: The little-endian ARM ELF produced for this article contained:
| Region or object | Linked address | Meaning in this fixture |
|---|---|---|
.data image | 0x08000400 | Eight source bytes: 03 00 00 00 05 00 00 00 |
boot_mode | 0x20000000 | First four destination bytes |
retry_limit | 0x20000004 | Next four destination bytes |
.bss | 0x20000008–0x20000010 | Eight bytes; end address excluded |
These are observed build outputs from Apple Clang 21.0.0 and GNU ld 2.47.20260726, targeting Cortex-M3. The addresses are chosen for this exercise, not discovered hardware requirements. No Cortex-M processor executed the ELF.
Why can a successful link still produce the wrong state?
Linking can establish addresses without executing the copy or clear operations. The original reset-handler convention is:
extern unsigned char _data_load[], _data_start[], _data_end[];
extern unsigned char _bss_start[], _bss_end[];
unsigned char *src = _data_load;
for (unsigned char *dst = _data_start; dst < _data_end; ++dst)
*dst = *src++;
for (unsigned char *dst = _bss_start; dst < _bss_end; ++dst)
*dst = 0;
This is toolchain-specific startup notation following the GNU linker-symbol convention. It is not a portable demonstration of pointer comparison across ordinary C objects. A production implementation must follow its compiler, runtime and architecture contracts, including alignment and any compiler-generated helper calls.
Notice the exclusive end addresses. Copying through _data_end with <= writes one byte into the next region. Starting the source at _data_start instead of _data_load reads RAM rather than the flash image. Trace the first and last destination bytes to catch the off-by-one error; inspect the source address to catch the RAM-to-RAM copy.
NOLOAD also needs a careful explanation. It controls an output section’s loading treatment; writing it in the script does not execute a zeroing loop. Our unsigned counters require zero values before application use. The mechanism that establishes those values remains part of the startup contract. GNU ld output-section types
Diagnose missing initialization with four snapshots
The local executable model uses ordinary byte arrays, not fixed hardware addresses. It deliberately fills 24 RAM-model bytes with 0xA5, then independently enables the eight-byte copy and eight-byte clear. The final eight bytes are guards that must remain unchanged.
Predict the result before running it. With neither operation, all four modeled words remain 0xA5A5A5A5. Clearing only makes ticks and faults zero but leaves both configuration values poisoned. Copying only restores 3 and 5 but leaves the counters poisoned. Performing both establishes all four expected values.
Local execution: All four branches passed 48 assertions, including guard checks. The poison value is a test input, not a claim that SRAM powers up with this pattern. The model demonstrates the ranges and failure signatures; it does not reproduce reset hardware, memory faults or debugger initialization.
Explain a diagnosis by connecting the observation to a check that separates the competing causes. If configuration values are wrong but counters are zero, inspect the copy source, destination and length first. If configuration is correct but counters are not, inspect the clear range. These are useful hypotheses, not proof that no other component wrote the memory.
For a real target, stop before the initialization loops and again immediately before application entry. Compare the exact ELF, flashed image, linker symbols and memory contents. A breakpoint reached after a debugger has initialized RAM cannot establish what an independent power-on boot did.
Ask whether the failure follows a cold boot, a software reset or a debugger restart. A retained RAM bank or a deliberately preserved diagnostic area can change what you observe. If a region is intentionally retained, exclude it from the ordinary clear range and document when its contents become valid. Reusing a previous boot’s bytes accidentally is different from a specified retention protocol.
Also separate static-storage initialization from automatic local variables. The existence of a startup clear loop does not initialize every future stack frame. Conversely, a global counter without an explicit initializer is not permission to consume arbitrary RAM contents. Our all-zero-byte clear is demonstrated for unsigned integer counters on this target; do not generalize it into a statement that every C pointer or floating representation is universally all-zero bits.
Use the map to discuss memory capacity honestly
An interviewer may ask whether an eight-byte initialized object costs eight or sixteen bytes. In this specific layout it occupies eight bytes of runtime RAM and eight bytes of flash payload. File padding, ELF metadata and programming-container size are separate accounting questions.
Likewise, “all constants live in flash” is too strong. Storage placement follows the implementation and script, and some systems copy executable or read-only sections into RAM. Inspect the map and actual accesses before generalizing from const syntax.
Official linker fact: The MEMORY command describes available regions and allows sections to be assigned to them. It does not describe every runtime allocation or prove that a stack cannot overflow. GNU ld MEMORY command
Our script reserves an illustrative 1,024-byte gap below the stack top and rejects static data that crosses it. That is a build-time layout check, not a measured worst-case stack bound. Interrupt nesting, call depth, local arrays and dynamic allocation require additional analysis. Explain which evidence supports each number.
A useful follow-up is to deliberately violate one layout condition. Reduce the allowed code-image boundary or RAM capacity, relink and check that the intended assertion fails. Then restore the script and confirm the original map. That verifies the guard itself; repeatedly building the successful configuration would never show whether the failure path protects anything. Keep this build evidence separate from the runtime poison-and-guard model.
Why does a volatile counter still lose an increment?
Assume one privileged Cortex-M3 core, ordinary RAM, a main routine and one maskable interrupt handler. Each increments the same counter once. Assume each aligned word load and store is indivisible for this scenario; the complete read-modify-write still consists of multiple steps.
One possible schedule is: main reads zero; the ISR reads zero and writes one; main resumes and writes its previously computed one. Two increments occurred, but the stored count is one. The bug needs neither a torn word nor an exotic cache explanation.
Compiler documentation: GCC explains that accesses to ordinary memory are not ordered merely by accessing a volatile object. Its guidance does not turn volatile into an atomic operation or a general publication barrier. Our example was compiled with Clang, so GCC’s page is a comparison reference, not a report of the compiler we ran. GCC volatile guidance
The original Python interleaving model inserts the ISR before main’s load, between load and store, or after store. It produces final counts 2, 1, and 2. That demonstrates the lost-update schedule under stated assumptions. It is not a simulator of all Cortex-M interrupt entry behavior or an experiment establishing C thread-data-race semantics.
Repair the critical section without changing its caller’s mask
For the limited single-core scenario, one possible platform-specific repair is a short critical section that excludes the relevant interrupt:
uint32_t saved = __get_PRIMASK();
__disable_irq();
/* Use the supported compiler/platform access contract here. */
counter++;
__set_PRIMASK(saved);
Official architecture interface: CMSIS provides PRIMASK access and interrupt-disable helpers. Masking can prevent taking an interrupt while it becomes pending; it does not erase the event or exclude every exception. Privilege requirements and the actual core matter. Arm CMSIS core-register access
Preserve the saved value. Unconditionally enabling interrupts at the end can break a caller that entered with interrupts already masked. In our model, prior mask zero allows a pending ISR after restoration, reaching two; prior mask one remains one and leaves that ISR deferred. The latter is correct preservation of the caller’s state, not proof of a missing interrupt.

This snippet still needs the target’s documented compiler-ordering contract. Masking this core’s interrupt does not stop DMA, another core or an excluded exception, and it is not a blanket cache-coherency solution. If those actors share the data, explain their ownership and synchronization separately. Keep the protected work short and justify any latency cost rather than assuming it is harmless.
Practice five related questions with explicit boundaries
These PracHub records provide related practice, not predictions of an exact interview. Some cover hosted systems or different firmware environments; state the changed assumptions before reusing an answer.
| PracHub question | Practice focus | Boundary to explain |
|---|---|---|
| Explain C printf pointers and memory layout | Distinguish pointers, objects and representation | Hosted memory layout differs from our linked microcontroller image. |
| Reason About C Pointer and Array Declarations | Read declarations before dereferencing | Linker symbols add implementation conventions beyond ordinary arrays. |
| Reason About Endianness, Ring Buffers, and Stack Behavior in Firmware | Connect representation to firmware state | A ring-buffer protocol needs its own producer/consumer contract. |
| Explain BIOS and UEFI Firmware Mechanisms | Separate boot stages and responsibilities | PC firmware is not the Cortex-M3 reset sequence. |
| Explain and Diagnose a Segmentation Fault | Use evidence to isolate invalid accesses | Hosted fault reporting cannot be assumed on bare metal. |
Rehearse one answer in this order: expected value, linked location, responsible operation, counterexample, then verification. Extend into the Apple firmware debugging article for board and protocol diagnosis, or the RTOS scheduling article for scheduler interactions.
Start with the C pointer-declaration question: write down the type of each expression, explain one dereference aloud, then identify which assumptions change for linker-defined addresses. The goal is an explanation another engineer can check, with a clear boundary between source semantics, toolchain output and hardware behavior.
Sources and Further Reading
- GNU ld: output-section load addresses
- GNU ld: LOADADDR and other built-in functions
- GNU ld: MEMORY regions
- GNU ld: output-section types and NOLOAD
- GCC: volatile accesses and ordinary-memory ordering
- Arm CMSIS: Cortex-M core-register access
Technical references checked October 11, 2026. Local linking and models are identified above; no physical-board startup or interrupt-latency test was performed.
Comments (0)