aarch64: fix BYE cold-restart crash — PSCI SYSTEM_RESET via HVC, not SMC

Root cause of the aarch64 BYE cold-restart exception (present since at
least 2026-08-08, ESR_EL1=0x02000000/EC=0 "Unknown reason"), found via
live gdb single-stepping through the actual crash: arch_cold_reset()
issued PSCI SYSTEM_RESET via `smc #0`, but QEMU's aarch64 virt machine
booted with AAVMF (UEFI firmware, no genuine EL3/TrustZone secure
monitor) serves PSCI via HVC, not SMC -- nothing exists to answer an
SMC call, so it trapped as an illegal instruction straight into the
kernel's own exception handler. Not memory corruption, not a race --
a wrong conduit for this boot configuration.

Fix: smc #0 -> hvc #0. Function ID and calling convention unchanged.

Getting to this required first discovering that starkernel_kernel.elf
is not the binary that actually runs -- MONOLITHIC_BUILD links
kernel_main() directly into starkernel_loader.efi, a completely
separate, differently-linked build artifact. Every earlier gdb
breakpoint attempt this session failed because it used addresses from
the wrong file. Real addresses (UEFI-chosen ImageBase + linker-map
RVA) let gdb catch the crash live for the first time.

Verified: full aarch64 acceptance pass, 30/30 stress-campaign reps
PASS (unaffected -- this bug only manifested on BYE), and BYE now
exits cleanly with no exception for the first time in this
investigation.

Full writeup in FABRIC-2.md Section I.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
Robert Allan James
2026-08-18 18:24:26 -04:00
co-authored by Claude Sonnet 5
parent bc31916461
commit b24a5a6e25
12 changed files with 368452 additions and 78539 deletions
+12 -4
View File
@@ -127,13 +127,21 @@ void arch_cold_reset(void)
arch_disable_interrupts();
/* PSCI SYSTEM_RESET (function 0x84000009). SYSTEM_RESET takes no
* arguments and has no SMC64 variant defined by the PSCI spec -- only
* the SMC32 encoding is valid. The prior 0xC4000009 (SMC64 convention)
* was not a real PSCI function ID, causing an unhandled exception
* (EC=0, "Unknown reason") on the SMC call instead of a reset. */
* the SMC32 encoding is valid (fixed 2026-08-18, was 0xC4000009).
*
* FABRIC-2.md Section I, 2026-08-18: live gdb tracing (using the real
* UEFI-relocated runtime address, not the standalone kernel.elf's
* link-time address -- see that section for why those differ) proved
* the SMC call itself traps: PC does not fall through to the wfi loop
* below, it jumps straight into this kernel's own exception vector.
* This QEMU aarch64 boot (AAVMF UEFI firmware, no genuine EL3/TrustZone
* secure monitor) has nothing to answer an SMC -- PSCI here is served
* via HVC (hypervisor call, EL2) instead. Switched conduits; the
* function ID and calling convention are unchanged. */
__asm__ volatile (
"mov x0, #0x84000000\n"
"movk x0, #0x0009\n"
"smc #0\n"
"hvc #0\n"
::: "x0", "memory"
);
for (;;) __asm__ volatile ("wfi" ::: "memory");