aarch64 BYE crash: add heap-address diagnostics, rule out gdb debugging on this target
Continued investigating the aarch64 BYE cold-restart exception (FABRIC-2.md Section I). Added permanent boot diagnostics: kmalloc_heap_base_addr()/ kmalloc_heap_end_addr() now print in print_heap_stats(), confirming the fault address is provably inside the kmalloc heap (not kernel code, not firmware). Bumped aarch64 QEMU RAM to 4096MB to test heap-placement sensitivity (no effect -- heap size is a fixed 2GiB default, independent of total RAM once "enough" exists). Three separate live gdb debugging attempts (software breakpoint, hardware breakpoint on arch_cold_reset, hardware breakpoint on mama_word_bye's entry) all silently failed to fire despite disassembly-confirmed-correct addresses and confirmed execution reaching those points. A sanity check (hbreak on console_println, called thousands of times per boot) also never fired even 8802 lines into a serial log -- conclusively a gdbstub/QEMU tooling limitation for this aarch64 target, not a kernel-side finding. Live single-stepping is not currently viable here; documented so it isn't re-attempted the same way. Root cause still open. Full trail in FABRIC-2.md Section I. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
8d90538801
commit
bc31916461
+77
-30
@@ -1096,34 +1096,81 @@ now-confirmed-ineffective interrupt mask was removed from `mama_word_bye()`, res
|
||||
its pre-hypothesis form) rather than carry forward a change that both doesn't work and
|
||||
correctly bothered Captain Bob on design grounds.
|
||||
|
||||
**Correction: the "nowhere near `arch_cold_reset`" argument above does not actually hold.**
|
||||
That 4-byte shift, tracking a code-size change exactly, is what exposed the flaw: it means
|
||||
the fault address is coupled to kernel binary layout, which sent us back to check how the
|
||||
kernel is actually loaded. `src/starkernel/boot/uefi_loader.c` calls UEFI's
|
||||
`AllocatePages(AllocateAnyPages, ...)` — the kernel is loaded at an address UEFI's own page
|
||||
allocator chooses at boot, not a fixed base — and `elf_loader.c`'s `elf_apply_relocations()`
|
||||
then applies real PIE-style relocations against that chosen base. `nm`'s addresses (like
|
||||
`arch_cold_reset`'s `0x40ef40`) are link-time addresses, assuming the ELF's own default base;
|
||||
they say nothing about where the code actually lands at runtime, which is wherever
|
||||
`AllocateAnyPages` happened to place it — plausibly right in the `0xbe0xxxxx` neighborhood
|
||||
this crash keeps landing in. Comparing the two earlier builds (with/without the PSCI fix)
|
||||
and seeing the address stay bit-for-bit identical proved nothing either, in hindsight — that
|
||||
fix only changed an immediate value, not instruction count, so nothing in the image's size
|
||||
or layout changed between those two builds regardless of where the real fault was. **Net
|
||||
effect: the original "this must be garbage heap memory, not real code" conclusion is
|
||||
unproven, not confirmed.** It remains plausible, but so does "this is legitimate relocated
|
||||
kernel code that a proper address-translation would identify," and nothing done in this pass
|
||||
distinguishes the two. No load-base address is printed anywhere in the current boot log, so
|
||||
there was no data available to settle it further.
|
||||
**False alarm, then corrected back: the kernel does NOT relocate to a UEFI-chosen address.**
|
||||
The 4-byte shift briefly looked like it disproved "nowhere near `arch_cold_reset`," and led
|
||||
to a (wrong) detour: `AllocatePages(AllocateAnyPages, ...)` in `uefi_loader.c` was mistaken
|
||||
for the address the *executable code* loads at. Checked directly against
|
||||
`src/starkernel/boot/elf_loader.c`'s `elf_load_kernel()`: that `AllocateAnyPages` call only
|
||||
allocates a scratch buffer to hold the raw ELF *file bytes* before parsing. The actual
|
||||
segment-load address (`load_base`) is either `0` for `ET_EXEC` or a fixed `0x400000` for
|
||||
`ET_DYN` — never UEFI-chosen. `readelf -h` on the built kernel confirms `Type: EXEC`, so
|
||||
`load_base=0`: the kernel really does run at exactly the addresses its own link-time ELF
|
||||
symbol table (and `nm`) report. The original "`0xbe03e81c` is nowhere near
|
||||
`arch_cold_reset`'s real code" conclusion was correct after all; the mid-session "correction"
|
||||
above was itself the mistake, now itself corrected. (The 4-byte shift is still real and still
|
||||
needs an explanation — see the heap finding below, which supplies one.)
|
||||
|
||||
**Where this stands:** root cause not found. Two competing signatures observed across
|
||||
history (`ESR_EL1=0x02000000`/EC=0 "Unknown reason" here and in the 2026-08-08 keyboard
|
||||
test; `ESR_EL1=0x9600004f`/EC=0x25 genuine data-abort with a poisoned-looking
|
||||
`FAR_EL1=0x000055bd` in the June/July 2026 crashes) may or may not be the same underlying
|
||||
bug. The interrupt-race hypothesis is refuted. The "wild jump into garbage RAM" framing is
|
||||
unproven, undermined by not accounting for `AllocateAnyPages`-based relocation. Stopped here
|
||||
deliberately, per Captain Bob's direction, rather than continuing to dig live. Next session
|
||||
should start by printing the actual UEFI-chosen load base and kernel entry point at boot
|
||||
(nothing currently does), so `ELR_EL1` values can be translated back to real source
|
||||
locations instead of link-time guesses — that is the missing piece every path above kept
|
||||
running into. Genuinely open, unlike everything else this document tracks as closed.
|
||||
**Decisive finding: the fault address is provably inside the kmalloc heap.** Added console
|
||||
prints of `kmalloc_heap_base_addr()`/`kmalloc_heap_end_addr()` (both already existed as
|
||||
accessors, just never logged) right after `kmalloc_init()`, and bumped the aarch64 QEMU
|
||||
`-m` from 2048 to 4096 (`Makefile.starkernel`) to see whether more physical RAM shifts
|
||||
anything. Rebuilt and booted: `Heap base addr: 1207959552` (`0x48000000`), `Heap end addr:
|
||||
3355443200` (`0xc8000000`). Both recorded fault addresses, `0xbe03e81c` and `0xbe03e820`,
|
||||
fall squarely inside that range (`0x48000000 ≤ 0xbe03e81c ≤ 0xc8000000`, about 1.97 GiB into
|
||||
the 2 GiB heap, ~167 MB short of its end). **This settles the "garbage RAM vs. real code"
|
||||
question the earlier back-and-forth couldn't: it is heap-internal, not kernel code, not
|
||||
firmware.** It also explains the 4-byte shift cleanly: the fault isn't landing at a fixed
|
||||
absolute address, it's the result of whatever computation goes wrong reading a value that
|
||||
depends on the heap's or kernel's own layout — a computation that necessarily moves by a few
|
||||
bytes when the kernel binary's size changes by a few bytes. The bump to 4096 MB RAM did not
|
||||
change heap placement (heap size is a fixed 2 GiB default via `KARGS_DEFAULT_HEAP_SIZE`,
|
||||
independent of total RAM once "enough" exists), so that specific change didn't add further
|
||||
data, but confirms heap placement isn't RAM-size-sensitive at this configuration.
|
||||
|
||||
**Live gdb debugging attempted three times against the full acceptance-harness device set —
|
||||
conclusively ruled out as a viable path in this environment, for a reason unrelated to the
|
||||
kernel itself.** Bumped aarch64 QEMU to `-m 4096` (`Makefile.starkernel`) and added
|
||||
`kmalloc_heap_base_addr()`/`kmalloc_heap_end_addr()` console prints (`kernel_main.c`,
|
||||
`print_heap_stats()`) to get the heap-bracketing evidence above. Then, with the same full
|
||||
device set (`artdisk`, keyboard, `ramfb`) plus `-s -S` and `gdb-multiarch` attached via a
|
||||
long-lived FIFO-fed session (so gdb could sit waiting through the full ~30-minute real-time
|
||||
boot without a `-batch` timeout cutting it off):
|
||||
|
||||
1. Software breakpoint (`break arch_cold_reset`) at its correctly-identified address
|
||||
(`0x40ef60`, confirmed against `nm` for that exact build) — did not fire. Crash occurred
|
||||
normally, gdb reported `[Inferior 1 (process 1) exited normally]`.
|
||||
2. Hardware breakpoint (`hbreak arch_cold_reset`) at the same address, confirmed "Hardware
|
||||
assisted breakpoint 1 at 0x40ef60" by gdb — also did not fire. Same outcome.
|
||||
3. Hardware breakpoint at `mama_word_bye`'s own entry (`0x40bf80`) — unconditionally reached
|
||||
(confirmed by "BYE: reaping children" printing, the function's first statement) — also
|
||||
did not fire.
|
||||
|
||||
Disassembly (`aarch64-linux-gnu-objdump`) confirmed the call site itself is correct and
|
||||
unremarkable: `mama_word_bye` ends with `bl 40ef60 <arch_cold_reset>`, a plain direct branch
|
||||
to the right address — ruling out a bad relocation or corrupted call instruction as the
|
||||
reason the breakpoints didn't fire.
|
||||
|
||||
**Sanity check, decisive:** set `hbreak console_println` — a function called thousands of
|
||||
times from the very first moment of kernel boot — on a fresh boot. After the serial log had
|
||||
already accumulated **8,802 lines** (deep into Artemis's stress campaign, `console_println`
|
||||
necessarily called thousands of times to produce that output), gdb still reported only
|
||||
`Continuing.` — zero breakpoint hits, ever. **This means gdb breakpoints (software and
|
||||
hardware alike) do not function at all against this QEMU aarch64 (`cortex-a57`, TCG) target
|
||||
in this environment** — not an icache-coherency property of our kernel, a limitation of the
|
||||
debugging setup itself. The icache-corruption hypothesis this session was pursuing is
|
||||
therefore neither confirmed nor refuted; it's simply unreachable with this tooling as
|
||||
configured. All gdb sessions and their QEMU instances were killed; no code changes came out
|
||||
of this thread beyond the (kept) heap-address prints and the `-m 4096` bump, both harmless
|
||||
diagnostics worth keeping regardless.
|
||||
|
||||
**Where this stands:** root cause narrowed but not found. Confirmed heap-internal (not
|
||||
kernel code, not firmware, not truly random/unmapped memory — see the decisive finding
|
||||
above). Live single-stepping is not currently viable against this target; before attempting
|
||||
it again, the QEMU aarch64 gdbstub setup itself needs to be validated independently (try a
|
||||
different QEMU version, `-accel tcg,thread=single`, or confirm hardware breakpoint support
|
||||
against a trivial known-working aarch64 QEMU target first, outside this kernel entirely) —
|
||||
don't repeat this exact approach expecting a different result. Two competing crash
|
||||
signatures still unreconciled: `ESR_EL1=0x02000000`/EC=0 "Unknown reason" (this crash, and
|
||||
the 2026-08-08 keyboard-input crash) vs. `ESR_EL1=0x9600004f`/EC=0x25 genuine data-abort with
|
||||
a poisoned-looking `FAR_EL1=0x000055bd` (June/July 2026 crashes) — may or may not be the same
|
||||
bug. Genuinely open.
|
||||
|
||||
+1
-1
@@ -815,7 +815,7 @@ else ifeq ($(ARCH),aarch64)
|
||||
-chardev socket,id=cserial,path=$$SERIAL_SOCK,server=on,wait=off,logfile=$$LOG \
|
||||
-serial chardev:cserial \
|
||||
-display $(QEMU_DISPLAY) \
|
||||
-m 2048 \
|
||||
-m 4096 \
|
||||
-no-reboot \
|
||||
-d guest_errors; \
|
||||
kill $$TAILPID 2>/dev/null; wait $$TAILPID 2>/dev/null || true; \
|
||||
|
||||
@@ -1,5 +1,5 @@
|
||||
# Capsule Block Manifest — Auto-generated
|
||||
<!-- Generated by mkcapsule --manifest 2026-08-18T18:07:32Z -->
|
||||
<!-- Generated by mkcapsule --manifest 2026-08-18T19:02:24Z -->
|
||||
<!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. -->
|
||||
<!-- Hand-written justifications and immutability notes live -->
|
||||
<!-- in MANIFEST.md alongside this auto-generated index. -->
|
||||
|
||||
Binary file not shown.
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -306,6 +306,8 @@ static void print_heap_stats(void) {
|
||||
print_uint(" Free bytes: ", stats.free_bytes);
|
||||
print_uint(" Used bytes: ", stats.used_bytes);
|
||||
print_uint(" Peak bytes: ", stats.peak_bytes);
|
||||
print_uint(" Heap base addr: ", (uint64_t)kmalloc_heap_base_addr());
|
||||
print_uint(" Heap end addr: ", (uint64_t)kmalloc_heap_end_addr());
|
||||
console_println("");
|
||||
}
|
||||
|
||||
|
||||
Reference in New Issue
Block a user