Fix real WIREBIND crash: stale active-VM pointer dispatched after blocking read (FABRIC-3.md §XII.3)
The interpreter_enabled guard added in the previous commit (662ef44) was a
real but incomplete fix -- re-running the exact repro against it still
panicked (this time as a raw #PF page fault), proving something deeper
was wrong.
Root cause, found via targeted console_puts probes (not GDB --
starkernel_kernel.elf's symbols don't correspond to the actual running
starkernel_loader.efi binary for this monolithic build, same gotcha
already on record from the 2026-08-18 aarch64 investigation):
sk_repl_run()'s main loop captures `active` once, before calling
sk_console_readline(), which then blocks for the next full line. If the
identity `active` points at is killed while that read is still blocked,
the bailout meant to catch this (sk_console_identity_present()) only
checks a generic "is anyone attached" boolean, not "is the specific
identity active belonged to still attached" -- a fast detach of one
identity followed by attach of a different one never produces an
observable gap in that boolean, so the bailout never fires. The stale
`active`, now pointing at freed memory, gets dispatched into.
Fix: re-resolve `active` fresh from g_repl_active_vm immediately before
dispatch, right after sk_console_readline() returns. One line, no
registry lookup, no dereference of the stale pointer -- closes the race
regardless of whether the bailout catches it first.
Verified: rebuilt amd64 clean, reproduced the exact same attach/USE/
detach/attach/USE sequence against the fixed build -- clean switch, no
fault, exerciser runs correctly afterward.
FABRIC-3.md §XII.2 also corrected to stop claiming the interpreter_
enabled guard alone closed the crash -- it didn't, per the above.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
662ef44e59
commit
70421bdd43
@@ -1251,6 +1251,25 @@ void sk_repl_run(VM *vm)
|
||||
continue;
|
||||
}
|
||||
|
||||
/* Found live 2026-09-10: `active` was captured once, above, before
|
||||
* sk_console_readline() blocked for this line -- but that call can
|
||||
* block for an arbitrarily long time, during which the VM `active`
|
||||
* points at can be killed (WIREBIND detach) and its memory freed.
|
||||
* The n<0 bailout above is supposed to catch a logout mid-read, but
|
||||
* it only fires on sk_console_identity_present() -- a generic "is
|
||||
* ANYONE attached" boolean, not "is the specific identity `active`
|
||||
* belonged to still attached" -- so a fast detach-then-reattach of
|
||||
* a *different* identity while this call was blocked (n still 0)
|
||||
* never trips it: presence reads true throughout, no gap is ever
|
||||
* observed. The stale `active` then gets dispatched into freed
|
||||
* memory. Confirmed live via targeted probes: g_repl_active_vm is
|
||||
* correctly reset to NULL by the kill/teardown path the moment it
|
||||
* happens, but this loop iteration's *local* `active` was already
|
||||
* snapshotted and never re-read. Re-resolve fresh from the global
|
||||
* right before dispatch -- cheap, and closes the race regardless
|
||||
* of whether the bailout above catches it first. */
|
||||
active = g_repl_active_vm ? g_repl_active_vm : vm;
|
||||
|
||||
sk_repl_dispatch_line(active, input);
|
||||
|
||||
/* ABORT stops mid-line but leaves the flag set for the caller to
|
||||
|
||||
Reference in New Issue
Block a user