Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging investigation into "identity 00 attaches slowly/stalls when Zuse never attached first this boot"): blk_migration_idle_check() was generalized earlier today to walk every attached device slot uniformly instead of hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan can only early-exit once it finds a devblock that is BOTH "hot" (claimed and worn) AND "free" -- and a just-attached, never-claimed USB identity drive can never satisfy the "hot" half by design (claiming only ever happens via blk_firsttouch_claim(), which only ever targets first_disk_slot()). So the scan ran to completion -- the drive's entire ~16,000 devblocks, mostly cache misses over slow emulated USB/BOT -- every single idle tick, forever, blocking sk_repl_idle() (and therefore the console and the storage-attach message round-trip) each time. Fix, in src/block_subsystem.c: a new has_ever_claimed flag on blk_dev_slot_t (set in blk_set_meta(), the single choke point every BLK_FLAG_CLAIMED transition passes through) skips the scan entirely, O(1), for any slot nothing has ever claimed -- the common case for a freshly-attached drive. A new migration_scan_lbn resume cursor bounds *any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks examined, picking up where the previous tick left off instead of restarting from start_lbn every time -- restores this function's own documented "coarse cadence, cheap early-exit" design intent for every device, not just the one it used to hardcode. Also along the way (kept, all real improvements, verified live): - src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle() had zero yield hints in their MMIO-polling loops; added arch_relax() to both (matches virtio_blk.c below) -- a tight loop of nothing but MMIO reads can starve TCG's own host-side timer injection under QEMU. - src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M iterations with zero logging on timeout; a single real (still not fully root-caused) timeout cost 31+ minutes of CPU before this was caught. Reduced to 1M and added a log line naming the failing sector, turning a silent, effectively-unbounded stall into a fast, loud failure -- callers already tolerate BLKIO_EIO. - src/starkernel/repl.c: blk_migration_idle_check() deferred for any idle tick where a storage-attach message round-trip is still pending, to keep the two block-subsystem-touching paths from interleaving; the existing MSG-TICK pump now checks the target VM's own dictionary for MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have -- or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity); fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_ ENABLED=1 path (unused today, but needed live to reproduce this bug with no identity attached at all). - src/starkernel/capsule/capsule_mint.c: dropped the dead S" common:messaging.4th" EXEC / MSG-CD-INIT lines from MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has no legitimate use for a messaging vocabulary it can never call. Status: the pathological CPU-climbing scan is confirmed fixed (verified live: CPU stays flat across an extended run instead of climbing without bound). The WIREBIND storage-attach message round-trip still does not complete promptly in the "Zuse never attached, other identity attaches first" scenario -- a separate, still-open issue in the message-delivery path itself, not the scan. Tracked as follow-on work. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
63b8b3bc29
commit
1a263555e2
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user