9bcc70647b39439b362bd740260db1328145709f
1
Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
1a263555e2 |
Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging investigation into "identity 00 attaches slowly/stalls when Zuse never attached first this boot"): blk_migration_idle_check() was generalized earlier today to walk every attached device slot uniformly instead of hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan can only early-exit once it finds a devblock that is BOTH "hot" (claimed and worn) AND "free" -- and a just-attached, never-claimed USB identity drive can never satisfy the "hot" half by design (claiming only ever happens via blk_firsttouch_claim(), which only ever targets first_disk_slot()). So the scan ran to completion -- the drive's entire ~16,000 devblocks, mostly cache misses over slow emulated USB/BOT -- every single idle tick, forever, blocking sk_repl_idle() (and therefore the console and the storage-attach message round-trip) each time. Fix, in src/block_subsystem.c: a new has_ever_claimed flag on blk_dev_slot_t (set in blk_set_meta(), the single choke point every BLK_FLAG_CLAIMED transition passes through) skips the scan entirely, O(1), for any slot nothing has ever claimed -- the common case for a freshly-attached drive. A new migration_scan_lbn resume cursor bounds *any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks examined, picking up where the previous tick left off instead of restarting from start_lbn every time -- restores this function's own documented "coarse cadence, cheap early-exit" design intent for every device, not just the one it used to hardcode. Also along the way (kept, all real improvements, verified live): - src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle() had zero yield hints in their MMIO-polling loops; added arch_relax() to both (matches virtio_blk.c below) -- a tight loop of nothing but MMIO reads can starve TCG's own host-side timer injection under QEMU. - src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M iterations with zero logging on timeout; a single real (still not fully root-caused) timeout cost 31+ minutes of CPU before this was caught. Reduced to 1M and added a log line naming the failing sector, turning a silent, effectively-unbounded stall into a fast, loud failure -- callers already tolerate BLKIO_EIO. - src/starkernel/repl.c: blk_migration_idle_check() deferred for any idle tick where a storage-attach message round-trip is still pending, to keep the two block-subsystem-touching paths from interleaving; the existing MSG-TICK pump now checks the target VM's own dictionary for MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have -- or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity); fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_ ENABLED=1 path (unused today, but needed live to reproduce this bug with no identity attached at all). - src/starkernel/capsule/capsule_mint.c: dropped the dead S" common:messaging.4th" EXEC / MSG-CD-INIT lines from MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has no legitimate use for a messaging vocabulary it can never call. Status: the pathological CPU-climbing scan is confirmed fixed (verified live: CPU stays flat across an extended run instead of climbing without bound). The WIREBIND storage-attach message round-trip still does not complete promptly in the "Zuse never attached, other identity attaches first" scenario -- a separate, still-open issue in the message-delivery path itself, not the scan. Tracked as follow-on work. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78 |