Fix pathological migration scan that stalled WIREBIND identity attach

Root cause (found by a fresh subagent after an extended live-debugging
investigation into "identity 00 attaches slowly/stalls when Zuse never
attached first this boot"): blk_migration_idle_check() was generalized
earlier today to walk every attached device slot uniformly instead of
hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan
can only early-exit once it finds a devblock that is BOTH "hot" (claimed
and worn) AND "free" -- and a just-attached, never-claimed USB identity
drive can never satisfy the "hot" half by design (claiming only ever
happens via blk_firsttouch_claim(), which only ever targets
first_disk_slot()). So the scan ran to completion -- the drive's entire
~16,000 devblocks, mostly cache misses over slow emulated USB/BOT --
every single idle tick, forever, blocking sk_repl_idle() (and therefore
the console and the storage-attach message round-trip) each time.

Fix, in src/block_subsystem.c: a new has_ever_claimed flag on
blk_dev_slot_t (set in blk_set_meta(), the single choke point every
BLK_FLAG_CLAIMED transition passes through) skips the scan entirely,
O(1), for any slot nothing has ever claimed -- the common case for a
freshly-attached drive. A new migration_scan_lbn resume cursor bounds
*any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks
examined, picking up where the previous tick left off instead of
restarting from start_lbn every time -- restores this function's own
documented "coarse cadence, cheap early-exit" design intent for every
device, not just the one it used to hardcode.

Also along the way (kept, all real improvements, verified live):
- src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle()
  had zero yield hints in their MMIO-polling loops; added arch_relax()
  to both (matches virtio_blk.c below) -- a tight loop of nothing but
  MMIO reads can starve TCG's own host-side timer injection under QEMU.
- src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M
  iterations with zero logging on timeout; a single real (still not
  fully root-caused) timeout cost 31+ minutes of CPU before this was
  caught. Reduced to 1M and added a log line naming the failing sector,
  turning a silent, effectively-unbounded stall into a fast, loud
  failure -- callers already tolerate BLKIO_EIO.
- src/starkernel/repl.c: blk_migration_idle_check() deferred for any
  idle tick where a storage-attach message round-trip is still pending,
  to keep the two block-subsystem-touching paths from interleaving; the
  existing MSG-TICK pump now checks the target VM's own dictionary for
  MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead
  of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have --
  or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity);
  fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_
  ENABLED=1 path (unused today, but needed live to reproduce this bug
  with no identity attached at all).
- src/starkernel/capsule/capsule_mint.c: dropped the dead
  S" common:messaging.4th" EXEC / MSG-CD-INIT lines from
  MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has
  no legitimate use for a messaging vocabulary it can never call.

Status: the pathological CPU-climbing scan is confirmed fixed (verified
live: CPU stays flat across an extended run instead of climbing without
bound). The WIREBIND storage-attach message round-trip still does not
complete promptly in the "Zuse never attached, other identity attaches
first" scenario -- a separate, still-open issue in the message-delivery
path itself, not the scan. Tracked as follow-on work.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
This commit is contained in:
Robert Allan James
2026-09-09 12:08:03 -04:00
co-authored by Claude Sonnet 5
parent 63b8b3bc29
commit 1a263555e2
30 changed files with 153358 additions and 33 deletions
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff