Root cause (found by a fresh subagent after an extended live-debugging
investigation into "identity 00 attaches slowly/stalls when Zuse never
attached first this boot"): blk_migration_idle_check() was generalized
earlier today to walk every attached device slot uniformly instead of
hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan
can only early-exit once it finds a devblock that is BOTH "hot" (claimed
and worn) AND "free" -- and a just-attached, never-claimed USB identity
drive can never satisfy the "hot" half by design (claiming only ever
happens via blk_firsttouch_claim(), which only ever targets
first_disk_slot()). So the scan ran to completion -- the drive's entire
~16,000 devblocks, mostly cache misses over slow emulated USB/BOT --
every single idle tick, forever, blocking sk_repl_idle() (and therefore
the console and the storage-attach message round-trip) each time.
Fix, in src/block_subsystem.c: a new has_ever_claimed flag on
blk_dev_slot_t (set in blk_set_meta(), the single choke point every
BLK_FLAG_CLAIMED transition passes through) skips the scan entirely,
O(1), for any slot nothing has ever claimed -- the common case for a
freshly-attached drive. A new migration_scan_lbn resume cursor bounds
*any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks
examined, picking up where the previous tick left off instead of
restarting from start_lbn every time -- restores this function's own
documented "coarse cadence, cheap early-exit" design intent for every
device, not just the one it used to hardcode.
Also along the way (kept, all real improvements, verified live):
- src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle()
had zero yield hints in their MMIO-polling loops; added arch_relax()
to both (matches virtio_blk.c below) -- a tight loop of nothing but
MMIO reads can starve TCG's own host-side timer injection under QEMU.
- src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M
iterations with zero logging on timeout; a single real (still not
fully root-caused) timeout cost 31+ minutes of CPU before this was
caught. Reduced to 1M and added a log line naming the failing sector,
turning a silent, effectively-unbounded stall into a fast, loud
failure -- callers already tolerate BLKIO_EIO.
- src/starkernel/repl.c: blk_migration_idle_check() deferred for any
idle tick where a storage-attach message round-trip is still pending,
to keep the two block-subsystem-touching paths from interleaving; the
existing MSG-TICK pump now checks the target VM's own dictionary for
MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead
of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have --
or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity);
fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_
ENABLED=1 path (unused today, but needed live to reproduce this bug
with no identity attached at all).
- src/starkernel/capsule/capsule_mint.c: dropped the dead
S" common:messaging.4th" EXEC / MSG-CD-INIT lines from
MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has
no legitimate use for a messaging vocabulary it can never call.
Status: the pathological CPU-climbing scan is confirmed fixed (verified
live: CPU stays flat across an extended run instead of climbing without
bound). The WIREBIND storage-attach message round-trip still does not
complete promptly in the "Zuse never attached, other identity attaches
first" scenario -- a separate, still-open issue in the message-delivery
path itself, not the scan. Tracked as follow-on work.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78