Commit Graph
4 Commits
Author SHA1 Message Date
Robert Allan JamesandClaude Sonnet 5 1a263555e2 Fix pathological migration scan that stalled WIREBIND identity attach
Root cause (found by a fresh subagent after an extended live-debugging
investigation into "identity 00 attaches slowly/stalls when Zuse never
attached first this boot"): blk_migration_idle_check() was generalized
earlier today to walk every attached device slot uniformly instead of
hardcoding first_disk_slot() (Artemis's own disk). But its per-slot scan
can only early-exit once it finds a devblock that is BOTH "hot" (claimed
and worn) AND "free" -- and a just-attached, never-claimed USB identity
drive can never satisfy the "hot" half by design (claiming only ever
happens via blk_firsttouch_claim(), which only ever targets
first_disk_slot()). So the scan ran to completion -- the drive's entire
~16,000 devblocks, mostly cache misses over slow emulated USB/BOT --
every single idle tick, forever, blocking sk_repl_idle() (and therefore
the console and the storage-attach message round-trip) each time.

Fix, in src/block_subsystem.c: a new has_ever_claimed flag on
blk_dev_slot_t (set in blk_set_meta(), the single choke point every
BLK_FLAG_CLAIMED transition passes through) skips the scan entirely,
O(1), for any slot nothing has ever claimed -- the common case for a
freshly-attached drive. A new migration_scan_lbn resume cursor bounds
*any* slot's per-tick cost to MIGRATION_SCAN_BUDGET (256) devblocks
examined, picking up where the previous tick left off instead of
restarting from start_lbn every time -- restores this function's own
documented "coarse cadence, cheap early-exit" design intent for every
device, not just the one it used to hardcode.

Also along the way (kept, all real improvements, verified live):
- src/starkernel/usb/xhci.c: xhci_wait_bit()/xhci_bot_wait_for_idle()
  had zero yield hints in their MMIO-polling loops; added arch_relax()
  to both (matches virtio_blk.c below) -- a tight loop of nothing but
  MMIO reads can starve TCG's own host-side timer injection under QEMU.
- src/starkernel/virtio/virtio_blk.c: vblk_io()'s spin bound was 33M
  iterations with zero logging on timeout; a single real (still not
  fully root-caused) timeout cost 31+ minutes of CPU before this was
  caught. Reduced to 1M and added a log line naming the failing sector,
  turning a silent, effectively-unbounded stall into a fast, loud
  failure -- callers already tolerate BLKIO_EIO.
- src/starkernel/repl.c: blk_migration_idle_check() deferred for any
  idle tick where a storage-attach message round-trip is still pending,
  to keep the two block-subsystem-touching paths from interleaving; the
  existing MSG-TICK pump now checks the target VM's own dictionary for
  MSG-TICK (vm_find_word + acl_allow) before VM-EXECing into it, instead
  of spamming "VM-EXEC: ERROR" every tick for a VM that doesn't have --
  or isn't allowed -- the word (e.g. a FORTH-79/83-locked-down identity);
  fixed a real -Wunused-variable build break in the EMERGENCY_CONSOLE_
  ENABLED=1 path (unused today, but needed live to reproduce this bug
  with no identity attached at all).
- src/starkernel/capsule/capsule_mint.c: dropped the dead
  S" common:messaging.4th" EXEC / MSG-CD-INIT lines from
  MINT_RESTRICTED_PERSONALITY -- a FORTH-79/83-locked-down identity has
  no legitimate use for a messaging vocabulary it can never call.

Status: the pathological CPU-climbing scan is confirmed fixed (verified
live: CPU stays flat across an extended run instead of climbing without
bound). The WIREBIND storage-attach message round-trip still does not
complete promptly in the "Zuse never attached, other identity attaches
first" scenario -- a separate, still-open issue in the message-delivery
path itself, not the scan. Tracked as follow-on work.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014Ec88YKxxhZGG1RNnune78
2026-09-09 12:08:03 -04:00
Robert Allan JamesandClaude Sonnet 5 0bae928aad Mint the 8 identity thumbdrives (bob, 00-06) with real MINT data
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
zuse-thumb-ident.img was already minted; bob-thumb-ident.img (Captain
Bob / rajames / rajames440@gmail.com) and 00-thumb-ident.img through
06-thumb-ident.img (full_name=username=the number, no email/phone) are
now real minted identities, not blank images.

Done via the actual kernel MINT word at a live console (Zuse
authenticated, drives hot-plugged one at a time via the QEMU HMP
monitor) -- no Python orchestration, per direct instruction: a small
StarForth word, MINT-NUM, wraps MINT for the 7 uniform numeric
identities, driven by plain bash + socat against the monitor/serial
unix sockets.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 22:42:41 -04:00
Robert Allan JamesandClaude Sonnet 5 9e81de3f43 xHCI/BOT driver: genuine multi-device support (FABRIC-3.md §VII)
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Per-slot registry (xhci_msc_slot_t/dev->msc_slots, sized off the
controller's own reported max_slots) replaces the single-device scalar
fields the driver carried since Milestones 2e-2h. Boot-time port scan no
longer stops at the first connected device; a connect/disconnect that
arrives while the Command Ring is busy is now queued and drained instead
of dropped. blkio_usb.c and repl.c's own single-device state (device
descriptor buffers, blkio_dev_t, attach bookkeeping) became per-slot
registries the same way.

Live multi-device testing (not just compiling) surfaced a second, more
severe bug outside the original plan: transfer_purpose and next_action
were also single scalars shared across the whole controller. Two devices
enumerating concurrently could have one's completion silently overwrite
the other's still-outstanding one, permanently stalling it with no error.
Fixed by moving both per-slot and, critically, reading the Transfer Event
TRB's own real Slot ID field instead of trusting external bookkeeping.

Verified live, all three architectures, mandatory clean-qemu acceptance:
existing single-device path unchanged, and two devices attached
simultaneously (amd64) both progress independently through enumeration
without corrupting or stalling each other.

Also in this pass (implemented and verified in earlier turns this
session, committed together per direct instruction):
- Headless-until-login console policy: no prompt/banner until a real
  identity logs in via an attached thumbdrive (WIREBIND or Zuse, neither
  special), reusing EMERGENCY_CONSOLE_ENABLED as the debug/recovery
  escape hatch (now default-off).
- KILL/g_repl_active_vm dangling-pointer fix: killing the VM the console
  is currently USE'd onto now detaches back to Hera first, matching the
  existing EJECT/UNCLEAN precedent.

FABRIC-3.md §VII/§VIII carry full closure notes for all three.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 22:14:14 -04:00
Robert Allan JamesandClaude Sonnet 5 c3db963164 FABRIC-3.md §VII: xHCI/BOT single-device architecture — plan only, halted
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Documents the full scope of the driver's single-device-at-a-time state
(connect/enumerate, control-transfer, and BOT state machines), the
persistent-vs-in-flight distinction that bounds the fix, the concurrency
target, and a numbered punch list. No driver code has been touched —
implementation is explicitly held pending go-ahead.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018EjXFo7mPXjUMjfJeuUUz4
2026-09-05 20:45:20 -04:00