3(arch) x 9(identity) x 3(rep) randomized full-factorial std79 DoE: 81/81 correct (FABRIC-3.md §XV)
Formal successor to §XII's ad hoc exerciser campaign, requested as a genuine randomized full-factorial design matching this project's own DoE methodology (capsules/doe.4th's Fisher-Yates shuffle), and written entirely in FORTH per explicit request -- not host-orchestrated shell scripting. experiments/std79-doe/std79-doe.fth: builds a 27-cell (9 identity x 3 replicate) run matrix, Fisher-Yates shuffles it with a fixed seed (matching doe.4th's own default), then dispatches each of the 24 exerciser test cases directly into the target identity's own live VM via VM-EXEC -- no console USE redirection, no per-trial host interaction. The zuse case runs as directly-compiled native code (RUN-TEST-NATIVE) rather than VM-EXEC targeting "Hera" herself: VM-EXEC's own vm_state_push/pop only saves rsp/exit_colon/ecw_nesting, not input_buffer/input_pos, so a self-targeting call while this capsule's own vm_interpret call is still mid-line would risk exactly the class of bug the idle-tick reentrancy guards exist for. amd64: all 27 trials ran with all 9 identities simultaneously live in one boot -- clean, zero mismatches. aarch64: hit a real, uninvestigated bug attaching all 9 simultaneously -- the 6th live VM's birth stalled for 90+ minutes at 100%+ CPU with a Stadium COOL-dispatch counter already at 335,000+ cycles, versus tens of thousands at the same checkpoint for earlier identities. Not root-caused here (flagged in FABRIC-3.md for later); worked around by splitting into 3 boots of 3 simultaneously-live identities each, every boot sharing the same seed so the master 27-slot shuffle is identical, filtered per boot by a new ACTIVE-LO/ACTIVE-HI range (EXEC-STD79-DOE's signature: seed lo hi -- ). run_id is always the slot's true position in the master shuffle, so trial order stays comparable across boots -- standard DoE blocking. riscv64: same 3-boot pattern, clean. Grand total: 81/81 trials correct, 0 mismatches, across all three architectures, all nine identities, all three replicates. Also flagged (FABRIC-3.md §XIV, not fixed): attaching several WIREBIND identities near-simultaneously (whether via rapid hotplug or all present from boot) causes the kernel to silently detect only some of them -- confirmed at the host/QMP level that every device was genuinely present. Worked around throughout this campaign by attaching one identity at a time with confirmed waits; real hardware hotplug could hit the same gap, so it's a genuine robustness concern, not just a test-harness inconvenience. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
70db955ac9
commit
e2abc56306
+118
@@ -2424,3 +2424,121 @@ precedent. Every other probe added during this investigation (`block_words.c`,
|
||||
`block_subsystem.c`, `capsule_birth.c`, `vm_core.c`, `mama_forth_words.c`, and the rest of
|
||||
`repl.c`'s temporary brackets) was reverted after use, per the established "write, run once,
|
||||
capture, revert" discipline.
|
||||
|
||||
## XIV. Concurrent WIREBIND attach detection gap — found live, NOT fixed, flagged for later
|
||||
2026-09-11
|
||||
|
||||
Found while setting up the std79 DoE (§XV): attaching several USB identity thumbdrives to
|
||||
distinct xHCI ports at (or near) the same real-time instant reliably causes the kernel to detect
|
||||
and birth only *some* of them -- the rest never trigger `WIREBIND: <id> attached` at all, even
|
||||
after minutes of continued idle-tick servicing. Confirmed at the host level via QMP
|
||||
`query-block` that every device genuinely was present and recognized by QEMU itself (this is not
|
||||
a QEMU-side failure) -- the gap is entirely in the guest kernel's own per-tick attach-pending
|
||||
detection, which apparently only reliably surfaces one (or a small handful) of several
|
||||
simultaneously-pending port events, silently dropping the rest rather than queuing or retrying
|
||||
them.
|
||||
|
||||
**Two reproductions, same result:**
|
||||
1. Eight `device_add` calls fired in rapid QMP succession (~1s apart) after boot: only 2 of 8
|
||||
ever attached.
|
||||
2. All eight identity drives present from the very first instant of boot itself (via
|
||||
`QEMU_EXTRA`, matching how Zuse's own drive is already a boot-time device, not a hotplug one)
|
||||
-- the same gap: only 1 of 8 was ever detected, even though this is arguably the *most*
|
||||
realistic "hardware already plugged in at power-on" scenario a real deployment would hit, not
|
||||
an artificial hotplug stress pattern.
|
||||
|
||||
**What reliably works:** attaching one identity at a time, waiting for its own confirmed
|
||||
`WIREBIND: <id> attached` line before sending the next `device_add` -- exactly the pattern
|
||||
already established for the §XII.4 exerciser campaign. This avoids the gap entirely by never
|
||||
presenting the kernel with more than one newly-pending port at once.
|
||||
|
||||
**Not root-caused or fixed here** -- out of scope for the DoE task that surfaced it, and
|
||||
flagged by the user as a real defect worth fixing "eventually," not immediately: real hardware
|
||||
hotplug events can genuinely arrive at the same instant (e.g. a USB hub with several drives
|
||||
already inserted, connected all at once), so this gap is a real robustness concern for actual
|
||||
identity thumbdrive usage, not just a test-harness inconvenience. Worth investigating alongside
|
||||
the existing xHCI multi-device history (§VII, and the 2026-09-05/06/07 fixes referenced there)
|
||||
-- possibly a recurrence or a related gap in the same per-slot attach-pending scan, given the
|
||||
similar shape ("several ports become pending at once, only some get serviced").
|
||||
|
||||
## XV. 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial DoE — self-contained
|
||||
FORTH capsule, replacing the ad hoc exerciser-feeding process (2026-09-11)
|
||||
|
||||
Requested as a formal successor to §XII's ad hoc exerciser campaign: a genuine randomized
|
||||
full-factorial design (architecture × identity, 3 replicates each) matching this project's own
|
||||
DoE methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle), rather than convenience
|
||||
batching. Explicitly requested "all written in FORTH" -- not host-orchestrated shell scripting.
|
||||
|
||||
**Design:** `experiments/std79-doe/std79-doe.fth`. All 9 identities (`zuse` plus the 8 WIREBIND
|
||||
users) attached *once*, simultaneously live for the whole run -- not per-trial hotplug (see §XIV:
|
||||
attaching many devices near-simultaneously hits a real kernel detection gap, worked around here
|
||||
by attaching them one at a time with confirmed waits before the DoE capsule runs at all, so the
|
||||
run itself never touches the host/QMP layer). The capsule itself is architecture-fixed per boot
|
||||
(matching the existing hard convention -- architecture can't be experimentally randomized without
|
||||
separate binaries/boots); the *identity* × *replicate* factor (27 cells) is Fisher-Yates
|
||||
shuffled with a fixed seed (`12345`, matching `doe.4th`'s own default seed for methodological
|
||||
consistency) and the same permutation is reused on all three architectures for the cleanest
|
||||
possible cross-arch comparison.
|
||||
|
||||
**Driving mechanism:** `VM-EXEC ( cmd-caddr cmd-u vm-name-caddr vm-name-u -- )` (already used
|
||||
internally by the MSG-TICK pump) dispatches each of the 24 exerciser test lines directly into
|
||||
the target identity's own live VM, synchronously, tagged with that VM's own console name --
|
||||
no `USE`/console redirection needed at all, and no per-trial host interaction. One case
|
||||
(`zuse`, factor level 0) is deliberately **not** dispatched via `VM-EXEC` targeting `"Hera"`
|
||||
herself: `VM-EXEC`'s own `vm_state_push`/`vm_state_pop` (`mama_forth_words.c`) only save
|
||||
`rsp`/`exit_colon`/`ecw_nesting`, not `input_buffer`/`input_pos` -- a self-targeting call while
|
||||
the DoE capsule's own `vm_interpret` call is still mid-line would risk exactly the class of bug
|
||||
the idle-tick reentrancy guards (`g_mama_interpreting`, `repl.c`) exist for, just uncovered by
|
||||
`VM-EXEC`'s narrower state stack. Sidestepped entirely: the same 24 test cases are duplicated as
|
||||
directly-compiled FORTH (`RUN-TEST-NATIVE`), called in-process with no `vm_interpret` recursion
|
||||
at all, for the `zuse` case only.
|
||||
|
||||
**amd64 result: clean, verified programmatically, all 9 identities simultaneously live for the
|
||||
whole 27-trial run.** run_id 0-26, each identity appearing exactly 3 times (confirmed via
|
||||
frequency count). All 27 trials' 34 printed result values diffed against the first trial as
|
||||
baseline: **zero mismatches**. `T14`-equivalent (`-123 456 M* SWAP D.`) correctly prints `-56088`
|
||||
in every trial -- confirms §XII.4 bug 1's fix (`M*`, commit `9a09949`) holds under this DoE's own
|
||||
`VM-EXEC`/native dispatch, not just the original exerciser's raw-text-feed path.
|
||||
|
||||
**aarch64: a second real bug found, worked around, not root-caused.** Attaching all 9 identities
|
||||
simultaneously (matching amd64's approach) hit a severe wall partway through: identities
|
||||
`bob`/`rajames`, `00`-`03` attached normally (each ~1-2 minutes), but the 6th simultaneously-live
|
||||
VM (`04`) never completed -- QEMU stayed at 100%+ CPU, genuinely computing (confirmed via
|
||||
`ps` CPU-time climbing steadily, not a deadlock), but produced next to no new console output for
|
||||
over 90 minutes before the attempt was abandoned. The `Stadium: dispatch cell=NNNNN
|
||||
behaviour=COOL` background patron-churn counter (visible in the log immediately before the
|
||||
stall) had already reached **335,000+** cycles at that point -- tens of thousands more than the
|
||||
same checkpoint shows for earlier, successful identities -- suggesting a genuine algorithmic
|
||||
blowup (likely superlinear in live-VM count) in the Stadium COOL dispatch path specifically on
|
||||
aarch64, not just "TCG is slower." amd64 attached all 8 non-`zuse` identities with no such cliff
|
||||
at any point. **Not root-caused here** -- flagged as a real defect (see the discussion below);
|
||||
this session's priority was completing the requested DoE, not chasing a new performance bug.
|
||||
|
||||
**Design revised to work around it:** each architecture other than amd64 runs as **3 boots of 3
|
||||
simultaneously-live identities each** (`zuse`+`rajames`+`00`, `01`+`02`+`03`, `04`+`05`+`06`),
|
||||
comfortably below whatever threshold triggers the aarch64 stall. `EXEC-STD79-DOE`'s signature
|
||||
changed from `( seed -- )` to `( seed lo hi -- )`: every boot still runs `INIT-MATRIX`/
|
||||
`SHUFFLE-MATRIX` with the *same* seed, producing the *identical* 27-slot master permutation every
|
||||
time -- only a new `ACTIVE-LO`/`ACTIVE-HI` range filter (`ACTIVE?`) differs per boot, so each
|
||||
boot walks the *full* master order and silently skips any slot outside its own live subset. The
|
||||
recorded `run_id` is always the slot's true position in the master shuffle (the loop index `I`
|
||||
itself, never a separately-incremented counter), so trial order stays meaningful and directly
|
||||
comparable across boots -- textbook DoE "blocking": randomized within each block, blocked across
|
||||
blocks by a practical constraint, not confounded with the treatment.
|
||||
|
||||
**aarch64 and riscv64 results (3 boots each): clean, verified programmatically.** All 27
|
||||
run_ids (0-26) present exactly once across the 3 boots for each architecture, each of the 9
|
||||
identities appearing exactly 3 times. Zero mismatches against baseline on either architecture.
|
||||
|
||||
**Grand total across all three architectures: 81/81 trials correct, 0 mismatches.** Every
|
||||
identity, every replicate, every architecture, byte-identical to the single established
|
||||
baseline. Raw logs and this verification are in `experiments/std79-doe/results-20260911/`.
|
||||
|
||||
**Follow-on item, not investigated further here:** the aarch64 Stadium/COOL dispatch stall.
|
||||
Worth a dedicated investigation later -- candidates to check first: whether `Stadium:
|
||||
dispatch cell=NNNNN` count genuinely scales worse than linearly with live VM count (the visible
|
||||
symptom), and whether it's aarch64-specific in a real sense or just the first architecture where
|
||||
TCG overhead pushed an existing superlinear cost over a wall-clock threshold nobody had hit
|
||||
before (amd64's own per-tick cost may have the same underlying shape, just cheap enough per-tick
|
||||
under native-ish TCG codegen that it never became visible in any campaign run so far).
|
||||
|
||||
|
||||
Reference in New Issue
Block a user