3(arch) x 9(identity) x 3(rep) randomized full-factorial std79 DoE: 81/81 correct (FABRIC-3.md §XV)
Formal successor to §XII's ad hoc exerciser campaign, requested as a genuine randomized full-factorial design matching this project's own DoE methodology (capsules/doe.4th's Fisher-Yates shuffle), and written entirely in FORTH per explicit request -- not host-orchestrated shell scripting. experiments/std79-doe/std79-doe.fth: builds a 27-cell (9 identity x 3 replicate) run matrix, Fisher-Yates shuffles it with a fixed seed (matching doe.4th's own default), then dispatches each of the 24 exerciser test cases directly into the target identity's own live VM via VM-EXEC -- no console USE redirection, no per-trial host interaction. The zuse case runs as directly-compiled native code (RUN-TEST-NATIVE) rather than VM-EXEC targeting "Hera" herself: VM-EXEC's own vm_state_push/pop only saves rsp/exit_colon/ecw_nesting, not input_buffer/input_pos, so a self-targeting call while this capsule's own vm_interpret call is still mid-line would risk exactly the class of bug the idle-tick reentrancy guards exist for. amd64: all 27 trials ran with all 9 identities simultaneously live in one boot -- clean, zero mismatches. aarch64: hit a real, uninvestigated bug attaching all 9 simultaneously -- the 6th live VM's birth stalled for 90+ minutes at 100%+ CPU with a Stadium COOL-dispatch counter already at 335,000+ cycles, versus tens of thousands at the same checkpoint for earlier identities. Not root-caused here (flagged in FABRIC-3.md for later); worked around by splitting into 3 boots of 3 simultaneously-live identities each, every boot sharing the same seed so the master 27-slot shuffle is identical, filtered per boot by a new ACTIVE-LO/ACTIVE-HI range (EXEC-STD79-DOE's signature: seed lo hi -- ). run_id is always the slot's true position in the master shuffle, so trial order stays comparable across boots -- standard DoE blocking. riscv64: same 3-boot pattern, clean. Grand total: 81/81 trials correct, 0 mismatches, across all three architectures, all nine identities, all three replicates. Also flagged (FABRIC-3.md §XIV, not fixed): attaching several WIREBIND identities near-simultaneously (whether via rapid hotplug or all present from boot) causes the kernel to silently detect only some of them -- confirmed at the host/QMP level that every device was genuinely present. Worked around throughout this campaign by attaching one identity at a time with confirmed waits; real hardware hotplug could hit the same gap, so it's a genuine robustness concern, not just a test-harness inconvenience. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
70db955ac9
commit
e2abc56306
@@ -0,0 +1,25 @@
|
||||
# std79 DoE — 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial
|
||||
|
||||
The formal successor to `experiments/std79-exerciser/`'s ad hoc campaign (see FABRIC-3.md §XII),
|
||||
requested as a genuine randomized full-factorial design matching this project's own DoE
|
||||
methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle) rather than convenience
|
||||
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
|
||||
FABRIC-3.md §XV for the full design writeup, including a real aarch64 bug found and worked
|
||||
around along the way (not root-caused — flagged there for later).
|
||||
|
||||
`std79-doe.fth` is not a capsule loaded via `EXEC` — feed it as raw text to a running REPL
|
||||
(e.g. `socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth`), same as the original exerciser, then
|
||||
invoke `EXEC-STD79-DOE ( seed lo hi -- )` once loaded. `lo`/`hi` select which identity-index
|
||||
range (0-8) this boot's live VMs cover — pass `0 8` if all 9 identities are simultaneously
|
||||
attached in one boot (works on amd64); pass a narrower range (e.g. `0 2`, `3 5`, `6 8`) to run
|
||||
the DoE across multiple smaller boots when attaching all 9 at once isn't practical (see §XV —
|
||||
this is how aarch64 and riscv64 were actually run). Every boot must use the *same* seed so
|
||||
`INIT-MATRIX`/`SHUFFLE-MATRIX` reproduce the identical 27-slot master permutation; only the
|
||||
`lo`/`hi` filter differs, so `run_id` (always the slot's true position in the master shuffle,
|
||||
0-26) stays directly comparable across boots.
|
||||
|
||||
`results-20260911/` holds the raw captured serial output for the full campaign — `amd64-doe-
|
||||
raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
|
||||
`{aarch64,riscv64}-doe-batch{1,2,3}-raw.log` (27 trials each, split across 3 boots of 3
|
||||
identities per the workaround above). **Result: 81/81 trials correct, 0 mismatches** — every
|
||||
identity, every replicate, every architecture, byte-identical to a single established baseline.
|
||||
Reference in New Issue
Block a user