Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s eviction-fallback scan iterated the entire stadium_ncells array filtered by owner, not the calling VM's own resident cells as its own doc comment claimed. Combined with stadium_grant_quota() always splitting from Hera's shrinking free list and stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule executes, this compounded into a real O(n) blowup — catastrophic specifically on aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886 cells, measured from boot logs). Fixed by threading a real per-VM doubly-linked resident-cell list (StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[]) so the fallback scan is bounded by that VM's own resident count, not the global cell array size. Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one continuous boot per architecture, all 9 identities simultaneously live throughout (the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now completes in ~34s; full boot-to-DoE-complete in ~290s. Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and the project memory note) that described the symptom as a runaway "335,000+ cycles" dispatch counter — those were cell array indices, not an event count. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
36 lines
2.7 KiB
Markdown
36 lines
2.7 KiB
Markdown
# std79 DoE — 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial
|
||
|
||
The formal successor to `experiments/std79-exerciser/`'s ad hoc campaign (see FABRIC-3.md §XII),
|
||
requested as a genuine randomized full-factorial design matching this project's own DoE
|
||
methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle) rather than convenience
|
||
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
|
||
FABRIC-3.md §XV for the design writeup and §XVI for the real aarch64 Stadium/COOL scaling bug
|
||
that was found, root-caused, and fixed (`src/starkernel/vm/stadium.c`, commit pending) — the
|
||
3-boot batching workaround below is now historical only; a single boot with all 9 identities
|
||
simultaneously live works cleanly on all three architectures as of the fix.
|
||
|
||
`std79-doe.fth` is not a capsule loaded via `EXEC` — feed it as raw text to a running REPL
|
||
(e.g. `socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth`), same as the original exerciser, then
|
||
invoke `EXEC-STD79-DOE ( seed lo hi -- )` once loaded. `lo`/`hi` select which identity-index
|
||
range (0-8) this boot's live VMs cover — pass `0 8` for a single boot with all 9 identities
|
||
simultaneously attached (now confirmed working on all three architectures, see §XVI); a narrower
|
||
range (e.g. `0 2`, `3 5`, `6 8`) still works too, for running the DoE across multiple smaller
|
||
boots if ever needed for an unrelated reason. Every boot must use the *same* seed so
|
||
`INIT-MATRIX`/`SHUFFLE-MATRIX` reproduce the identical 27-slot master permutation; only the
|
||
`lo`/`hi` filter differs, so `run_id` (always the slot's true position in the master shuffle,
|
||
0-26) stays directly comparable across boots.
|
||
|
||
`results-20260911/` holds the raw captured serial output from the *original* campaign (pre-fix):
|
||
`amd64-doe-raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
|
||
`{aarch64,riscv64}-doe-batch{1,2,3}-raw.log` (27 trials each, split across 3 boots of 3
|
||
identities, the then-necessary workaround). **Result: 81/81 trials correct, 0 mismatches.**
|
||
|
||
`results-20260911-stadium-fix/` holds the *post-fix* rerun — one boot per architecture, all 9
|
||
identities simultaneously live in every boot, same seed throughout. **Result: 81/81 trials
|
||
correct, 0 mismatches, byte-identical output to the original campaign** — and the DOE-RUN
|
||
header sequence (run_id/id_idx/id_label/rep assignment) is md5-identical across all three raw
|
||
logs, confirming the master shuffle is genuinely architecture-independent. Total per-architecture
|
||
wall-clock (boot + all 9 identity attaches + full 27-trial run, one continuous QEMU session):
|
||
amd64 ~165s, aarch64 ~290s, riscv64 ~161s — aarch64's identity 04 attach alone, which stalled
|
||
90+ minutes before the fix, now completes in ~34s.
|