Files
LithosAnanake/experiments/std79-doe/README.md
T
Robert Allan JamesandClaude Sonnet 5 9eff122090
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run
Fix aarch64 Stadium/COOL O(ncells) scan; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVI)
Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s
eviction-fallback scan iterated the entire stadium_ncells array filtered by owner,
not the calling VM's own resident cells as its own doc comment claimed. Combined
with stadium_grant_quota() always splitting from Hera's shrinking free list and
stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule
executes, this compounded into a real O(n) blowup — catastrophic specifically on
aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap
kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886
cells, measured from boot logs).

Fixed by threading a real per-VM doubly-linked resident-cell list
(StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[])
so the fallback scan is bounded by that VM's own resident count, not the global
cell array size.

Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one
continuous boot per architecture, all 9 identities simultaneously live throughout
(the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials
correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw
logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now
completes in ~34s; full boot-to-DoE-complete in ~290s.

Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and
the project memory note) that described the symptom as a runaway "335,000+ cycles"
dispatch counter — those were cell array indices, not an event count.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-11 16:37:44 -04:00

36 lines
2.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# std79 DoE — 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial
The formal successor to `experiments/std79-exerciser/`'s ad hoc campaign (see FABRIC-3.md §XII),
requested as a genuine randomized full-factorial design matching this project's own DoE
methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle) rather than convenience
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
FABRIC-3.md §XV for the design writeup and §XVI for the real aarch64 Stadium/COOL scaling bug
that was found, root-caused, and fixed (`src/starkernel/vm/stadium.c`, commit pending) — the
3-boot batching workaround below is now historical only; a single boot with all 9 identities
simultaneously live works cleanly on all three architectures as of the fix.
`std79-doe.fth` is not a capsule loaded via `EXEC` — feed it as raw text to a running REPL
(e.g. `socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth`), same as the original exerciser, then
invoke `EXEC-STD79-DOE ( seed lo hi -- )` once loaded. `lo`/`hi` select which identity-index
range (0-8) this boot's live VMs cover — pass `0 8` for a single boot with all 9 identities
simultaneously attached (now confirmed working on all three architectures, see §XVI); a narrower
range (e.g. `0 2`, `3 5`, `6 8`) still works too, for running the DoE across multiple smaller
boots if ever needed for an unrelated reason. Every boot must use the *same* seed so
`INIT-MATRIX`/`SHUFFLE-MATRIX` reproduce the identical 27-slot master permutation; only the
`lo`/`hi` filter differs, so `run_id` (always the slot's true position in the master shuffle,
0-26) stays directly comparable across boots.
`results-20260911/` holds the raw captured serial output from the *original* campaign (pre-fix):
`amd64-doe-raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
`{aarch64,riscv64}-doe-batch{1,2,3}-raw.log` (27 trials each, split across 3 boots of 3
identities, the then-necessary workaround). **Result: 81/81 trials correct, 0 mismatches.**
`results-20260911-stadium-fix/` holds the *post-fix* rerun — one boot per architecture, all 9
identities simultaneously live in every boot, same seed throughout. **Result: 81/81 trials
correct, 0 mismatches, byte-identical output to the original campaign** — and the DOE-RUN
header sequence (run_id/id_idx/id_label/rep assignment) is md5-identical across all three raw
logs, confirming the master shuffle is genuinely architecture-independent. Total per-architecture
wall-clock (boot + all 9 identity attaches + full 27-trial run, one continuous QEMU session):
amd64 ~165s, aarch64 ~290s, riscv64 ~161s — aarch64's identity 04 attach alone, which stalled
90+ minutes before the fix, now completes in ~34s.