Fix aarch64 Stadium/COOL O(ncells) scan; rerun std79 DoE clean, 81/81 (FABRIC-3.md §XVI)
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run

Root cause of the 90+ minute aarch64 VM-birth stall found in §XV: stadium_admit()'s
eviction-fallback scan iterated the entire stadium_ncells array filtered by owner,
not the calling VM's own resident cells as its own doc comment claimed. Combined
with stadium_grant_quota() always splitting from Hera's shrinking free list and
stadium_word_dispatch() calling stadium_admit() per distinct word a VM's capsule
executes, this compounded into a real O(n) blowup — catastrophic specifically on
aarch64 because its -m 4096 (vs 1024 on amd64/riscv64) inflates the kmalloc heap
kmalloc_init() bisects down to, which inflates stadium_ncells 4x (335,544 vs 83,886
cells, measured from boot logs).

Fixed by threading a real per-VM doubly-linked resident-cell list
(StadiumVMQuota.resident_head + stadium_resident_next[]/stadium_resident_prev[])
so the fallback scan is bounded by that VM's own resident count, not the global
cell array size.

Verified with a full rerun of the 3x9x3 std79 DoE campaign from scratch: one
continuous boot per architecture, all 9 identities simultaneously live throughout
(the 3-boot aarch64/riscv64 batching workaround is no longer needed). 81/81 trials
correct, 0 mismatches, DOE-RUN header sequence md5-identical across all three raw
logs. Identity 04's attach on aarch64, which stalled 90+ minutes before, now
completes in ~34s; full boot-to-DoE-complete in ~290s.

Corrects an earlier misreading (carried into §XV, std79-doe.fth's comments, and
the project memory note) that described the symptom as a runaway "335,000+ cycles"
dispatch counter — those were cell array indices, not an event count.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
This commit is contained in:
Robert Allan James
2026-09-11 16:37:44 -04:00
co-authored by Claude Sonnet 5
parent e2abc56306
commit 9eff122090
23 changed files with 94109 additions and 63 deletions
+20 -10
View File
@@ -4,22 +4,32 @@ The formal successor to `experiments/std79-exerciser/`'s ad hoc campaign (see FA
requested as a genuine randomized full-factorial design matching this project's own DoE
methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle) rather than convenience
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
FABRIC-3.md §XV for the full design writeup, including a real aarch64 bug found and worked
around along the way (not root-caused — flagged there for later).
FABRIC-3.md §XV for the design writeup and §XVI for the real aarch64 Stadium/COOL scaling bug
that was found, root-caused, and fixed (`src/starkernel/vm/stadium.c`, commit pending) — the
3-boot batching workaround below is now historical only; a single boot with all 9 identities
simultaneously live works cleanly on all three architectures as of the fix.
`std79-doe.fth` is not a capsule loaded via `EXEC` — feed it as raw text to a running REPL
(e.g. `socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth`), same as the original exerciser, then
invoke `EXEC-STD79-DOE ( seed lo hi -- )` once loaded. `lo`/`hi` select which identity-index
range (0-8) this boot's live VMs cover — pass `0 8` if all 9 identities are simultaneously
attached in one boot (works on amd64); pass a narrower range (e.g. `0 2`, `3 5`, `6 8`) to run
the DoE across multiple smaller boots when attaching all 9 at once isn't practical (see §XV —
this is how aarch64 and riscv64 were actually run). Every boot must use the *same* seed so
range (0-8) this boot's live VMs cover — pass `0 8` for a single boot with all 9 identities
simultaneously attached (now confirmed working on all three architectures, see §XVI); a narrower
range (e.g. `0 2`, `3 5`, `6 8`) still works too, for running the DoE across multiple smaller
boots if ever needed for an unrelated reason. Every boot must use the *same* seed so
`INIT-MATRIX`/`SHUFFLE-MATRIX` reproduce the identical 27-slot master permutation; only the
`lo`/`hi` filter differs, so `run_id` (always the slot's true position in the master shuffle,
0-26) stays directly comparable across boots.
`results-20260911/` holds the raw captured serial output for the full campaign `amd64-doe-
raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
`results-20260911/` holds the raw captured serial output from the *original* campaign (pre-fix):
`amd64-doe-raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
`{aarch64,riscv64}-doe-batch{1,2,3}-raw.log` (27 trials each, split across 3 boots of 3
identities per the workaround above). **Result: 81/81 trials correct, 0 mismatches** — every
identity, every replicate, every architecture, byte-identical to a single established baseline.
identities, the then-necessary workaround). **Result: 81/81 trials correct, 0 mismatches.**
`results-20260911-stadium-fix/` holds the *post-fix* rerun — one boot per architecture, all 9
identities simultaneously live in every boot, same seed throughout. **Result: 81/81 trials
correct, 0 mismatches, byte-identical output to the original campaign** — and the DOE-RUN
header sequence (run_id/id_idx/id_label/rep assignment) is md5-identical across all three raw
logs, confirming the master shuffle is genuinely architecture-independent. Total per-architecture
wall-clock (boot + all 9 identity attaches + full 27-trial run, one continuous QEMU session):
amd64 ~165s, aarch64 ~290s, riscv64 ~161s — aarch64's identity 04 attach alone, which stalled
90+ minutes before the fix, now completes in ~34s.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+18 -15
View File
@@ -123,21 +123,24 @@ VARIABLE SW-I VARIABLE SW-J VARIABLE SW-VI VARIABLE SW-VJ
VARIABLE CURR-ID
VARIABLE CURR-REP
( Boot-batching support, added 2026-09-11 (FABRIC-3.md SXV): aarch64 hit a )
( severe, apparently superlinear per-tick slowdown once ~9-10 VMs stayed )
( simultaneously live at once (Stadium COOL dispatch cell counters running )
( into the hundreds of thousands with zero forward progress for 90+ )
( minutes) -- not yet root-caused, flagged as a real defect worth chasing )
( separately. Worked around by running each architecture as 3 boots of 3 )
( identities each (zuse+rajames+00, 01+02+03, 04+05+06), every boot sharing )
( the SAME seed so INIT-MATRIX/SHUFFLE-MATRIX produce the identical master )
( 27-slot permutation every time -- only ACTIVE-LO/ACTIVE-HI differ, so )
( each boot walks the FULL master order and simply skips any slot whose )
( identity isn't in its own live subset. The recorded run_id is always I )
( itself (the slot's true position in the master shuffle), never a )
( separately-incremented counter, so row order stays meaningful and )
( comparable across all 3 boots -- classic DoE "blocking": randomized )
( within, blocked across, by a practical constraint. )
( Boot-batching support, added 2026-09-11 (FABRIC-3.md SXV). Originally a )
( workaround for a real aarch64 Stadium/COOL scaling bug -- root-caused and )
( fixed 2026-09-11/12 (FABRIC-3.md SXVI, src/starkernel/vm/stadium.c): )
( stadium_grant_quota() halves the granting VM's own free list on every )
( birth with no floor, and stadium_admit()'s eviction fallback used to scan )
( the ENTIRE global cell array (stadium_ncells, in the hundreds of )
( thousands on aarch64's larger RAM-scaled Stadium) filtered by owner, )
( instead of walking just the calling VM's own resident cells as its own )
( doc comment already claimed -- fixed by threading a real per-VM resident )
( list. A single boot with all 9 identities simultaneously live now works )
( cleanly on all three architectures (see results-20260911-stadium-fix/), )
( so ACTIVE-LO/ACTIVE-HI is no longer load-bearing -- kept only as general )
( flexibility: every boot shares the SAME seed so INIT-MATRIX/SHUFFLE-MATRIX )
( produce the identical master 27-slot permutation regardless of how many )
( boots this runs across; only the ACTIVE-LO/ACTIVE-HI filter differs if )
( split. The recorded run_id is always I itself (the slot's true position )
( in the master shuffle), never a separately-incremented counter, so row )
( order stays meaningful and comparable across any number of boots. )
VARIABLE ACTIVE-LO
VARIABLE ACTIVE-HI
: ACTIVE? ( idx -- flag )