Two real bugs found and fixed while attempting to raise N-REPS from 3 (both harmless only by coincidence at N-REPS=3, since 3 happened to equal the hardcoded/literal values): - N-RUNS was `27 CONSTANT`, not derived -- now `N-ID N-REPS * CONSTANT N-RUNS`. - START-REP's run_id decode used a literal `3 *` where it meant `N-REPS *` (confirmed against EXEC-STD79-DOE, the serialized baseline, which correctly uses N-REPS for the same decode). Re-verified at N-REPS=3 (amd64): byte-identical to the already- verified baseline -- 193s, 0 faults, 99/99 tokens every identity, K conserved on all 656 rows. Raising N-REPS to 6 was then tested and found to cause real, silent data loss at full 8-identity scale (rajames 0/198, 00 66/198, 01/02 99/198, 03-06 fully complete) -- same failure class as §XXI defect 2, just past the budget again since doubling N-REPS roughly doubles total MSG-SEND volume (24->48). Measured the reservoir's replenishment directly rather than assume from source: 40 sends drained it from 21735 to 1; 300s of pure idle time (zero further sends) brought it back to 2049 -- real, but only ~6.8 units/s, meaning a full refill would take on the order of 53 minutes against a campaign's few-hundred-second runtime. Raising N-REPS further needs a cheaper per-send cost or an explicit top-up, not reliance on ambient decay. Reverted N-REPS to 3 (fixes kept, they're correctness fixes independent of the value) rather than commit a silently-lossy result. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
std79 DoE — 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial
The formal successor to experiments/std79-exerciser/'s ad hoc campaign (see FABRIC-3.md §XII),
requested as a genuine randomized full-factorial design matching this project's own DoE
methodology (capsules/doe.4th's Fisher-Yates run-matrix shuffle) rather than convenience
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
FABRIC-3.md §XV for the design writeup, §XVI for the real aarch64 Stadium/COOL scaling bug that
was found, root-caused, and fixed (src/starkernel/vm/stadium.c), and §XVII for a second,
related Stadium bug (stadium_grant_quota()'s donor-floor) found while writing up §XVI and
fixed in a follow-up pass — the 3-boot batching workaround below is now historical only; a
single boot with all 9 identities simultaneously live works cleanly on all three architectures
as of both fixes.
std79-doe.fth is not a capsule loaded via EXEC — feed it as raw text to a running REPL
(e.g. socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth), same as the original exerciser, then
invoke EXEC-STD79-DOE ( seed lo hi -- ) once loaded. lo/hi select which identity-index
range (0-8) this boot's live VMs cover — pass 0 8 for a single boot with all 9 identities
simultaneously attached (now confirmed working on all three architectures, see §XVI); a narrower
range (e.g. 0 2, 3 5, 6 8) still works too, for running the DoE across multiple smaller
boots if ever needed for an unrelated reason. Every boot must use the same seed so
INIT-MATRIX/SHUFFLE-MATRIX reproduce the identical 27-slot master permutation; only the
lo/hi filter differs, so run_id (always the slot's true position in the master shuffle,
0-26) stays directly comparable across boots.
results-20260911/ holds the raw captured serial output from the original campaign (pre-fix):
amd64-doe-raw.log (all 27 trials, one boot, all 9 identities simultaneously live) and
{aarch64,riscv64}-doe-batch{1,2,3}-raw.log (27 trials each, split across 3 boots of 3
identities, the then-necessary workaround). Result: 81/81 trials correct, 0 mismatches.
results-20260911-stadium-fix/ holds the rerun after §XVI's O(ncells)-scan fix — one boot per
architecture, all 9 identities simultaneously live in every boot, same seed throughout.
Result: 81/81 trials correct, 0 mismatches, byte-identical output to the original campaign —
and the DOE-RUN header sequence (run_id/id_idx/id_label/rep assignment) is md5-identical across
all three raw logs, confirming the master shuffle is genuinely architecture-independent. Total
per-architecture wall-clock (boot + all 9 identity attaches + full 27-trial run, one continuous
QEMU session): amd64 ~165s, aarch64 ~290s, riscv64 ~161s — aarch64's identity 04 attach alone,
which stalled 90+ minutes before the fix, now completes in ~34s.
results-20260911-donor-floor-fix/ holds a further rerun after §XVII's donor-floor fix (same
discipline: any Stadium defect repair reruns the whole DoE from the top). Result: 81/81
trials correct, 0 mismatches, DOE-RUN header sequence md5-identical to every prior run. Total
wall-clock: amd64 ~161s, aarch64 ~280s (confirms no regression from §XVI's fix), riscv64 ~162s.
results-20260912-reconfirm/ holds a plain reconfirmation rerun the following day, no code
changes since §XVII's fix (e51a8d2) — same seed, same single-boot-per-architecture, all 9
identities simultaneously live. Result: 81/81 trials correct, 0 mismatches, shuffle sequence
still md5-identical to every prior run. Total wall-clock: amd64 ~151s, aarch64 ~276s, riscv64
~165s.
results-20260912-with-heartbeat-csv/ — EXEC-STD79-DOE now brackets its trial loop with
HB-ON/HB-OFF (doe_log.c's per-heartbeat-tick physics/timing CSV logger), so
scripts/extract_doe.sh's automatically-generated CSV finally carries real data for this
campaign — every prior run's extracted CSV was silently empty, since g_doe_log_enabled
defaults off and nothing had ever turned it on. Each directory holds both the raw log
(*-doe-raw.log) and its paired heartbeat CSV (*-heartbeat.csv, ~255-257 rows/architecture,
18 columns: tick_number, elapsed_ns, hot_word_count, avg_word_heat_q48, window_width,
apic_ticks, hera/hermes/artemis heat_q48, etc. — see doe_log.c's own header comment for the
full column list). Result: 27/27 trials correct on every architecture (verified via the
campaign's most distinctive result markers — the two 18-19 digit M*/M/MOD values and the
2147483648 2/ result all show exactly 27 occurrences, zero faults) — not verified via
exact substring/byte reconstruction, because turning HB-ON on for the whole run exposed a real,
if purely cosmetic, effect: the heartbeat tick's async CSV printer and the trial loop's own
console output share the same serial line with no locking between them, so CSV rows get spliced
mid-token into the trial output on the wire (confirmed live: a "missing" DOE-RUN,0 line turned
out to be DOE-RUN, and 0 ,6 ,04,2 on two separate physical log lines with a full CSV row
printed in between). The underlying FORTH execution itself is unaffected — values are correct,
nothing is corrupted in memory — only the printed character stream interleaves, so downstream
tooling that wants an exact reconstruction of the trial-output stream from these three raw logs
needs to account for that (strip [HADES][DOE ] ... fragments and rejoin) rather than assume
one physical line is one logical print.
analysis-20260912/ builds exactly that reconstruction and turns it into a 3×9 factorial
analysis: correlate_doe.py labels every heartbeat-tick row with the trial that was active when
it printed, combine.py merges all three architectures into combined.csv (768 rows, Q48.16
fields decoded to floats), and analysis.R computes per-cell (architecture × identity) means/SD
and a two-way ANOVA for each of 12 telemetry metrics, rendering one boxplot SVG per metric.
Superseded by report-20260912/ below as the primary deliverable — this first pass never
captured K (the fleet conservation invariant) and reported results as plain markdown rather than
this project's established LaTeX report style; kept as-is (never discard data), still useful for
the 12 secondary metrics' own writeup.
results-20260912-with-k/ — doe_log.c's per-tick CSV gained two more columns,
fleet_k_q48/fleet_conserved (vm_physics_fleet_heat_sum() over ALL live VMs — the genuine
fleet-wide conservation invariant, not reconstructable from the 3 named-Tripod-member heat
columns already present), requiring a kernel rebuild + fresh campaign rerun on all three
architectures. Same raw-log/heartbeat-CSV pairing as results-20260912-with-heartbeat-csv/.
report-20260912/ — the primary deliverable for the heartbeat/K dataset. Same
correlate/combine/R pipeline as analysis-20260912/, extended for K, feeding a hand-authored
LaTeX report (report-20260912/report/std79_doe_report.pdf) built in the same style as
experiments/bare_metal/analysis/report/bare_metal_doe_report.tex (abstract-first with headline
numbers, TOC, light/dark figure pairs, ANOVA tables, discussion, limitations). Headline
result: K = 1.0000000000 (Q48.16 raw 65536) on every one of 775 heartbeat-tick observations,
sd(K) = 0, 100% conserved, across all three architectures, nine identities, and three
replicates — zero deviation. See report-20260912/README.md for the full pipeline and the
report itself for the complete analysis and discussion, including why this also serves as a
free regression check on the two Stadium fixes (§XVI–XVII) that landed the same week.