Files
LithosAnanake/experiments/std79-doe/README.md
T
Robert Allan JamesandClaude Sonnet 5 16cc74243c
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run
std79 DoE campaign rerun post-Stage-4: 81/81, §XIV caught live by rerun's own harness bug (FABRIC-3.md §XXXI)
Reran the established 3x9x3 randomized full-factorial campaign
(std79-doe.fth) on all 3 architectures per standing project discipline
(any Stadium-adjacent change reruns the whole DoE from the top).

First amd64 attempt exposed a real test-harness bug that re-triggered
the already-known §XIV concurrent-attach gap: the sequential-attach wait
loop checked for any recent WIREBIND-attached line instead of the
specific identity requested, firing the next device_add before the
kernel finished the current one. Only 2 of 8 identities attached; the
campaign itself completed cleanly with graceful "VM-EXEC: VM not found"
refusals rather than corrupting anything. Fixed the wait loop, discarded
the invalid run's campaign result (its boot log kept for the record),
reran clean.

Corrected reruns: all 8 identities individually confirmed on all 3
architectures, zero VM-not-found errors, zero faults, 81/81 trials
correct against established baseline values. Raw logs archived at
experiments/std79-doe/results-20260915-stage4/.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BWpNjdwPtFLuVLaAq44L9K
2026-09-15 04:56:31 -04:00

116 lines
9.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# std79 DoE — 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial
The formal successor to `experiments/std79-exerciser/`'s ad hoc campaign (see FABRIC-3.md §XII),
requested as a genuine randomized full-factorial design matching this project's own DoE
methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle) rather than convenience
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
FABRIC-3.md §XV for the design writeup, §XVI for the real aarch64 Stadium/COOL scaling bug that
was found, root-caused, and fixed (`src/starkernel/vm/stadium.c`), and §XVII for a second,
related Stadium bug (`stadium_grant_quota()`'s donor-floor) found while writing up §XVI and
fixed in a follow-up pass — the 3-boot batching workaround below is now historical only; a
single boot with all 9 identities simultaneously live works cleanly on all three architectures
as of both fixes.
`std79-doe.fth` is not a capsule loaded via `EXEC` — feed it as raw text to a running REPL
(e.g. `socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth`), same as the original exerciser, then
invoke `EXEC-STD79-DOE ( seed lo hi -- )` once loaded. `lo`/`hi` select which identity-index
range (0-8) this boot's live VMs cover — pass `0 8` for a single boot with all 9 identities
simultaneously attached (now confirmed working on all three architectures, see §XVI); a narrower
range (e.g. `0 2`, `3 5`, `6 8`) still works too, for running the DoE across multiple smaller
boots if ever needed for an unrelated reason. Every boot must use the *same* seed so
`INIT-MATRIX`/`SHUFFLE-MATRIX` reproduce the identical 27-slot master permutation; only the
`lo`/`hi` filter differs, so `run_id` (always the slot's true position in the master shuffle,
0-26) stays directly comparable across boots.
`results-20260911/` holds the raw captured serial output from the *original* campaign (pre-fix):
`amd64-doe-raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
`{aarch64,riscv64}-doe-batch{1,2,3}-raw.log` (27 trials each, split across 3 boots of 3
identities, the then-necessary workaround). **Result: 81/81 trials correct, 0 mismatches.**
`results-20260911-stadium-fix/` holds the rerun after §XVI's O(ncells)-scan fix — one boot per
architecture, all 9 identities simultaneously live in every boot, same seed throughout.
**Result: 81/81 trials correct, 0 mismatches, byte-identical output to the original campaign**
and the DOE-RUN header sequence (run_id/id_idx/id_label/rep assignment) is md5-identical across
all three raw logs, confirming the master shuffle is genuinely architecture-independent. Total
per-architecture wall-clock (boot + all 9 identity attaches + full 27-trial run, one continuous
QEMU session): amd64 ~165s, aarch64 ~290s, riscv64 ~161s — aarch64's identity 04 attach alone,
which stalled 90+ minutes before the fix, now completes in ~34s.
`results-20260911-donor-floor-fix/` holds a further rerun after §XVII's donor-floor fix (same
discipline: any Stadium defect repair reruns the whole DoE from the top). **Result: 81/81
trials correct, 0 mismatches**, DOE-RUN header sequence md5-identical to every prior run. Total
wall-clock: amd64 ~161s, aarch64 ~280s (confirms no regression from §XVI's fix), riscv64 ~162s.
`results-20260912-reconfirm/` holds a plain reconfirmation rerun the following day, no code
changes since §XVII's fix (`e51a8d2`) — same seed, same single-boot-per-architecture, all 9
identities simultaneously live. **Result: 81/81 trials correct, 0 mismatches**, shuffle sequence
still md5-identical to every prior run. Total wall-clock: amd64 ~151s, aarch64 ~276s, riscv64
~165s.
`results-20260912-with-heartbeat-csv/``EXEC-STD79-DOE` now brackets its trial loop with
`HB-ON`/`HB-OFF` (`doe_log.c`'s per-heartbeat-tick physics/timing CSV logger), so
`scripts/extract_doe.sh`'s automatically-generated CSV finally carries real data for this
campaign — every prior run's extracted CSV was silently empty, since `g_doe_log_enabled`
defaults off and nothing had ever turned it on. Each directory holds both the raw log
(`*-doe-raw.log`) and its paired heartbeat CSV (`*-heartbeat.csv`, ~255-257 rows/architecture,
18 columns: tick_number, elapsed_ns, hot_word_count, avg_word_heat_q48, window_width,
apic_ticks, hera/hermes/artemis heat_q48, etc. — see `doe_log.c`'s own header comment for the
full column list). **Result: 27/27 trials correct on every architecture** (verified via the
campaign's most distinctive result markers — the two 18-19 digit `M*`/`M/MOD` values and the
`2147483648 2/` result all show exactly 27 occurrences, zero faults) — **not** verified via
exact substring/byte reconstruction, because turning HB-ON on for the whole run exposed a real,
if purely cosmetic, effect: the heartbeat tick's async CSV printer and the trial loop's own
console output share the same serial line with no locking between them, so CSV rows get spliced
mid-token into the trial output on the wire (confirmed live: a "missing" `DOE-RUN,0` line turned
out to be `DOE-RUN,` and `0 ,6 ,04,2` on two separate physical log lines with a full CSV row
printed in between). The underlying FORTH execution itself is unaffected — values are correct,
nothing is corrupted in memory — only the printed character stream interleaves, so downstream
tooling that wants an exact reconstruction of the trial-output stream from these three raw logs
needs to account for that (strip `[HADES][DOE ] ...` fragments and rejoin) rather than assume
one physical line is one logical print.
`analysis-20260912/` builds exactly that reconstruction and turns it into a 3×9 factorial
analysis: `correlate_doe.py` labels every heartbeat-tick row with the trial that was active when
it printed, `combine.py` merges all three architectures into `combined.csv` (768 rows, Q48.16
fields decoded to floats), and `analysis.R` computes per-cell (architecture × identity) means/SD
and a two-way ANOVA for each of 12 telemetry metrics, rendering one boxplot SVG per metric.
**Superseded by `report-20260912/` below as the primary deliverable** — this first pass never
captured K (the fleet conservation invariant) and reported results as plain markdown rather than
this project's established LaTeX report style; kept as-is (never discard data), still useful for
the 12 secondary metrics' own writeup.
`results-20260912-with-k/``doe_log.c`'s per-tick CSV gained two more columns,
`fleet_k_q48`/`fleet_conserved` (`vm_physics_fleet_heat_sum()` over ALL live VMs — the genuine
fleet-wide conservation invariant, not reconstructable from the 3 named-Tripod-member heat
columns already present), requiring a kernel rebuild + fresh campaign rerun on all three
architectures. Same raw-log/heartbeat-CSV pairing as `results-20260912-with-heartbeat-csv/`.
`report-20260912/`**the primary deliverable for the heartbeat/K dataset.** Same
correlate/combine/R pipeline as `analysis-20260912/`, extended for K, feeding a hand-authored
LaTeX report (`report-20260912/report/std79_doe_report.pdf`) built in the same style as
`experiments/bare_metal/analysis/report/bare_metal_doe_report.tex` (abstract-first with headline
numbers, TOC, light/dark figure pairs, ANOVA tables, discussion, limitations). **Headline
result: K = 1.0000000000 (Q48.16 raw 65536) on every one of 775 heartbeat-tick observations,
sd(K) = 0, 100% conserved, across all three architectures, nine identities, and three
replicates — zero deviation.** See `report-20260912/README.md` for the full pipeline and the
report itself for the complete analysis and discussion, including why this also serves as a
free regression check on the two Stadium fixes (§XVIXVII) that landed the same week.
`results-20260915-stage4/` — rerun after FABRIC-3.md §XXVIII Stage 4 (WIREBIND VMs as Stage 3
switch-signal participants + mark-and-defer tombstone reap), same discipline as every prior
Stadium-adjacent change: any relevant mechanism change reruns the whole DoE from the top, one
boot per architecture, all 9 identities simultaneously live, same seed (12345). **Result: 81/81
trials correct, 0 mismatches, 0 `VM-EXEC: VM not found` errors, zero `VM fault` halts, on all
three architectures.** Note for anyone rerunning this: the sequential-attach wait loop must
confirm each identity's own specific `WIREBIND: <name> attached and ready` line before
proceeding to the next — a loop that just checks for *any* recent WIREBIND-attached line will
spuriously "succeed" immediately and fire the next `device_add` before the kernel has actually
finished the current one, re-triggering §XIV's own known concurrent-attach detection gap (caught
live during this rerun: a first attempt with a buggy wait loop attached only 2 of 8 identities,
though the campaign itself still completed cleanly with graceful `VM-EXEC: VM not found`
refusals for the missing ones rather than corrupting anything — the harness bug, not a kernel
defect). This campaign still drives every identity via direct `VM-EXEC` dispatch, same as every
prior run — it does not itself exercise Stage 4's new switch-signal participation, which was
verified separately (see FABRIC-3.md §XXX for the live multiuser-attach + preemption-fleet
verification that does).