Files
LithosAnanake/experiments/std79-doe/README.md
T
Robert Allan JamesandClaude Sonnet 5 66ba21adb4
Build / build-amd64-iso (push) Canceled after 0s
Build / build-aarch64-iso (push) Canceled after 0s
Build / build-riscv64-img (push) Canceled after 0s
Log fleet_k_q48/fleet_conserved; K holds exactly, 775/775 ticks (FABRIC-3.md §XVIII)
doe_log.c's per-heartbeat-tick CSV gains two columns: fleet_k_q48
(vm_physics_fleet_heat_sum() over ALL live VMs -- the genuine fleet-wide
conservation invariant K, not reconstructable from the 3 named-Tripod-
member heat columns already logged, which omit every identity VM's own
heat) and fleet_conserved (vm_physics_conserved() as 0/1). Requested
explicitly after the first heartbeat-telemetry analysis pass
(analysis-20260912/) omitted K entirely.

Kernel rebuilt on all three architectures, full 3x9x3 campaign rerun
(results-20260912-with-k/). K = 1.0000000000 (Q48.16 raw 65536) on every
one of 775 heartbeat-tick observations, sd(K) = 0, 100% fleet_conserved,
across amd64/aarch64/riscv64, nine identities, three replicates -- zero
deviation. Also a free regression check on both recent Stadium fixes
(§XVI/§XVII): neither disturbed the reservoir-transfer accounting K
depends on.

Found and fixed a tooling wrinkle along the way: fleet_conserved, being
the CSV row's very last field with nothing after it to bound a regex
match, can have a resumed trial digit merge into it with zero separator
on the wire -- combine.py now derives it from fleet_k_q48 directly (same
epsilon vm_physics_conserved() uses) instead of trusting the raw field.
fleet_k_q48 itself is unaffected either way.

Full analysis, discussion, and light/dark SVG->PDF figures written up as
a proper LaTeX report (report-20260912/report/std79_doe_report.pdf),
following experiments/bare_metal/analysis/report/bare_metal_doe_report.tex's
established style -- supersedes analysis-20260912/'s markdown-only first
pass as the primary deliverable for this dataset (kept, not discarded).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 13:25:09 -04:00

98 lines
7.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# std79 DoE — 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial
The formal successor to `experiments/std79-exerciser/`'s ad hoc campaign (see FABRIC-3.md §XII),
requested as a genuine randomized full-factorial design matching this project's own DoE
methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle) rather than convenience
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
FABRIC-3.md §XV for the design writeup, §XVI for the real aarch64 Stadium/COOL scaling bug that
was found, root-caused, and fixed (`src/starkernel/vm/stadium.c`), and §XVII for a second,
related Stadium bug (`stadium_grant_quota()`'s donor-floor) found while writing up §XVI and
fixed in a follow-up pass — the 3-boot batching workaround below is now historical only; a
single boot with all 9 identities simultaneously live works cleanly on all three architectures
as of both fixes.
`std79-doe.fth` is not a capsule loaded via `EXEC` — feed it as raw text to a running REPL
(e.g. `socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth`), same as the original exerciser, then
invoke `EXEC-STD79-DOE ( seed lo hi -- )` once loaded. `lo`/`hi` select which identity-index
range (0-8) this boot's live VMs cover — pass `0 8` for a single boot with all 9 identities
simultaneously attached (now confirmed working on all three architectures, see §XVI); a narrower
range (e.g. `0 2`, `3 5`, `6 8`) still works too, for running the DoE across multiple smaller
boots if ever needed for an unrelated reason. Every boot must use the *same* seed so
`INIT-MATRIX`/`SHUFFLE-MATRIX` reproduce the identical 27-slot master permutation; only the
`lo`/`hi` filter differs, so `run_id` (always the slot's true position in the master shuffle,
0-26) stays directly comparable across boots.
`results-20260911/` holds the raw captured serial output from the *original* campaign (pre-fix):
`amd64-doe-raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
`{aarch64,riscv64}-doe-batch{1,2,3}-raw.log` (27 trials each, split across 3 boots of 3
identities, the then-necessary workaround). **Result: 81/81 trials correct, 0 mismatches.**
`results-20260911-stadium-fix/` holds the rerun after §XVI's O(ncells)-scan fix — one boot per
architecture, all 9 identities simultaneously live in every boot, same seed throughout.
**Result: 81/81 trials correct, 0 mismatches, byte-identical output to the original campaign**
and the DOE-RUN header sequence (run_id/id_idx/id_label/rep assignment) is md5-identical across
all three raw logs, confirming the master shuffle is genuinely architecture-independent. Total
per-architecture wall-clock (boot + all 9 identity attaches + full 27-trial run, one continuous
QEMU session): amd64 ~165s, aarch64 ~290s, riscv64 ~161s — aarch64's identity 04 attach alone,
which stalled 90+ minutes before the fix, now completes in ~34s.
`results-20260911-donor-floor-fix/` holds a further rerun after §XVII's donor-floor fix (same
discipline: any Stadium defect repair reruns the whole DoE from the top). **Result: 81/81
trials correct, 0 mismatches**, DOE-RUN header sequence md5-identical to every prior run. Total
wall-clock: amd64 ~161s, aarch64 ~280s (confirms no regression from §XVI's fix), riscv64 ~162s.
`results-20260912-reconfirm/` holds a plain reconfirmation rerun the following day, no code
changes since §XVII's fix (`e51a8d2`) — same seed, same single-boot-per-architecture, all 9
identities simultaneously live. **Result: 81/81 trials correct, 0 mismatches**, shuffle sequence
still md5-identical to every prior run. Total wall-clock: amd64 ~151s, aarch64 ~276s, riscv64
~165s.
`results-20260912-with-heartbeat-csv/``EXEC-STD79-DOE` now brackets its trial loop with
`HB-ON`/`HB-OFF` (`doe_log.c`'s per-heartbeat-tick physics/timing CSV logger), so
`scripts/extract_doe.sh`'s automatically-generated CSV finally carries real data for this
campaign — every prior run's extracted CSV was silently empty, since `g_doe_log_enabled`
defaults off and nothing had ever turned it on. Each directory holds both the raw log
(`*-doe-raw.log`) and its paired heartbeat CSV (`*-heartbeat.csv`, ~255-257 rows/architecture,
18 columns: tick_number, elapsed_ns, hot_word_count, avg_word_heat_q48, window_width,
apic_ticks, hera/hermes/artemis heat_q48, etc. — see `doe_log.c`'s own header comment for the
full column list). **Result: 27/27 trials correct on every architecture** (verified via the
campaign's most distinctive result markers — the two 18-19 digit `M*`/`M/MOD` values and the
`2147483648 2/` result all show exactly 27 occurrences, zero faults) — **not** verified via
exact substring/byte reconstruction, because turning HB-ON on for the whole run exposed a real,
if purely cosmetic, effect: the heartbeat tick's async CSV printer and the trial loop's own
console output share the same serial line with no locking between them, so CSV rows get spliced
mid-token into the trial output on the wire (confirmed live: a "missing" `DOE-RUN,0` line turned
out to be `DOE-RUN,` and `0 ,6 ,04,2` on two separate physical log lines with a full CSV row
printed in between). The underlying FORTH execution itself is unaffected — values are correct,
nothing is corrupted in memory — only the printed character stream interleaves, so downstream
tooling that wants an exact reconstruction of the trial-output stream from these three raw logs
needs to account for that (strip `[HADES][DOE ] ...` fragments and rejoin) rather than assume
one physical line is one logical print.
`analysis-20260912/` builds exactly that reconstruction and turns it into a 3×9 factorial
analysis: `correlate_doe.py` labels every heartbeat-tick row with the trial that was active when
it printed, `combine.py` merges all three architectures into `combined.csv` (768 rows, Q48.16
fields decoded to floats), and `analysis.R` computes per-cell (architecture × identity) means/SD
and a two-way ANOVA for each of 12 telemetry metrics, rendering one boxplot SVG per metric.
**Superseded by `report-20260912/` below as the primary deliverable** — this first pass never
captured K (the fleet conservation invariant) and reported results as plain markdown rather than
this project's established LaTeX report style; kept as-is (never discard data), still useful for
the 12 secondary metrics' own writeup.
`results-20260912-with-k/``doe_log.c`'s per-tick CSV gained two more columns,
`fleet_k_q48`/`fleet_conserved` (`vm_physics_fleet_heat_sum()` over ALL live VMs — the genuine
fleet-wide conservation invariant, not reconstructable from the 3 named-Tripod-member heat
columns already present), requiring a kernel rebuild + fresh campaign rerun on all three
architectures. Same raw-log/heartbeat-CSV pairing as `results-20260912-with-heartbeat-csv/`.
`report-20260912/`**the primary deliverable for the heartbeat/K dataset.** Same
correlate/combine/R pipeline as `analysis-20260912/`, extended for K, feeding a hand-authored
LaTeX report (`report-20260912/report/std79_doe_report.pdf`) built in the same style as
`experiments/bare_metal/analysis/report/bare_metal_doe_report.tex` (abstract-first with headline
numbers, TOC, light/dark figure pairs, ANOVA tables, discussion, limitations). **Headline
result: K = 1.0000000000 (Q48.16 raw 65536) on every one of 775 heartbeat-tick observations,
sd(K) = 0, 100% conserved, across all three architectures, nine identities, and three
replicates — zero deviation.** See `report-20260912/README.md` for the full pipeline and the
report itself for the complete analysis and discussion, including why this also serves as a
free regression check on the two Stadium fixes (§XVIXVII) that landed the same week.