Files
LithosAnanake/experiments/std79-doe/README.md
T
Robert Allan JamesandClaude Sonnet 5 934be5a257
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run
Add 3x9 factorial analysis of std79-doe heartbeat/physics telemetry
Correlates every HB-ON/HB-OFF heartbeat-tick CSV row (results-20260912-
with-heartbeat-csv/*-doe-raw.log) back to which trial (run_id, id_idx,
id_label, rep) was active when it printed, despite the async tick
printer splicing rows mid-token -- including mid a DOE-RUN marker itself
-- into the trial loop's own console output on the shared serial line.

Pipeline (analysis-20260912/, see its own README.md):
- correlate_doe.py: two-pass reconstruction per architecture (remove
  atomic CSV-row spans to rebuild the clean trial-output stream, map
  each removed row's offset back to the nearest preceding run_id
  marker); identity/rep looked up from a known-clean prior run's
  run_id mapping rather than re-parsed, since one aarch64 marker
  (trial 12, rajames rep 1) lost its id_idx digit to a zero-separator
  collision with an adjacent CSV field and is unrecoverable from that
  log alone -- its rows fold into trial 11 instead, documented as a
  known limitation.
- combine.py: merges all three architectures into combined.csv (768
  rows), decoding Q48.16 fields to floats and jitter_bits' IEEE754 bit
  pattern to real jitter_ns.
- analysis.R: per-cell (architecture x identity) means/SD and two-way
  ANOVA for each of 12 telemetry metrics, one boxplot SVG per metric,
  written up as ANALYSIS.md.

Key findings: identity significantly affects word-heat/window-sizing
metrics (expected -- different identities execute different word
sets), architecture significantly affects timing metrics (APIC
ticks/tick, timing variance, fleet heat -- expected, different QEMU
targets), zero architecture x identity interaction on any metric.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
2026-09-12 12:46:34 -04:00

83 lines
6.4 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# std79 DoE — 3(architecture) × 9(identity) × 3(replicate) randomized full-factorial
The formal successor to `experiments/std79-exerciser/`'s ad hoc campaign (see FABRIC-3.md §XII),
requested as a genuine randomized full-factorial design matching this project's own DoE
methodology (`capsules/doe.4th`'s Fisher-Yates run-matrix shuffle) rather than convenience
batching — and written entirely in FORTH, not host-orchestrated shell scripting. See
FABRIC-3.md §XV for the design writeup, §XVI for the real aarch64 Stadium/COOL scaling bug that
was found, root-caused, and fixed (`src/starkernel/vm/stadium.c`), and §XVII for a second,
related Stadium bug (`stadium_grant_quota()`'s donor-floor) found while writing up §XVI and
fixed in a follow-up pass — the 3-boot batching workaround below is now historical only; a
single boot with all 9 identities simultaneously live works cleanly on all three architectures
as of both fixes.
`std79-doe.fth` is not a capsule loaded via `EXEC` — feed it as raw text to a running REPL
(e.g. `socat - UNIX-CONNECT:<serial_sock> < std79-doe.fth`), same as the original exerciser, then
invoke `EXEC-STD79-DOE ( seed lo hi -- )` once loaded. `lo`/`hi` select which identity-index
range (0-8) this boot's live VMs cover — pass `0 8` for a single boot with all 9 identities
simultaneously attached (now confirmed working on all three architectures, see §XVI); a narrower
range (e.g. `0 2`, `3 5`, `6 8`) still works too, for running the DoE across multiple smaller
boots if ever needed for an unrelated reason. Every boot must use the *same* seed so
`INIT-MATRIX`/`SHUFFLE-MATRIX` reproduce the identical 27-slot master permutation; only the
`lo`/`hi` filter differs, so `run_id` (always the slot's true position in the master shuffle,
0-26) stays directly comparable across boots.
`results-20260911/` holds the raw captured serial output from the *original* campaign (pre-fix):
`amd64-doe-raw.log` (all 27 trials, one boot, all 9 identities simultaneously live) and
`{aarch64,riscv64}-doe-batch{1,2,3}-raw.log` (27 trials each, split across 3 boots of 3
identities, the then-necessary workaround). **Result: 81/81 trials correct, 0 mismatches.**
`results-20260911-stadium-fix/` holds the rerun after §XVI's O(ncells)-scan fix — one boot per
architecture, all 9 identities simultaneously live in every boot, same seed throughout.
**Result: 81/81 trials correct, 0 mismatches, byte-identical output to the original campaign**
and the DOE-RUN header sequence (run_id/id_idx/id_label/rep assignment) is md5-identical across
all three raw logs, confirming the master shuffle is genuinely architecture-independent. Total
per-architecture wall-clock (boot + all 9 identity attaches + full 27-trial run, one continuous
QEMU session): amd64 ~165s, aarch64 ~290s, riscv64 ~161s — aarch64's identity 04 attach alone,
which stalled 90+ minutes before the fix, now completes in ~34s.
`results-20260911-donor-floor-fix/` holds a further rerun after §XVII's donor-floor fix (same
discipline: any Stadium defect repair reruns the whole DoE from the top). **Result: 81/81
trials correct, 0 mismatches**, DOE-RUN header sequence md5-identical to every prior run. Total
wall-clock: amd64 ~161s, aarch64 ~280s (confirms no regression from §XVI's fix), riscv64 ~162s.
`results-20260912-reconfirm/` holds a plain reconfirmation rerun the following day, no code
changes since §XVII's fix (`e51a8d2`) — same seed, same single-boot-per-architecture, all 9
identities simultaneously live. **Result: 81/81 trials correct, 0 mismatches**, shuffle sequence
still md5-identical to every prior run. Total wall-clock: amd64 ~151s, aarch64 ~276s, riscv64
~165s.
`results-20260912-with-heartbeat-csv/``EXEC-STD79-DOE` now brackets its trial loop with
`HB-ON`/`HB-OFF` (`doe_log.c`'s per-heartbeat-tick physics/timing CSV logger), so
`scripts/extract_doe.sh`'s automatically-generated CSV finally carries real data for this
campaign — every prior run's extracted CSV was silently empty, since `g_doe_log_enabled`
defaults off and nothing had ever turned it on. Each directory holds both the raw log
(`*-doe-raw.log`) and its paired heartbeat CSV (`*-heartbeat.csv`, ~255-257 rows/architecture,
18 columns: tick_number, elapsed_ns, hot_word_count, avg_word_heat_q48, window_width,
apic_ticks, hera/hermes/artemis heat_q48, etc. — see `doe_log.c`'s own header comment for the
full column list). **Result: 27/27 trials correct on every architecture** (verified via the
campaign's most distinctive result markers — the two 18-19 digit `M*`/`M/MOD` values and the
`2147483648 2/` result all show exactly 27 occurrences, zero faults) — **not** verified via
exact substring/byte reconstruction, because turning HB-ON on for the whole run exposed a real,
if purely cosmetic, effect: the heartbeat tick's async CSV printer and the trial loop's own
console output share the same serial line with no locking between them, so CSV rows get spliced
mid-token into the trial output on the wire (confirmed live: a "missing" `DOE-RUN,0` line turned
out to be `DOE-RUN,` and `0 ,6 ,04,2` on two separate physical log lines with a full CSV row
printed in between). The underlying FORTH execution itself is unaffected — values are correct,
nothing is corrupted in memory — only the printed character stream interleaves, so downstream
tooling that wants an exact reconstruction of the trial-output stream from these three raw logs
needs to account for that (strip `[HADES][DOE ] ...` fragments and rejoin) rather than assume
one physical line is one logical print.
`analysis-20260912/` builds exactly that reconstruction and turns it into a real 3×9 factorial
analysis: `correlate_doe.py` labels every heartbeat-tick row with the trial that was active when
it printed, `combine.py` merges all three architectures into `combined.csv` (768 rows, Q48.16
fields decoded to floats), and `analysis.R` computes per-cell (architecture × identity) means/SD
and a two-way ANOVA for each of 12 telemetry metrics, rendering one boxplot SVG per metric. See
`analysis-20260912/README.md` for the pipeline and `analysis-20260912/ANALYSIS.md` for the
report itself — key findings: identity significantly affects word-heat/window-sizing metrics
(expected — different identities execute different word sets), architecture significantly
affects timing metrics (APIC ticks/tick, timing variance, fleet heat — expected, different QEMU
targets), and **zero architecture × identity interaction on any metric** — the two factors'
effects are clean and separable, not confounded.