Files
LithosAnanake/docs/working/architecture/DOE-LIBRARY-HOWTO-20260819.md
Robert Allan JamesandClaude Sonnet 5 d0a76420a5 Add DoE library HOWTO (cookbook entry 2), flag stale L8-DOE claims in bare_metal/README.md
Documents capsules/doe.4th's word-level DOE/EXEC-DOE entry points, CSV
format, and a verified-not-fixed caveat: the rep column doesn't track
actual repetition count when n-reps differs from the file's fixed N-REPS=30
constant (cfg is unaffected, only rep is misleading -- use run_id instead).

Also surfaces, but does not fix, a real staleness finding in
experiments/bare_metal/README.md: its documented L8-DOE/WL-HI/WL-LO
entry point and workload-dispatch mechanism does not exist anywhere in the
current capsule set. What "DoE" actually names today is three separate
mechanisms (doe.4th's word-level DOE, doe-campaign.4th's fleet-touch
campaigns, and artemis/init.4th's auto-run ART-STRESS-CAMPAIGN) -- this
HOWTO documents only the first, per explicit scope decision.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-19 05:11:12 -04:00

7.6 KiB
Raw Permalink Blame History

DoE Library HOWTO — capsules/doe.4th

Status: WORKING. Second cookbook entry, following TURTLE-GRAPHICS-HOWTO-20260819.md. Scope decided explicitly 2026-08-19: document doe.4th's word-level DoE only — see "Which DoE?" below for why.

Which DoE?

The name "DoE" currently refers to three separate, unrelated mechanisms in this repo, not one. Before writing this HOWTO, experiments/bare_metal/ README.md (marked "mandatory read before touching capsules") was checked against the actual current capsule set and found stale on exactly this point — it describes an L8-DOE ( seed reps -- ) entry point with 16 WL-LO/WL-HI workload-dispatch slots, auto-invoked from init.4th. None of L8-DOE, WL-HI, or WL-LO exist anywhere in capsules/ today, and init.4th does not call any DoE mechanism — it only loads lib.4th, fabric.4th, font.4th, and prints the boot banner. This is flagged here, not fixed — correcting README.md is a separate task.

What actually exists, as of 2026-08-19:

Mechanism File Entry point What it does
Word-level DoE (this HOWTO) doe.4th DOE / EXEC-DOE ( seed n-reps -- ) A single embedded arithmetic workload (DOE-WORK), run across the 16 L8 factor configurations, streaming a CSV row per run to serial. Not auto-run anywhere.
Compudynamics fleet campaign doe-campaign.4th CAMPAIGN / SMOKE-CAMPAIGN / THREE-VM-CAMPAIGN Spawns Hermes/Artemis and drives real VM-EXEC touches between them to measure fleet heat conservation (VM-CONSERVED?). Not auto-run anywhere.
Artemis stress campaign capsules/artemis/init.4th ART-STRESS-CAMPAIGN Runs unconditionally at the bottom of the file, so it fires automatically every time Artemis is born. This is the actual source of the live [Artemis][HADES][DOE ] CSV rows visible during every kernel boot — unrelated to either mechanism above, and the subject of FABRIC-2.md's item 4.6 fix.

Only the first is a self-contained "package/library" in the sense the cookbook wants — a capsule you load and call with your own parameters, not a multi-VM orchestration script.

What doe.4th is

A word-level DoE harness measuring L8 Jacquard mode selector behavior across a 2⁴ full-factorial design — four boolean factors (entropy, CV, temporal decay, stability), 16 configurations, run some number of reps per configuration in Fisher-Yates shuffled order, streaming one CSV row per run to serial via [HADES][DOE ]-style output. This is the same measurement approach experiments/bare_metal/README.md's CSV-format section documents correctly (its 15-column heartbeat-tick description is a different, lower-level CSV — see "Two different CSVs" below) — only the entry point and auto-invocation claims in that doc are wrong.

Loading and running it

S" doe.4th" EXEC
DOE

DOE is 12345 3 EXEC-DOE — a fixed convenience call (seed 12345, 3×16=48 total runs). For a custom seed/run-count:

S" doe.4th" EXEC
54321 5 EXEC-DOE   ( seed=54321, 5*16=80 total runs )

Verified on the hosted build (PLOT/framebuffer concerns don't apply here — this is pure arithmetic and serial text output, works identically hosted and kernel-side): both calls above run to DOE: complete with zero VM errors, correct run counts (48 and 80 rows respectively, run_id columns confirm 047 and 079).

CSV format

CSV-HEADER (doe.4th block 2101):

run_id,cfg,rep,ent_in,cv_in,tmp_in,stb_in,
l8_mode,win_div,infer_win,infer_dec_q,
infer_var_q,early_exit,bc_mean_q,bb_mean_q,fit_q

16 columns (the line breaks above are doe.4th's own CRLFs inside CSV-HEADER; the emitted header is one logical row — this is intentional multi-line output, not a formatting error in the source).

Column Meaning
run_id Sequential run counter, 0 to (n-reps × 16) 1.
cfg Which of the 16 factor configurations (015, bit-decoded: b3=entropy, b2=CV, b1=temporal decay, b0=stability).
rep See "Known caveat" below — does not mean "which repetition, 0 to n-reps1."
ent_in/cv_in/tmp_in/stb_in The Q48.16 factor values actually applied for this run (0 or the fixed HI constant per factor).
l8_mode L8 Jacquard selector's resulting mode after L8-UPDATE/L8-APPLY.
win_div WINDOW-DIVERSITY at end of run.
infer_win/infer_dec_q/infer_var_q/early_exit/fit_q INFER-RUN's output accessors (INFER-WINDOW@/INFER-DECAY@/INFER-VARIANCE@/INFER-EARLY-EXIT@/INFER-FIT@ — the same words covered by this session's earlier inference_words_test.c, Module 26 POST coverage).
bc_mean_q/bb_mean_q BAYES-CACHE-MEAN / BAYES-BUCKET-MEAN.

Known caveat — the rep column

RUN-MATRIX is allocated and shuffled across a fixed N-RUNS = 480 cells (N-CFG=16 × the compile-time N-REPS=30 constant), regardless of what n-reps value is actually passed to EXEC-DOE. EXEC-DOE's own loop correctly runs n-reps × 16 times (confirmed above — a 5-rep call produces exactly 80 rows), but each iteration reads I MATRIX@ from the full 480-cell shuffled range and decodes rep as val MOD 30. The result: cfg is correctly uniform across all 16 configurations regardless of n-reps, but rep is a essentially-random value in 029 rather than a genuine "which repetition" counter — verified directly: a 5-rep run (54321 5 EXEC-DOE) produced rep values including 27, 21, 24, 18 in its output, not values constrained to 04. This is existing production behavior in doe.4th, not something introduced or fixed here — reported per this repo's "report bugs, don't fix unless asked" rule. If you need a trustworthy per-config repetition index from a CSV, use run_id (unique, sequential) or compute your own from row order, not rep.

Two different CSVs

experiments/bare_metal/README.md's "CSV Format" section (15 columns: tick_number, elapsed_ns, ... variance_q48) documents a different, lower-level CSV — one heartbeat-tick row per [HADES][DOE ] line, emitted by the C-side heartbeat/DoE metrics machinery (src/heartbeat_export.c/doe_metrics.c), independent of which FORTH-level DoE mechanism (if any) is driving execution at the time. doe.4th's own 16-column per-run CSV (documented above) is emitted separately, directly by EMIT-ROW, using plain TYPE/EMIT to serial — the two coexist in the same log stream but answer different questions ("what did the timing look like this tick" vs. "what did this whole DoE run measure").

Verification performed

  • Hosted-build trace (./build/amd64/standard/starforth -s --log-error, doe.4th's blocks piped in with Block NNNN headers stripped, matching the same methodology used for turtle.4th): DOE (default 12345 3) and 54321 5 EXEC-DOE both complete with DOE: complete, zero VM errors, correct row counts (48 and 80).
  • Column count and header/row alignment checked directly against CSV-HEADER's literal text (16 fields, header and rows agree).
  • The rep-column caveat above is empirical, not inferred from reading the source alone — confirmed by comparing actual emitted rep values across two different n-reps calls.
  • Not yet re-verified against a live kernel boot's serial log (the hosted trace already exercises the identical FORTH code path; doe.4th uses no kernel-only words, so no further boot-side verification is expected to change this).