Documents capsules/doe.4th's word-level DOE/EXEC-DOE entry points, CSV format, and a verified-not-fixed caveat: the rep column doesn't track actual repetition count when n-reps differs from the file's fixed N-REPS=30 constant (cfg is unaffected, only rep is misleading -- use run_id instead). Also surfaces, but does not fix, a real staleness finding in experiments/bare_metal/README.md: its documented L8-DOE/WL-HI/WL-LO entry point and workload-dispatch mechanism does not exist anywhere in the current capsule set. What "DoE" actually names today is three separate mechanisms (doe.4th's word-level DOE, doe-campaign.4th's fleet-touch campaigns, and artemis/init.4th's auto-run ART-STRESS-CAMPAIGN) -- this HOWTO documents only the first, per explicit scope decision. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
7.6 KiB
DoE Library HOWTO — capsules/doe.4th
Status: WORKING. Second cookbook entry, following
TURTLE-GRAPHICS-HOWTO-20260819.md. Scope decided explicitly 2026-08-19:
document doe.4th's word-level DoE only — see "Which DoE?" below for why.
Which DoE?
The name "DoE" currently refers to three separate, unrelated mechanisms
in this repo, not one. Before writing this HOWTO, experiments/bare_metal/ README.md (marked "mandatory read before touching capsules") was checked
against the actual current capsule set and found stale on exactly this
point — it describes an L8-DOE ( seed reps -- ) entry point with 16
WL-LO/WL-HI workload-dispatch slots, auto-invoked from init.4th. None
of L8-DOE, WL-HI, or WL-LO exist anywhere in capsules/ today, and
init.4th does not call any DoE mechanism — it only loads lib.4th,
fabric.4th, font.4th, and prints the boot banner. This is flagged here,
not fixed — correcting README.md is a separate task.
What actually exists, as of 2026-08-19:
| Mechanism | File | Entry point | What it does |
|---|---|---|---|
| Word-level DoE (this HOWTO) | doe.4th |
DOE / EXEC-DOE ( seed n-reps -- ) |
A single embedded arithmetic workload (DOE-WORK), run across the 16 L8 factor configurations, streaming a CSV row per run to serial. Not auto-run anywhere. |
| Compudynamics fleet campaign | doe-campaign.4th |
CAMPAIGN / SMOKE-CAMPAIGN / THREE-VM-CAMPAIGN |
Spawns Hermes/Artemis and drives real VM-EXEC touches between them to measure fleet heat conservation (VM-CONSERVED?). Not auto-run anywhere. |
| Artemis stress campaign | capsules/artemis/init.4th |
ART-STRESS-CAMPAIGN |
Runs unconditionally at the bottom of the file, so it fires automatically every time Artemis is born. This is the actual source of the live [Artemis][HADES][DOE ] CSV rows visible during every kernel boot — unrelated to either mechanism above, and the subject of FABRIC-2.md's item 4.6 fix. |
Only the first is a self-contained "package/library" in the sense the cookbook wants — a capsule you load and call with your own parameters, not a multi-VM orchestration script.
What doe.4th is
A word-level DoE harness measuring L8 Jacquard mode selector behavior
across a 2⁴ full-factorial design — four boolean factors (entropy, CV,
temporal decay, stability), 16 configurations, run some number of reps per
configuration in Fisher-Yates shuffled order, streaming one CSV row per run
to serial via [HADES][DOE ]-style output. This is the same measurement
approach experiments/bare_metal/README.md's CSV-format section documents
correctly (its 15-column heartbeat-tick description is a different,
lower-level CSV — see "Two different CSVs" below) — only the entry point
and auto-invocation claims in that doc are wrong.
Loading and running it
S" doe.4th" EXEC
DOE
DOE is 12345 3 EXEC-DOE — a fixed convenience call (seed 12345,
3×16=48 total runs). For a custom seed/run-count:
S" doe.4th" EXEC
54321 5 EXEC-DOE ( seed=54321, 5*16=80 total runs )
Verified on the hosted build (PLOT/framebuffer concerns don't apply here
— this is pure arithmetic and serial text output, works identically hosted
and kernel-side): both calls above run to DOE: complete with zero VM
errors, correct run counts (48 and 80 rows respectively, run_id columns
confirm 0–47 and 0–79).
CSV format
CSV-HEADER (doe.4th block 2101):
run_id,cfg,rep,ent_in,cv_in,tmp_in,stb_in,
l8_mode,win_div,infer_win,infer_dec_q,
infer_var_q,early_exit,bc_mean_q,bb_mean_q,fit_q
16 columns (the line breaks above are doe.4th's own CRLFs inside
CSV-HEADER; the emitted header is one logical row — this is intentional
multi-line output, not a formatting error in the source).
| Column | Meaning |
|---|---|
run_id |
Sequential run counter, 0 to (n-reps × 16) − 1. |
cfg |
Which of the 16 factor configurations (0–15, bit-decoded: b3=entropy, b2=CV, b1=temporal decay, b0=stability). |
rep |
See "Known caveat" below — does not mean "which repetition, 0 to n-reps−1." |
ent_in/cv_in/tmp_in/stb_in |
The Q48.16 factor values actually applied for this run (0 or the fixed HI constant per factor). |
l8_mode |
L8 Jacquard selector's resulting mode after L8-UPDATE/L8-APPLY. |
win_div |
WINDOW-DIVERSITY at end of run. |
infer_win/infer_dec_q/infer_var_q/early_exit/fit_q |
INFER-RUN's output accessors (INFER-WINDOW@/INFER-DECAY@/INFER-VARIANCE@/INFER-EARLY-EXIT@/INFER-FIT@ — the same words covered by this session's earlier inference_words_test.c, Module 26 POST coverage). |
bc_mean_q/bb_mean_q |
BAYES-CACHE-MEAN / BAYES-BUCKET-MEAN. |
Known caveat — the rep column
RUN-MATRIX is allocated and shuffled across a fixed N-RUNS = 480
cells (N-CFG=16 × the compile-time N-REPS=30 constant), regardless of
what n-reps value is actually passed to EXEC-DOE. EXEC-DOE's own loop
correctly runs n-reps × 16 times (confirmed above — a 5-rep call
produces exactly 80 rows), but each iteration reads I MATRIX@ from the
full 480-cell shuffled range and decodes rep as val MOD 30. The
result: cfg is correctly uniform across all 16 configurations regardless
of n-reps, but rep is a essentially-random value in 0–29 rather than
a genuine "which repetition" counter — verified directly: a 5-rep run
(54321 5 EXEC-DOE) produced rep values including 27, 21, 24, 18
in its output, not values constrained to 0–4. This is existing
production behavior in doe.4th, not something introduced or fixed here —
reported per this repo's "report bugs, don't fix unless asked" rule. If you
need a trustworthy per-config repetition index from a CSV, use run_id
(unique, sequential) or compute your own from row order, not rep.
Two different CSVs
experiments/bare_metal/README.md's "CSV Format" section (15 columns:
tick_number, elapsed_ns, ... variance_q48) documents a different,
lower-level CSV — one heartbeat-tick row per [HADES][DOE ] line, emitted
by the C-side heartbeat/DoE metrics machinery
(src/heartbeat_export.c/doe_metrics.c), independent of which FORTH-level
DoE mechanism (if any) is driving execution at the time. doe.4th's own
16-column per-run CSV (documented above) is emitted separately, directly by
EMIT-ROW, using plain TYPE/EMIT to serial — the two coexist in the
same log stream but answer different questions ("what did the timing look
like this tick" vs. "what did this whole DoE run measure").
Verification performed
- Hosted-build trace (
./build/amd64/standard/starforth -s --log-error,doe.4th's blocks piped in withBlock NNNNheaders stripped, matching the same methodology used forturtle.4th):DOE(default12345 3) and54321 5 EXEC-DOEboth complete withDOE: complete, zero VM errors, correct row counts (48 and 80). - Column count and header/row alignment checked directly against
CSV-HEADER's literal text (16 fields, header and rows agree). - The
rep-column caveat above is empirical, not inferred from reading the source alone — confirmed by comparing actual emittedrepvalues across two differentn-repscalls. - Not yet re-verified against a live kernel boot's serial log (the hosted
trace already exercises the identical FORTH code path;
doe.4thuses no kernel-only words, so no further boot-side verification is expected to change this).