These are alternate boot personality capsules, not files doe.4th dispatches from -- confirmed by reading doe.4th itself (it generates its own synthetic workload internally, DOE-WORK) and mkcapsule.c (only the exact filename "init.4th" is special-cased as the active MAMA_INIT capsule). "workload-N.4th" names them for what they are without colliding with that reserved name. Renamed the 10 files (git mv, preserving history) and their own self-referential header comments (also fixed a pre-existing typo, "init-4.th" -> "workload-4.4th"), updated capsules/README.md, capsules/MANIFEST.md (21 references), .claude/CLAUDE.md, and a tools/mkcapsule.c comment. Also fixed experiments/bare_metal/README.md's "Adding a Custom Workload Capsule" section, discovered stale while doing this rename: it documented a WL-HI/WL-LO dispatch table and a wl_id CSV column that don't exist anywhere in the current capsules/ tree or doe.4th's own CSV header -- corrected to describe what's actually there (no pluggable workload dispatch; a custom workload is run by substituting it in as the boot's own init.4th). Verified: mkcapsule --lint clean (35 files, 0 violations), all 3 architectures build with zero new warnings, amd64 boots to zuse)ok> with an unchanged dict_hash/capsule_hash from every prior boot this session (0xc8f4b09e36f4fc4a / 0x1ef4939ed32ec1e6) and zero faults. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
296 lines
9.8 KiB
Markdown
296 lines
9.8 KiB
Markdown
# Bare-Metal DoE Experiment
|
||
|
||
This directory holds data and analysis from the LithosAnanke kernel's
|
||
Design of Experiments (DoE) runs — blind full-factorial 2⁴ experiments
|
||
that measure the L8 Jacquard mode selector's effect on the Steady-State
|
||
Machine across all three supported architectures (amd64, aarch64, riscv64).
|
||
|
||
---
|
||
|
||
## Directory Layout
|
||
|
||
```
|
||
experiments/bare_metal/
|
||
├── runs/ ← timestamped canonical CSVs from every acceptance run
|
||
├── latest/ ← arch-named copies of the most recent run (human-readable)
|
||
│ ├── amd64.csv
|
||
│ ├── aarch64.csv
|
||
│ └── riscv64.csv
|
||
└── analysis/
|
||
├── charts/ ← generated SVG/PNG charts
|
||
├── report/ ← timestamped LaTeX / Markdown reports
|
||
└── tables/ ← generated summary tables
|
||
```
|
||
|
||
`runs/` is the canonical archive. `latest/` is the eyeball-friendly shortcut
|
||
— always the most recent run per architecture, overwritten on each new run.
|
||
|
||
---
|
||
|
||
## Running the DoE
|
||
|
||
The DoE runs automatically when the kernel boots because `init.4th` calls it.
|
||
The standard acceptance command runs all three architectures sequentially:
|
||
|
||
```bash
|
||
make -f Makefile.starkernel ARCH=amd64 clean qemu
|
||
make -f Makefile.starkernel ARCH=aarch64 clean qemu
|
||
make -f Makefile.starkernel ARCH=riscv64 clean qemu
|
||
```
|
||
|
||
Each run executes 48 trials (16 L8 configs × 3 reps, Fisher-Yates shuffled),
|
||
captures ~1.27 million heartbeat rows per architecture, and writes two files:
|
||
|
||
| File | Path |
|
||
|------|------|
|
||
| Timestamped canonical CSV | `experiments/bare_metal/runs/doe-<arch>-<YYYYMMDD-HHMMSS>.csv` |
|
||
| Latest convenience copy | `experiments/bare_metal/latest/<arch>.csv` |
|
||
|
||
**Run architectures sequentially, never in parallel.** All three QEMU
|
||
instances use `accel=tcg` (software emulation). Concurrent runs compete for
|
||
host CPU and corrupt the timing signal that the DoE is measuring.
|
||
|
||
---
|
||
|
||
## Disabling the DoE
|
||
|
||
To boot into the REPL without running the experiment, comment out the last
|
||
two lines of `capsules/init.4th`:
|
||
|
||
```forth
|
||
Block 2049
|
||
( first init.4th )
|
||
: STAR 42 EMIT ;
|
||
: STARS 0 DO STAR LOOP ;
|
||
: MARGIN 30 SPACES ;
|
||
: BAR MARGIN 5 STARS CR ;
|
||
: BLIP MARGIN STAR CR ;
|
||
: F CR BAR BLIP BAR BLIP BLIP CR ;
|
||
( S" Hermes" BIRTH )
|
||
( S" Artemis" BIRTH )
|
||
( S" doe.4th" EXEC ) ← comment this out
|
||
( 123456 3 L8-DOE ) ← comment this out
|
||
```
|
||
|
||
The kernel will boot to the `ok>` REPL with no experiment running.
|
||
Commenting both lines leaves `doe.4th` unloaded so none of its words
|
||
(`L8-DOE`, `WL-NAME`, etc.) are defined, which is the cleanest state for
|
||
interactive sessions.
|
||
|
||
---
|
||
|
||
## Changing the Seed and Rep Count
|
||
|
||
The DoE entry point is `L8-DOE ( seed reps -- )`.
|
||
The call in `init.4th` is:
|
||
|
||
```forth
|
||
123456 3 L8-DOE
|
||
```
|
||
|
||
- **Seed** — any non-zero integer. The same seed always produces the same
|
||
shuffled run order, so results are reproducible. Change the seed to
|
||
explore a different permutation; different seeds are statistically
|
||
equivalent but verify shuffle-independence.
|
||
|
||
- **Reps** — trials per L8 configuration (1–200). 3 reps × 16 configs =
|
||
48 runs, which takes roughly 25–30 minutes per architecture under TCG.
|
||
Increase for higher statistical power; decrease for quick smoke checks.
|
||
|
||
```forth
|
||
( quick smoke check — 1 rep, 16 runs total )
|
||
42 1 L8-DOE
|
||
|
||
( full study — 10 reps, 160 runs )
|
||
987654 10 L8-DOE
|
||
```
|
||
|
||
---
|
||
|
||
## What Are Capsules?
|
||
|
||
A **capsule** is a named blob of FORTH-79 source text stored in `capsules/`.
|
||
The kernel's `EXEC` word loads a capsule by filename and interprets it as
|
||
FORTH source. `BIRTH` (commented out in `init.4th`) would instead spawn an
|
||
isolated child VM whose sole personality is that capsule's code.
|
||
|
||
There are two roles:
|
||
|
||
| Role | Who uses it | What it does |
|
||
|------|-------------|--------------|
|
||
| **Init capsule** | Mama VM at boot | Defines the VM's vocabulary and behavior |
|
||
| **Workload capsule** | DoE machinery | Provides a computational task to time |
|
||
|
||
`init.4th` is the Mama VM's init capsule — executed exactly once at kernel
|
||
boot, and the only file `mkcapsule` special-cases by exact filename as the
|
||
active `MAMA_INIT` capsule. The numbered files (`workload-0.4th` …
|
||
`workload-9.4th`) and the L8 variant files (`init-l8-*.4th`) are alternate
|
||
personality capsules, not a workload dispatched by `doe.4th`'s own DoE —
|
||
`doe.4th` generates its own synthetic arithmetic workload internally
|
||
(`DOE-WORK`) and has no pluggable per-workload dispatch table today. To
|
||
exercise one of the numbered capsules' own workload, substitute it in as
|
||
the boot's `init.4th` (e.g. `cp capsules/workload-3.4th capsules/init.4th`
|
||
before building) rather than wiring it into `doe.4th`.
|
||
|
||
---
|
||
|
||
## `.4th` File Structure
|
||
|
||
Every `.4th` file must follow StarForth's block format. The block system
|
||
maps source text to 1024-byte logical blocks; the `Block NNNN` header tells
|
||
the loader which block slot to fill.
|
||
|
||
**Mandatory rules:**
|
||
|
||
1. The first line of each logical block must be `Block NNNN` (capital B,
|
||
single space, decimal integer).
|
||
2. Block numbers must be unique within a single capsule file.
|
||
3. Blocks are loaded in file order and executed top-to-bottom.
|
||
4. Each block can hold up to 1024 bytes of source text.
|
||
5. Comments use `( ... )` — parentheses with spaces inside.
|
||
6. Word definitions use `: NAME ... ;` — standard FORTH-79.
|
||
|
||
**Minimal capsule skeleton:**
|
||
|
||
```forth
|
||
Block 3100
|
||
( My capsule description )
|
||
|
||
: MY-WORD ( -- )
|
||
42 . CR ;
|
||
|
||
MY-WORD
|
||
```
|
||
|
||
**Multi-block capsule:**
|
||
|
||
```forth
|
||
Block 3100
|
||
( Block 1: helpers )
|
||
: HELPER ( n -- n*2 ) 2 * ;
|
||
|
||
Block 3101
|
||
( Block 2: main logic )
|
||
: MAIN ( -- )
|
||
10 0 DO I HELPER . CR LOOP ;
|
||
|
||
MAIN
|
||
```
|
||
|
||
The block number namespace is shared across all loaded capsules.
|
||
Convention used in this repository:
|
||
|
||
| Range | Contents |
|
||
|-------|----------|
|
||
| 2048–2099 | init.4th (Mama VM boot sequence) |
|
||
| 2100–2199 | doe.4th (DoE machinery) |
|
||
| 3000–3999 | Reserved for user-defined workload capsules (see below) |
|
||
| 4000+ | User-defined capsules |
|
||
|
||
---
|
||
|
||
## Adding a Custom Workload Capsule
|
||
|
||
**Step 1 — Create the file.**
|
||
|
||
Add `capsules/my-workload.4th` using block numbers in the 4000+ range:
|
||
|
||
```forth
|
||
Block 4000
|
||
( my-workload.4th - description of what this measures )
|
||
|
||
: MY-COMPUTE ( n -- )
|
||
0 SWAP 0 DO I 3 * + LOOP DROP ;
|
||
|
||
Block 4001
|
||
( main entry point )
|
||
: RUN-MY-WORKLOAD ( -- )
|
||
500 0 DO I MY-COMPUTE LOOP ;
|
||
|
||
RUN-MY-WORKLOAD
|
||
```
|
||
|
||
The last line should execute the workload so `EXEC` runs it immediately when
|
||
the capsule is loaded.
|
||
|
||
**Step 2 — Run it.**
|
||
|
||
`doe.4th` has no pluggable workload dispatch table today — its own
|
||
factorial (entropy/CV/temporal-decay/stability × reps) drives a single
|
||
self-contained synthetic workload (`DOE-WORK`, Block 2103), not a file
|
||
picked from a list. To measure your own workload's physics behavior
|
||
instead, run it standalone as the boot's own capsule rather than trying
|
||
to wire it into `doe.4th`'s factorial:
|
||
|
||
```bash
|
||
cp capsules/my-workload.4th capsules/init.4th # substitutes it in
|
||
make -f Makefile.starkernel ARCH=amd64 clean qemu
|
||
```
|
||
|
||
(`mkcapsule` special-cases the exact filename `init.4th` as the one
|
||
capsule flagged `MAMA_INIT` and loaded at boot — this is genuinely a
|
||
substitution, not a selection from a list, so restore the real `init.4th`
|
||
afterward, e.g. `git checkout capsules/init.4th`.)
|
||
|
||
Your workload's heartbeat rows will appear in the CSV the same way any
|
||
boot's do — match by the `DOE-RUN` marker lines described below, once your
|
||
capsule's own trial-loop word emits one the same way `doe.4th`'s does.
|
||
|
||
---
|
||
|
||
## CSV Format
|
||
|
||
Each row emitted by the `[HADES][DOE ]` serial tag is one heartbeat tick
|
||
during a workload execution. Extract with:
|
||
|
||
```bash
|
||
grep -aP '\[HADES\]\[DOE \]' logs2/qemu-amd64-<timestamp>.log \
|
||
| sed 's/.*\[DOE \] //' > my.csv
|
||
```
|
||
|
||
Columns (15 total):
|
||
|
||
| # | Name | Type | Description |
|
||
|---|------|------|-------------|
|
||
| 1 | `tick_number` | uint32 | Monotonic heartbeat counter |
|
||
| 2 | `elapsed_ns` | uint64 | Nanoseconds since run start |
|
||
| 3 | `tick_interval_ns` | uint64 | Interval from prior tick |
|
||
| 4 | `cache_hits_delta` | uint32 | Hot-words cache hits this tick |
|
||
| 5 | `bucket_hits_delta` | uint32 | Bucket hits this tick |
|
||
| 6 | `word_executions_delta` | uint32 | Words executed this tick |
|
||
| 7 | `hot_word_count` | uint64 | Words with heat ≥ threshold |
|
||
| 8 | `avg_word_heat_q48` | uint64 | Mean heat (raw Q48.16 integer) |
|
||
| 9 | `window_width` | uint32 | L8's target rolling window size |
|
||
| 10 | `actual_window_size` | uint32 | True analysis width: min(total_executions, window_width) |
|
||
| 11 | `predicted_label_hits` | uint32 | ANOVA early-exit confirmations (L8 validation signal) |
|
||
| 12 | `jitter_bits` | uint64 | Estimated jitter (IEEE 754 bit pattern) |
|
||
| 13 | `apic_ticks` | uint64 | APIC timer monotonic count |
|
||
| 14 | `time_trust_q48` | uint64 | Time-trust score (Q48.16) |
|
||
| 15 | `variance_q48` | uint64 | Timing variance (Q48.16) |
|
||
|
||
`avg_word_heat_q48` is a raw fixed-point integer. To convert to a human-readable
|
||
heat value: `avg_word_heat = avg_word_heat_q48 / 65536.0`.
|
||
|
||
`jitter_bits` is the IEEE 754 double-precision bit pattern of the jitter in
|
||
nanoseconds. In R: `readBin(as.raw(…), "double")`. In Python:
|
||
`struct.unpack('d', struct.pack('Q', n))[0]`.
|
||
|
||
---
|
||
|
||
## Interpreting `predicted_label_hits`
|
||
|
||
This column is the feedback-loop closure signal.
|
||
|
||
Each non-zero value means the inference engine ran ANOVA on the current
|
||
execution window and confirmed the L8 selector's config choice correlated
|
||
with the subsequent execution pattern — an "early exit" because the
|
||
statistical test converged without needing all data.
|
||
|
||
- **High rate** → L8 chose well; the system settled quickly into a stable regime.
|
||
- **Low rate** → L8 is still searching; the workload is novel or transient.
|
||
- **Zero throughout** → The workload ended before the inference engine had
|
||
enough data, or the window is too small to trigger ANOVA.
|
||
|
||
This is the metric that closes the loop between "L8 made a choice" and
|
||
"that choice was actually validated by what the VM did next."
|