# FABRIC-3.md — the Stadium, continued again **Status:** Living working document, opened 2026-08-25 as the successor to `FABRIC-2.md` (now closed/archival — see its own header). This document does not repeat `FABRIC-2.md`'s design argument or history; it restates only outcomes, with pointers back to the section that derived them. Read `FABRIC.md` for the original "why," `FABRIC-2.md` for everything derived through 2026-08-25, this document for what's left as of that date onward. **Provenance.** Everything in Section A below is a full, non-sampled carry-forward of every open (`- [ ]`) item in `FABRIC-2.md` as of 2026-08-25 — 51 items, confirmed by `grep -c '^- \[ \]' FABRIC-2.md`, none dropped (Section A itself holds 49: the other 2, `FABRIC-2.md` §F.3's own two checkbox lines, were pure summaries cross-referencing items already listed individually elsewhere — 4.4s/1.11/4.3/§17.4 and 5.1/ACL-RWT re-measurement, both of which are carried forward as their own individual items above — not distinct content, confirmed by diffing item text programmatically before writing this document, not assumed). Extracted mechanically (a script pulling each checkbox item's own text, stopping at the first blank line rather than the next checkbox, to avoid pulling in unrelated already-resolved narrative that happened to sit between two open items in the source document) and spot-checked against the original. Item numbers/labels are carried forward unchanged, for traceability — this is not a renumbering or a re-prioritization. Section groupings match `FABRIC-2.md`'s own (documentation debt, xHCI WRITE(10), Milestone 3–9 punch lists, etc.) — items are relocated, not reorganized. **How to use this document going forward.** New findings, new punch-list items, and new decisions get added here, not to `FABRIC-2.md`. Follow the same discipline `FABRIC-2.md` §(intro) established for how work gets picked up, closed, and recorded. --- ## A. Carried forward from FABRIC-2.md (51 items, all still open as of 2026-08-25) ### From FABRIC-2.md §A — Blocked or scoped, not started - [ ] **1.11 — Dirty-event granularity.** Leaning region-based. Blocked on item 4.3 — settled as part of the console migration, not speculatively before it. *Refs (FABRIC.md):* §17.5, §23.2, §23.4 #1. - [ ] **4.3 — Console.** Umbrella item; settles 1.11 as part of the work. Nearly everything under it (4.3.1–4.3.7f, 4.4–4.4ac) is done — the parent stays open only because 4.4s below is still blocked and nothing has formally closed the umbrella. *Refs (FABRIC.md):* §17.5, §27. - [ ] **4.4s — `(user)` prompt segment.** Scoped, blocked, not started. Extends 4.4's prompt format. *Refs (FABRIC.md):* §27.8, 4.4. - [ ] **5.1 — Re-run the DoE on the new substrate.** A green POST suite is not evidence that determinism holds under the Stadium migration — needs its own campaign. Not started. ### From FABRIC-2.md §D — Design questions still genuinely open - [ ] **§17.4 — the framebuffer utility's internal heat/decay dynamics are undesigned.** Explicitly "Open, deferred": not a Stadium patron, but what physics (if any) governs it internally was never designed. Not blocking anything. Blocked on item 1.11 specifically (dirty-event granularity), not "the framebuffer work" in general — see `FABRIC-2.md` §D's own 2026-08-13 refinement of this item before assuming it's ripe. ### From FABRIC-2.md §E — Documentation debt - [x] **Already done — stale carry-forward, closed 2026-08-26.** This item was never marked complete when the actual work concluded. `FABRIC-2.md` Sections M through T (2026-08-19 to 2026-08-21) already re-ran this measurement — and in the process found the original "ACL-RWT" name itself was wrong: the Rolling-Window-of-Truth mechanism it was named after was dead code, removed 2026-07-08 (`ACL-RECHECK-RW` was never reachable — `acl_recheck()` only ever looks up the 11-char `ACL-RECHECK`, never the 14-char `-RW` variant). What the original campaign actually measured was the live `ACL-TTL` mechanism under a misleading name. **Final, accepted result (`FABRIC-2.md` §T): +0.0603% ACL-TTL enforcement overhead, architecture-independent (identical across amd64/aarch64/riscv64), fully deterministic (CV=0.000%)** — see also memory `project_acl_ttl_overhead_final.md`. No new campaign run here; this entry corrects the bookkeeping, not the measurement. ### From FABRIC-2.md §J — Maintainability sweep (2026-08-18) - [ ] `docs/lithosananke/ROADMAP.md` and `M7.1.md` — stale `Branch: lithosananke` (no such branch exists post-split), `M7.1.md`'s "Status: Design Complete" (shipped and live, not just designed), `ROADMAP.md`'s self-contradiction (M8 marked OBSOLETE in one place, still a live success criterion in another), and its stale "AHCI driver" claim for M9 (real implementation is `virtio_blk.c`) — not fixed, flagged. - [ ] Top-level `ROADMAP.md` (StarForth-era, "Phase 0 Complete... Phase 1 Starting," dated 2025-12-14) — badly stale, no historical/superseded banner to warn a reader. Not fixed. - [ ] `docs/03-architecture/word-acl/DESIGN.md` says ACL Phase 7 (LithosAnanke kernel parity) is still "remaining" — direct contradiction with `.claude/CLAUDE.md`, which states Phase 7 is independently verified complete. Not fixed. - [x] Confirmed accurate, not stale (2026-08-26): `VM-FLEET-ATTRACTOR-DESIGN-20260705.md`'s claim that `doe-campaign.4th` is "broken and being superseded." Live-ran `SMOKE-CAMPAIGN` from the current capsule (amd64) — it completes without error and correctly conserves fleet heat (`VM-PHYSICS: conserved=CONSERVED`, `fleet_heat_sum=65536`). Initially read that as contradicting the doc's claim; **Captain Bob corrected this directly — it's broken.** "Doesn't crash" and "runs" are not the same claim: the doc's actual argument is that the capsule has no real controlled-experimental-factor mechanism (no manual heat-injection point under the current design, so it cannot drive the fleet through controlled scenarios the way a DoE campaign needs to), a methodological gap a clean execution trace doesn't surface or disprove. Doc's claim stands; not touched. - [x] Tracked (2026-08-26) — real proof-modeling gap, not just stale prose, so not fixed here. Confirmed by reading `StarForth_Loop4_Pipeline.thy`'s own comment (lines 127–137): `pipeline_metrics_state`'s `pm_last_accuracy_num`/`pm_last_accuracy_den` fraction pair doesn't correspond to anything in the real C struct — `include/vm.h`'s `PipelineGlobalMetrics` has a single `double last_checked_accuracy` field, no num/den pair anywhere. The `.thy` file's own comment already scopes the real fix correctly: "a full field-level pass over `pipeline_metrics_state` is its own separate task" — matches this project's standing caution that each remaining Isabelle gap needs its own subsystem model, not a documentation-sprint patch. The punch-list ask was tracking this outside the buried `.thy` comment, which it now has here — the actual re-model stays unattempted, on purpose. ### From FABRIC-2.md §X, Milestone 2 — USB hardware stack - [ ] Decide and implement where the hotplug event surfaces to the rest of the kernel — likely a callback registered by whatever owns the home-blocks logic, not xHCI code calling into `block_subsystem.c` directly (matching the existing "kernel/Artemis decoupling boundary" pattern already documented in `block_subsystem.c`). **Partially addressed by Milestone 2h's `blkio_usb.c`/connect-time attach wiring (`FABRIC-2.md`, 2026-08-25) — worth re-checking whether that closes this item outright before treating it as still fully open.** - [ ] Implement CBW/data/CSW for SCSI WRITE(10) — this is where the earlier "read/write, unquestionable" requirement actually gets satisfied. Still the single biggest functional gap in the xHCI driver — blocks writing to a real USB thumb drive at all (`blkio_usb.c` is read-only today specifically because of this). - [ ] Implement basic error/stall recovery (CSW failure status, endpoint stall clear) — at minimum enough to not wedge the controller on a single bad transfer. ### From FABRIC-2.md §X, Milestone 3 — Block subsystem extensions - [ ] Implement the CA-signed-cert verification path (Milestone 6 dependency — the cert chain validator doesn't exist yet either). - [ ] Implement the first-touch allocation function: given a verified identity pubkey and a requested block count, either read an existing range from the drive's map or claim a new one at `g.total_user_lbn` and write it back. *(Single-block relocation itself — the mechanism this would allocate ranges for — is done: `blk_subsys_relocate_block()`/ `RELOCATE-BLOCK`, `FABRIC-2.md`, commit `36d832f`. This item is about the identity→range allocation that decides what to relocate blocks* into*, still unbuilt.)* - [ ] Design the on-drive block-map format (Section U item 4) — what it records (block ranges claimed? individual block liveness? something else), how it's serialized. - [ ] Implement writing the block-map to a drive. - [ ] Implement reading/validating the block-map from a drive on insertion. - [ ] Design the migration state machine (Section U item 5) — states, transition triggers. Session direction, 2026-08-25: **ACL manages *when* to relocate** (capacity pressure, or a compudynamics heat/cold signal); migration itself is expected to be rare, not routine. The `physics_hotwords_cache.c`-reuse question is settled differently than originally framed — see this document's new §B below (Stadium unification), which reframes block/word placement as a `compudynamics.c`-driven decision generically, not a `physics_hotwords_cache.c` (`DictEntry*`-hardcoded) reuse question specifically. - [ ] Decide and implement unclean-removal handling (Section U's explicitly flagged open question — never answered) — at minimum, detect a mid-flush disconnect via Milestone 2e's disconnect signal and decide what state that leaves affected blocks in. ### From FABRIC-2.md §X, Milestone 4 — Drive/credential security - [x] **Designed (2026-08-26), not yet implemented — Phase 8 kickoff.** Mirrors two existing precedents exactly: `CAPSULE_MAGIC_PACK`'s bit-packed magic (`include/starkernel/capsule.h`) and `blk_volume_meta_t`'s magic+version+fields+pad-to-4096 structural convention (`include/block_subsystem.h`). Lives at the first 4KiB devblock of the GPT metadata partition (the ~1GB partition decided 2026-08-22) — the header format doesn't depend on the still-missing GPT parser; it's just what gets written starting at that partition's first devblock once something can locate it. Deliberately narrow: identifies/authenticates the drive only, does **not** invent the block-map or credential/cert formats (both separate, still-open items below) — reserves offset/size pointers to where they'll live instead of embedding them. ```c #define HOMEBLOCKS_SIG_MAGIC 0x4248414CULL /* 'LAHB' -- LithosAnanke Home Blocks, * same little-endian ASCII packing as * CAPSULE_DESC_MAGIC's 'CAPS' */ #define HOMEBLOCKS_SIG_VERSION_0 0 #define HOMEBLOCKS_SIG_PACK(ver) \ (HOMEBLOCKS_SIG_MAGIC | ((uint64_t)(ver) << 32)) #define HOMEBLOCKS_SIG_GET_MAGIC(m) ((uint32_t)((m) & 0xFFFFFFFFULL)) #define HOMEBLOCKS_SIG_GET_VERSION(m) ((uint8_t)(((m) >> 32) & 0xFF)) typedef struct { uint64_t magic; /* HOMEBLOCKS_SIG_PACK(...) */ uint8_t drive_uuid[16]; /* unique per-mint instance id -- Phase 8 mints multiple * distinct drives, needs something to tell them apart */ uint64_t minted_time_ns; uint64_t metadata_devblocks; /* size of this GPT metadata partition, in 4KiB devblocks * -- sanity/bounds check against the GPT entry once a * parser exists */ uint32_t cert_offset; /* devblock offset within this partition where the * CA-signed cert blob starts; 0 = not yet minted */ uint32_t cert_devblocks; uint32_t blockmap_offset; /* devblock offset where the block-map (Milestone 3, * format still undesigned) starts */ uint32_t blockmap_devblocks; uint64_t hdr_crc; /* REAL from day one, not a placeholder like * blk_volume_meta_t's "unused yet" hdr_crc -- this * header's whole job is gating a warn-and-refuse * security check below, so the crc has to actually work */ uint8_t _pad[4096 - 64]; /* pad to one devblock; 64 = sum of the fields above */ } homeblocks_sig_t; ``` `cert_offset`/`cert_devblocks` are 0 on first design pass — the real CA hasn't been generated yet (Milestone 6), so this reserves the *shape* of where a cert will attach without committing to a cert format that doesn't exist. Same reasoning for `blockmap_offset`/`blockmap_devblocks` against Milestone 3's still-open block-map design. **Implemented (2026-08-26):** `include/starkernel/homeblocks_sig.h` — `homeblocks_sig_t` + `HOMEBLOCKS_SIG_PACK`/`_GET_MAGIC`/`_GET_VERSION` macros, a C99 compile-time size assertion (same discipline `stadium.h`'s own header-size checks use), and the identical field layout shown above. Verified standalone: `sizeof(homeblocks_sig_t) == 4096`, compiles clean under `-std=c99 -Wall -Wextra -Werror`. Not yet consumed by any code — nothing in the block subsystem or xHCI driver reads or writes it yet, so no functional kernel change and no 3-arch acceptance boot needed for this step; that starts with the signature-check implementation, the next punch-list item below. - [x] **Implemented (2026-08-26): the check function itself, real and complete — not yet wired to any write path.** `include/starkernel/homeblocks_sig.h` + `src/starkernel/homeblocks_sig.c`: `homeblocks_sig_check(dev, sig_start_fblock, out_sig)` reads the 4 consecutive 1KB `blkio` forth-blocks the 4KB header spans, verifies magic → version → CRC-64 in order, returns one of `HOMEBLOCKS_SIG_OK`/`_BLANK`/`_BAD_VERSION`/ `_BAD_CRC`/`_READ_ERROR`. Reuses `block_subsystem.c`'s existing CRC-64/ISO (`compute_crc64`, previously `static`/file-local, now exposed) rather than a second CRC implementation — same algorithm already proven via per-block checksums. Takes the header's starting block as a plain parameter rather than resolving it internally: this function verifies a signature given a location; finding that location (GPT-partition-relative, once a parser exists) stays the caller's job, not invented here. **Verified against the actual shipped code**, not a reimplementation: a standalone host test links the real `homeblocks_sig.c` against a fake in-memory `blkio_dev` and exercises all four outcomes — blank media → `BLANK`, a correctly-minted header → `OK` (round-trips `drive_uuid`/`minted_time_ns` correctly), a flipped CRC → `BAD_CRC`, an unrecognized version → `BAD_VERSION`. All four pass. A full QEMU-hotplug live test isn't proportionate yet — nothing calls this function from the live kernel path (deliberately; wiring it into the attach path is the next item below), so a live boot check has nothing to exercise. Clean zero-warning compile and clean boot on all three architectures confirms no build/link regression from exposing `compute_crc64` and adding the new source file to every kernel build. - [x] **Implemented (2026-08-26): the "warn" half, live and wired.** Wired into `sk_repl_idle()`'s USB hotplug attach handler (`repl.c`), right between `blkio_usb_open_msc()` succeeding and `blk_subsys_attach_device()` — calls `homeblocks_sig_check(&usb_blk_dev, 0, &sig)` and logs a distinct message per outcome (recognized / blank-or-foreign / bad-version / bad-crc / read-error). **The "refuse" half is deliberately not implemented — there is nothing real to gate yet.** `blkio_usb.c` has no SCSI `WRITE(10)` support at all (Milestone 2's biggest open item), so there is no write path today to refuse; the only thing attach currently enables is read-only access, which is also the general-purpose USB block I/O path this repo already relies on for unrelated testing, not exclusively a home-blocks identity workflow. Refusing attach on blank media would have broken that legitimate use without protecting anything real — building that gate now would be enforcement with no live consumer, the same "don't build ahead of a real caller" reasoning `EXPIRE`'s deferral used. Refuse belongs on the write path once `WRITE(10)` exists to give it something to gate against. `sig_start_fblock` is hardcoded to `0` at the call site — correct for today's unpartitioned raw test/real media (no GPT parser exists yet), explicitly flagged in the code comment as the one place that will need to change to a real GPT-partition-relative lookup once that parser lands, isolated from `homeblocks_sig.c`'s own location-agnostic check logic. **Verified live**, not just compiled: hot-attached `disk/usb-thumbdrive-test.img` (blank media, no `LAHB` magic) through a running amd64 instance's QMP socket (`blockdev-add` + `device_add usb-storage,bus=xhci0.0`) — captured exactly 4 real TUR+READ10 BOT cycles (matching the header's 4 forth-block span) followed by `xhci: USB drive not recognized (blank or foreign media) -- read-only general use only`, then normal attach completing successfully afterward (no regression — no `MSC block-subsystem attach failed`). Conservation intact, no panic. Clean zero-warning compile and clean boot on all three architectures. - [x] **Resolved (2026-08-26): reuse `acl_pinned` directly, no new flag needed.** The open question assumed credential data "isn't a dictionary word" — but the design already chosen for it (`ZUSE-CERT-LO`/`HI`, `ACL-CA-KEY-LO`/`HI`) are `CONSTANT` words, i.e. real `DictEntry`s. `acl_pinned`'s enforcement is more general than assumed: `vm_create_word()` (`dictionary_management.c:394-404`) — the single choke point every word-defining construct goes through — unconditionally refuses to let *anything* shadow a pinned name ("Pin is permanent: no word may shadow a pinned entry — ever, by anyone"), a real general redefinition guard, not just an ACL-mode-change lock. `zuse.4th`'s `ACL-ZUSE-BOOT` already pins both cert constants today — the mechanism is already wired for this. **Real gap surfaced along the way, not yet fixed:** `ACL-ZUSE-BOOT` self-activates and pins `ZUSE-CERT-LO`/`HI` unconditionally on *every* boot — before any legitimate minting step could ever run, permanently locking in the `0` placeholder on the very first boot. The boot sequence needs to distinguish "already minted, pin it" from "not yet minted, don't pin yet" before minting can work at all. **This directly shaped the next design pass (2026-08-26, Captain Bob):** a dedicated `disk/zuse.img` QEMU test thumbdrive, "bleachable" back to pristine/unminted state for repeated first-boot testing; a one-time first-boot mint-Zuse flow (mint → write real cert → blow the fuse → *then* pin, resolving the gap above); a separate, ongoing `S" name" MINT` word for an authenticated Zuse session to mint additional regular users; and a Zuse recovery path, explicitly flagged as unresolved and risky if rushed — not to be designed casually, since the earlier "no software recovery, mint a new one" rule existed specifically to close a hole a careless recovery path could reopen. **First piece implemented (2026-08-26): the `zuse.img` bleach mechanism.** `disk/zuse.img` (64MB, blank, matching the existing USB-fixture convention exactly — see `disk/README.md`) + `scripts/bleach_zuse_img.sh` (idempotent reset back to blank). Verified live: hot-attached via QMP (same method as the warn-and-refuse verification above) — reads back as `HOMEBLOCKS_SIG_BLANK` (`xhci: USB drive not recognized (blank or foreign media)`), correctly simulating a genuine first boot. Deliberately flat/raw, not GPT-partitioned, matching `homeblocks_sig_check()`'s current `sig_start_fblock=0` call site — both move to a real GPT-partition-relative offset together once a parser exists, not attempted here. No kernel code touched this step (host-side test tooling only), so no 3-arch acceptance boot needed — single live amd64 QMP-hotplug confirmation is the right verification tier. **Correction (Captain Bob, 2026-08-26): drives are not bound to any particular size.** 64MB was only ever this fixture's arbitrary test-convenience size, matching `usb-thumbdrive-test.img`'s existing BOT-driver-testing precedent — never a real-world constraint. Audited for anywhere this might have implied otherwise: the actual format (`homeblocks_sig_t`) already carries its own `metadata_devblocks` field and hardcodes no size anywhere, confirmed clean. `scripts/bleach_zuse_img.sh` gained a `--size-mb` override so this was never a hidden assumption baked into the tooling either. The earlier `project_usb_thumbdrive_gpt_layout.md` memory's "16GB reference size" phrasing (already hedged as tentative, but risked reading as a target) corrected to state the point explicitly — the GPT layout's design point is the *proportions* (small metadata partition, everything else block storage), not any absolute size. **Still open, not attempted:** the mint-then-pin boot-sequence fix itself, the `MINT` word, and the Zuse recovery path. **Correction, supersedes the "reuse `acl_pinned`" resolution above (2026-08-26): a pinned `CONSTANT` is not actually tamper-proof.** `ACL-PIN`/`acl_pinned` only guards against *redefinition* — `vm_create_word()`'s pin check blocks a second `: ZUSE-CERT-LO ... ;`, but nothing stops `' ZUSE-CERT-LO >BODY !` from overwriting the same word's data field in place. A `CONSTANT`'s value lives in its data field, so the earlier design left the cert mutable from FORTH despite being "pinned." Found while starting the mint-then-pin boot-sequence fix itself; fixing that gap came first since building a real mint flow on top of a tamperable store would just re-open the hole later. **Fixed (2026-08-26): moved cert storage out of the dictionary entirely.** New `VM` struct fields (`include/vm.h`): `zuse_cert_lo`/`zuse_cert_hi` (the cert value) + `zuse_cert_installed` (one-time fuse bit). New `vm_zuse_cert_install(vm, lo, hi)` (`src/vm.c`) — C-only, no FORTH word wraps it, returns `-1` on a second call rather than silently re-installing (a second call is a caller bug, not a runtime condition to recover from). No FORTH store word exists or should exist for these fields, closing the `>BODY` path structurally rather than by convention. Three new read-only C primitives (`src/word_source/starforth_words.c`, same shape as the existing `HEARTBEAT-TICKS@`): `ZUSE-CERT-LO@`, `ZUSE-CERT-HI@`, `ZUSE-CERT-INSTALLED?`. `capsules/zuse.4th`'s old `ZUSE-CERT-LO`/`HI` `CONSTANT` words (and `ACL-ZUSE-BOOT`'s two now-pointless `ACL-PIN` calls on them) deleted outright rather than left as dead/insecure scaffolding — `mkcapsule --lint` clean (31/31) after the edit. `vm_zuse_cert_install()` has no caller yet: the real mint flow still doesn't exist (Milestone 6 CA + the `MINT` word are both still open), and calling it with a placeholder value would just be a stub wearing the shape of a fix — so this stays an honest, complete slice (storage + read accessors) with the actual mint-then- pin sequence still explicitly open, not faked. **Verified:** hosted `make` build clean, zero warnings; clean boot to `ok>` on all three architectures (amd64/aarch64/riscv64), Stadium conservation intact (43691/21845/65536) on all three, no panics or guest errors. ACL is opt-in (`init.4th`'s `S" ACL.4th" EXEC` commented out by default) so the new words weren't exercised live from the REPL this pass — compile/lint/boot verification only. **MINT word design, picked up 2026-08-26.** Before scoping `MINT` itself, found a real conflict with an existing, deliberate decision: `include/starkernel/ed25519.h` is verify- only by design — "this kernel never signs or generates keys (no entropy source to do so safely anyway); signing happens in the host-side build tool" (`FABRIC-2.md`, Milestone 6, 2026-08-22). But the vision for `MINT` (§D below) is an *interactive*, on-device `S" name" MINT` word — an authenticated Zuse session signing a new user's cert live, at runtime. A kernel that structurally never signs can't do that as envisioned. Raised directly; **decided (Captain Bob, 2026-08-26): give the kernel a real signing capability** rather than reshape `MINT` around verify-only. This reopens the prior "no entropy source" constraint deliberately, not by accident. **Phase A — `virtio-rng`, done 2026-08-26.** Checked what entropy is actually available before choosing a design: `include/starkernel/vm_uuid.h` already found, for VM UUIDs, that amd64 has RDRAND and riscv64 has the Zkr extension, but QEMU's aarch64 CPU models (including `max`) expose neither RNDR nor any RNG property at all — confirmed directly against QEMU 10.2.1. That's why VM UUIDs use a deterministic PRNG uniformly instead of a per-arch split; that same choice is **not safe for Ed25519 keygen** — a seed drawn from a known value makes the private key predictable. Decided: add a `virtio-rng` device instead of a per-arch RDRAND/Zkr split with a weaker aarch64 fallback — QEMU supplies real host entropy identically on all three arches, closing the aarch64 gap directly (QEMU-only; real hardware at Milestone 8 needs a real per-arch RNG driver, a separate later problem). New `include/starkernel/virtio_rng.h` + `src/starkernel/virtio/virtio_rng.c`, transport plumbing (PCI capability walk, common-cfg feature negotiation, split virtqueue) mirroring the existing `virtio_blk.c` exactly — same device family, same quirks. Simpler shape than block: one virtqueue, one device-writable descriptor, no request header or status byte (the entropy device has none); `virtio_rng_get_bytes()` loops internally since the device may return fewer bytes than requested per round. `-object rng-random,id=rng0,filename=/dev/urandom` + `-device virtio-rng-pci` added to all three arches' QEMU invocations (`Makefile.starkernel`). Wired into boot (`kernel_main.c`, right after the existing `virtio_blk_find_artemis()` call site, same graceful-noop-on-absence precedent). **Verified live, not just compiled:** a temporary probe (written, run once, captured, reverted — per this project's standing probe convention) pulled 16 real bytes through the full request/notify/poll/used-ring round trip on all three architectures and printed them: amd64 `be9223909b86a8ccbfff705ccae2caa6`, aarch64 `861487df65a6d7a26b2c9c34f5ff96ee`, riscv64 `fec51d80e8c169a35aad9d908eab2f39` — three different values, confirming real entropy, not a stale or repeated buffer. Probe reverted; permanent code is just the driver + init call. A second, final 3-arch acceptance boot ran against that reverted code (not the probe build) to confirm the shipped state itself is clean. Clean zero-warning compile and clean boot to `ok>` on all three architectures, Stadium conservation intact (43691/21845/65536), no panics or guest errors on either pass. **Still open: Phase B (real Ed25519 keygen/signing, seeded from this entropy) and Phase C (the `MINT` word itself, cert format, and whether Zuse's own keypair needs to chain to the Milestone 6 offline root CA or is a self-sovereign instance-local root of trust).** **Phase B — real Ed25519 keygen/signing, done 2026-08-26.** Extended `include/starkernel/ed25519.h`/`src/starkernel/crypto/ed25519.c` (previously verify-only) with `ed25519_keygen(seed, pubkey_out)` and `ed25519_sign(seed, msg, msg_len, sig_out)`, per RFC 8032 §5.1.5/5.1.6, reusing every point-arithmetic primitive verify already had (`scalar_mult`, `point_compress`, the base-point constants) — no new curve code, only the seed-expansion/clamping and per-message nonce derivation verify never needed. Signing is deterministic (nonce derived from seed+message, not fresh randomness): only keygen ever touches entropy, via a caller-supplied seed (`virtio_rng_get_bytes()`, Phase A) — keygen itself still generates nothing and trusts the caller for randomness quality, matching this file's original design philosophy exactly. New `scalar_muladd()` (`scalar25519.c`/`.h`) for signing's `S = (k*a + r) mod L` step, the one piece of scalar arithmetic verify never needed (verify only ever reduced or compared, never multiplied scalars). Schoolbook 256×256-bit multiply into a `u128` wide accumulator with exactly one final carry-propagation pass — deliberately the same shape as `fe25519.c`'s existing field multiply, because that file's own history records a real bug from trying to carry mid-accumulation instead of in one final pass; structurally can't repeat that mistake this way. Reduces the result via the existing, already-proven `scalar_reduce512()` rather than writing new modular-reduction logic. **Verified against an independent implementation, not self-consistency** — this project's own standing lesson (two real, invisible-by-inspection bugs in the original from-scratch field arithmetic, an off-by-one-hex-digit hand-transcribed SHA-512 vector) means a passing self-check proves nothing on its own. Built a throwaway host test harness (compiled, run, discarded — the crypto files have no `__STARKERNEL__` gate, so they link as an ordinary Linux binary) against Python's `cryptography` library (OpenSSL-backed). Six trials — five random seed/message pairs (message lengths 1, 32, 255, 1000 bytes) plus the empty-message case — every one produced a byte-for-byte identical public key and signature to the independent implementation, not just a signature this codebase's own verify accepted. **Verified on-target too:** clean zero-warning compile of the crypto files on all three architectures, and a full 3-arch QEMU acceptance boot (amd64/aarch64/riscv64) — all clean to `ok>`, Stadium conservation intact, no panics or guest errors. Nothing calls `ed25519_keygen()`/`ed25519_sign()` from the live kernel path yet (Phase C's job); this pass is compile/link/boot-regression verification for the crypto library itself. **Still open: Phase C** — the `MINT` word itself, cert format, and whether Zuse's own keypair needs to chain to the Milestone 6 offline root CA or is a self-sovereign instance-local root of trust. **Phase C scoping, 2026-08-26.** Three findings before any code: 1. **Resolved, not a real conflict: Zuse doesn't need Milestone 6's CA.** That CA chain is specifically for *capsule/code signing* (root → snakeoil intermediate → per-capsule Ed25519 signatures verified at capsule-load time) — a different trust domain from *user identity*. The vision's own framing ("we mint one and only one Zuse user and blow a fuse ... the only way around is a new system") already implies Zuse's authority comes from being the unique first-boot mint on *this instance*, not from an external chain. **Decided: Zuse is a self-sovereign, instance-local root of trust**, keypair generated on-device from real entropy (Phase A+B). Regular users, minted later via `MINT`, get certs signed by *Zuse's* key, not the Milestone 6 CA — two independent PKI domains. 2. **A real gap in this session's own earlier work:** `vm_zuse_cert_install(vm, lo, hi)` (the very first change this session made, before Phase A existed) only holds two `uint64_t` (16 bytes) — sized against the old placeholder `ZUSE-CERT-LO`/`HI` FORTH-cell design, not against what a real Ed25519 keypair needs (32-byte pubkey alone, well over 100 bytes for a full cert). Needs expanding before Phase C can store anything real. 3. **A genuine blocker, found by asking where the cert would actually live:** "mint once, ever" requires surviving reboots, but `Makefile.starkernel`'s `qemu` target copied a fresh, pristine `OVMF_VARS.fd` on *every* invocation (not just after `clean`) — so a UEFI-NVRAM-based cert (the real-hardware-compatible option, and this codebase already has a live `SetVariable`/`GetVariable` precedent via `SF_VAR_REBOOT_TRIES`/`SF_VAR_BOOT_ARGS`) would never actually persist under this project's own normal test workflow. Digging further: aarch64's `qemu` recipe had no persistent NVRAM store *at all* — a single combined `-bios $AAVMF_CODE` argument, no separate writable VARS pflash drive like amd64/riscv64 have. **Decided (on request): fix the harness rather than switch substrates.** amd64/riscv64: the VARS-template copy is now conditional on the destination not already existing, so `clean` (which deletes the whole `build/$(ARCH)/kernel` tree, `OVMF_VARS.fd`/`RISCV_VARS.fd` included) is the bleach step, and a bare `make qemu` now preserves NVRAM across runs — exactly matching the existing "always pass `clean` before `qemu`" acceptance convention, no new bleach script needed. aarch64: restructured to split CODE(ro)/VARS(rw) pflash drives matching the other two (host has `/usr/share/AAVMF/AAVMF_VARS.fd` alongside the existing `AAVMF_CODE.fd`), with a graceful fallback to the old single-`-bios` mode (and a console note) on a host that only has non-split firmware packaged, so this doesn't regress environments without one. **Verified live:** all three architectures still boot clean to `ok>` with the new pflash arrangement, Stadium conservation intact, no panics or guest errors — this is infrastructure-only (no cert code yet), so a plain boot-regression check is the right verification tier. **Still open:** the actual cert struct (expanding past the 16-byte placeholder), the first-boot mint-vs-already-minted boot sequence using `SetVariable`/`GetVariable`, and the `MINT` word itself. **Cert struct expanded (2026-08-26):** `vm_zuse_cert_install()` (both `src/vm.c`'s hosted copy and a new kernel-side duplicate in `src/starkernel/vm/vm_core.c` -- the kernel build's `VM_EXCLUDE` list drops `src/vm.c` entirely, same reason `vm_set_base()` already has two independent copies) now takes a real 32-byte seed + 32-byte pubkey instead of the old 16-byte placeholder. FORTH-side `ZUSE-CERT-LO@`/`HI@` replaced with `ZUSE-PUBKEY@ ( i -- u )` (8-byte LE chunk `i`, 0..3, of the public half only -- the seed has no FORTH access at all). `ACL-ZUSE-BOOT` now checks `ZUSE-CERT-INSTALLED?` before authenticating rather than authenticating unconditionally. Verified: clean compile and clean boot on all three architectures. **NVRAM persistence attempt: crashed, root-caused, reverted -- do not retry as designed.** First attempt placed the mint-or-load `GetVariable`/`SetVariable` logic right after `virtio_rng_init()` (before `capsule_birth_mama()`); it page-faulted (`CR2` inside the OVMF flash MMIO window, a supervisor write to a not-present page) partway through boot. Moved the same logic to the one place in this codebase already calling `SetVariable` post- `ExitBootServices` successfully (`SF_VAR_REBOOT_TRIES`, much later in boot) — **identical crash, same RIP and CR2** — which disproved the "too early in boot" theory outright: it isn't a timing issue. **Localized precisely (advisor-directed, one boot, debug markers around each call):** `GetVariable` returns fine. `SetVariable` **with real 64-byte data** never returns — that's the exact fault site. The pre-existing `SF_VAR_REBOOT_TRIES` call that looked like a working precedent is actually a **delete of a variable that's never existed** (`size=0, data=NULL`) — a fundamentally different, much cheaper internal path than a real data write, so it proved nothing about real persistence being safe. **Root cause: this kernel's VMM never maps whatever memory region OVMF's variable service needs to actually write flash-backed variable data** — a real gap in UEFI runtime-services support, not specific to Zuse. Fixing it for real means walking the UEFI memory map for the relevant regions and mapping them into the kernel's own page tables, and per Section U's own note, the flash window's location is firmware/arch-specific (OVMF's differs from AAVMF's and EDK2-riscv64's), so "walk the map and map everything" is not guaranteed 3-arch-uniform even once attempted. **Second, independent finding (not a bug, a design flaw in the persistence choice): storing the raw 32-byte seed in NVRAM was a real defect regardless of the crash.** `SetVariable` was called with `EFI_VARIABLE_RUNTIME_ACCESS`, meaning any later-loaded UEFI application or the booted OS itself could read Zuse's private key straight out of NVRAM. For an irrevocable "one and only one Zuse, ever" root of trust, that undermines the property the design exists to provide — this would have needed fixing even had the crash not happened. **Decision needed, not yet made:** given virtio-blk writes are already proven working on all three architectures in this repo (`vblk_write`, Artemis's own persistence across runs), a dedicated file-backed system-identity disk (mirroring `disk/artemis.img`'s existing pattern, separate from Artemis's internal storage and separate from home-blocks USB thumbdrives) is the substrate with no open unknowns today — recommended over either fixing the UEFI flash-mapping gap (real but large, unscoped VMM work) or accepting the NVRAM approach as originally designed (has the exposed-seed defect regardless). Not decided or built yet. **Reverted to a known-safe state:** all Zuse mint/NVRAM code removed from `kernel_main.c` (only two harmless includes remain), `init.4th`'s `ACL.4th` line back to its documented commented-out default. Verified clean compile and clean boot on all three architectures in this reverted state. **Substrate corrected (Captain Bob, 2026-08-26): no files, ever — this OS's entire reason for being is anti-POSIX, anti-file.** The "dedicated system-identity disk" recommendation above was framed in file/filesystem language by mistake; corrected before any code was written. The only real persistence primitives here are content-addressed capsules and raw LBN-numbered blocks (`block_subsystem.c`) — never a filesystem, never file paths. Saved as `feedback_no_files_anti_posix.md` so this isn't re-learned next session. **Design, agreed on request: a growable metadata fence at the TOP of a device's block space, mirroring the bottom BAM reservation from the opposite end.** `block_subsystem.c`'s BAM already reserves the bottom `BLK_DISK_SYS_RESERVED` (32) blocks of every attached device, invisible to FORTH's `BLOCK`/`BUFFER`. Zuse's cert (and future system metadata) gets a second reservation at the *top* of the same device, starting at `BLK_META_FENCE_INIT` (128) blocks and growing downward as needed — the two reservations grow toward each other from opposite ends, never colliding, same shape as a stack/heap. Explicitly never RAM-backed (the fast-RAM/ramdrive LBN ranges are documented as volatile in this same file's own header comment — losing Zuse's identity to a RAM eviction is exactly the failure this is designed against). Reuses Artemis's own already-attached, already-proven virtio-blk device — no new device attachment. Rejected reusing BAM's own bottom-reserved zone directly: those 32 blocks are fully claimed by BAM/volume-metadata bookkeeping, not free space. **Step 1 (field round-trip) implemented and verified 2026-08-26, allocator not yet touched.** New `meta_fence_blocks` field in `blk_volume_meta_t`, appended after `reloc_devblocks` and carved from `_pad[]` — identical graceful-default technique the `reloc_start`/`reloc_devblocks` fields already established (a pre-existing formatted volume reads the field back as 0 via its zeroed former padding, not a format-breaking change). Added a compile-time `_Static_assert(sizeof(blk_volume_meta_t) == 4096, ...)`, same discipline `homeblocks_sig.h` already uses — caught a real bug immediately: the hand-summed `_pad[]` size formula was off by 4 bytes (a compiler-inserted alignment gap before `tracked_blocks` that the manual byte-count missed), found via `offsetof()` rather than by re-deriving the arithmetic by hand again, consistent with this project's standing rule to never trust a hand-derived numeric claim in this class of code. Worked against disposable clones throughout, never the real `disk/artemis.img` (`ARTDISK=...` is `?=`-overridable) — `disk/artemis-metafence-fresh.img` (blank, exercises the fresh-format path) and `disk/artemis-metafence-test.img` (a copy of the pre-existing `artemis.img`, exercises the graceful-default-on-reload path), both kept as regression fixtures per `disk/README.md`'s existing convention (mirrors `artemis-reloc-test.img` exactly). **Verified independently via direct byte reads of the disk image, not the kernel's own self-report** (`log_message(LOG_INFO, ...)` turned out not to reach serial output at all in this build — an unrelated, pre-existing log-level gap, not a regression): fresh format writes `meta_fence_blocks=128` at header byte offset 184; a second boot without reformatting reads it back unchanged; the pre-existing old-format image correctly reads back 0. Full 3-arch acceptance boot against the real, untouched `disk/artemis.img` also clean — conservation intact, no panics. **Step 2 (allocator + read/write accessors), done 2026-08-26.** Units corrected from "Forth 1 KiB blocks" to 4 KiB devblocks (matching `bam_devblocks`/`reloc_devblocks`) before anything depended on the original meaning — a clean fix, not a migration, since nothing consumed the field yet. This let the fence fold directly into `compute_totals_from_B()`'s existing `payload4k` calculation (`total_devblocks - 1 - B - R - F`, F = `meta_fence_blocks`) instead of needing a second, separate subtraction against `user_blocks` — `total_blocks`, `user_blocks`, and `free_blocks` all shrink correctly for free, in both the fresh-format and reload code paths, from this one formula change. New `blk_meta_zone_read()`/`blk_meta_zone_write()` (`block_subsystem.c`/`.h`) — raw, unpacked 4 KiB devblock I/O (no Forth-block packing, same shape as the header/BAM/reloc-table regions), addressed by `devblock_from_top` counting down from the device's last physical devblock, refusing (not silently clamping) if the index isn't within the on-disk `meta_fence_blocks`. No FORTH word wraps either — C-only, same discipline as `vm_zuse_cert_install()` itself, which will be this zone's first real tenant. **Verified independently at every step, never trusting the kernel's own report:** - Capacity math: read a freshly-formatted image's header bytes directly and independently recomputed the expected `total_blocks` in a separate Python script using the same formula — exact match (22647, down from what it would have been without the fence). - Accessor correctness: a temporary probe (written, run, captured, reverted) wrote a known 256-byte-repeating pattern via `blk_meta_zone_write(0, ...)`, read it back via `blk_meta_zone_read(0, ...)`, and compared in-memory (`PASS`) — then, independently, read the raw image file at the exact expected physical byte offset (`(total_devblocks-1)*4096`) and confirmed the pattern landed there byte-for-byte. - `log_message(LOG_INFO, ...)` still doesn't reach serial output in this build (same pre-existing gap noted in Step 1) — all verification here used `console_println` (which does reach serial) for the temporary probe, and direct file reads for everything else. Full 3-arch acceptance boot (real, untouched `disk/artemis.img`, probe code fully reverted) clean on all three architectures — conservation intact, no panics. **Still open:** wiring `vm_zuse_cert_install()`'s seed+pubkey to actually persist through these new accessors (the zone exists and works; nothing writes Zuse's cert into it yet), and the `MINT` word itself. **Step 3 (Zuse's cert wired to the fence), done 2026-08-26 -- first-boot mint-then-load is real, end to end.** New `include/starkernel/zuse_cert_devblock.h`: a small, standalone on-disk record format (`zuse_cert_devblock_t` -- magic + version + 32-byte seed + 32-byte pubkey + a real CRC-64/ISO from day one, same "real from day one" discipline `homeblocks_sig_t` already established, since this gates a real security check) occupying devblock_from_top=0 of the fence. Deliberately its own header, not inlined at the boot-time call site: the still-open `MINT` word will be a second consumer of this exact format later. `kernel_main.c`'s mint-or-load logic moved from the crashed NVRAM approach to this: read devblock 0 of the fence, and if magic/version/CRC all check out, install the existing cert; otherwise, if `virtio_rng` is ready, mint a fresh one (Phase A+B) and write it. Runs right after `virtio_rng_init()`, well before `capsule_birth_mama()` -- unlike the crashed NVRAM attempt, raw block I/O against Artemis's already-proven virtio-blk device has no boot-timing risk at all, so the earlier "re-invoke `ACL-ZUSE-BOOT` after Mama birth" workaround is no longer needed; `ACL.4th`/`zuse.4th`'s self-activating `ACL-ZUSE-BOOT` sees a correctly-populated cert on its one, ordinary first pass. **Verified live, independently, across every real scenario, never trusting the kernel's own report:** - **Fresh mint** (blank `disk/artemis-metafence-fresh.img`): boot logs `Zuse: minted, fuse blown`; the on-disk record at the exact expected physical offset independently decodes to magic bytes `b'ZUSE'`, version 1, a real 32-byte seed and pubkey, and a CRC that an independent from-scratch Python re-implementation of the exact CRC-64/ISO algorithm (table generation included, not just the check) confirms byte-for-byte. - **Reload** (reboot the same now-minted image, no reformat): boot logs `Zuse: cert loaded from block fence`; the on-disk seed and pubkey are byte-for-byte identical to the first boot's -- genuinely "mint once, ever," not a silent re-mint. - **Graceful refusal on a pre-fence volume** (`disk/artemis-metafence-test.img`, `meta_fence_blocks=0`): `blk_meta_zone_read`/`write` both correctly refuse (no space to read or write), so the kernel mints a cert for RAM/this-boot-only use and honestly reports `Zuse: minted but fence write FAILED (not persistent)` -- no crash, no silent data loss, no corruption of a device with no fence at all. - **Real disk regression check:** the same graceful-refusal path exercised identically against the real, untouched `disk/artemis.img` (which has no fence yet either) on all three architectures -- clean boot, conservation intact, no panics, `disk/artemis.img` itself reverted afterward (no committed churn). **Phase 8's core arc is now functionally complete:** real entropy (Phase A) → real signing (Phase B) → real, anti-file, block-native persistence (Phase C) → a working first-boot mint that survives reboots. **Still open:** the ongoing `S" name" MINT` word for an authenticated Zuse session to mint additional regular users (needs `zuse_cert_devblock_t`-format certs signed by Zuse's own key, not just installed) -- the real remaining piece of the original vision. ### From FABRIC-2.md §X, Milestone 5 — Console/VM key-match binding - [ ] Settle the still-open question: reuse `ACL-PIN`/`acl_allow` directly, or build a separate key-matching primitive — `ACL-PIN` gates word execution specifically and nothing today gates console-session-to-VM ownership, so this decision needs to happen before any code gets written here. - [ ] Design the key/lock data shape (what the console presents, what the VM carries, how they're compared). - [ ] Wire drive insertion (Milestone 2e's hotplug signal, post-identity-authentication) to a call into `capsule_birth_baby()` (confirmed a real, callable, on-demand birth path already) to spin up or re-attach that identity's VM. - [ ] Implement the actual attach/bind step — extending `sk_repl_set_active_vm()` (confirmed to exist, currently an unguarded raw pointer-set) with the key-match check from above, so a console can only bind to the one VM whose lock matches its key. - [ ] Implement detach behavior on console disconnect or VM teardown. ### From FABRIC-2.md §X, Milestone 6 — Kernel/capsule PKI signing chain - [ ] Generate (offline, outside the kernel/repo entirely) the real root CA keypair — "stays unrevocable," never embedded, never loaded by any kernel code. - [ ] Generate the "snakeoil" intermediate certificate, signed by that real root CA (this is a real CA-signed intermediate, not a self-signed/untrusted cert despite the name — "snakeoil" names its informal/private-project status). - [ ] Embed the already-CA-signed snakeoil intermediate as a capsule blob at build time (mechanically proven already via the font-capsule precedent — no new embedding infrastructure needed, just a new payload). **Bootstrapping resolved: no kernel-boot-time verification of a hardcoded CA public key is needed at all** — trust is established once, at build time, by whoever holds the real root CA and produces the build. - [ ] Add a signing step to the `mkcapsule` build tool (or a separate signing tool) that produces a signature alongside each capsule's existing xxHash64. - [ ] Extend `MANIFEST_AUTO.md`'s generation to add a signature-status column, matching the existing xxHash64 column's generation pattern. - [ ] Implement magic-number-based content-type detection (Section U item 14) — a shared primitive, also usable for Milestone 4's foreign-drive check. - [x] **Root CA + snakeoil intermediate generated 2026-08-26**, entirely offline, in a sibling directory outside this repo (`/home/rajames/CLionProjects/lithosananke-ca/`, not tracked by git here). Ed25519, OpenSSL 3.5.5. Root: self-signed, 20-year validity (2026–2046), `CN=LithosAnanke Root CA`. Intermediate: a real CA-signed cert (not self-signed despite the name), 10-year validity, `CA:TRUE, pathlen:0` (can sign capsules, can't mint further intermediates), chain verified (`openssl verify` returns `OK`). Both private keys `chmod 600`. The root key never touches this repo or any kernel code, per the design's own requirement. - [x] **Snakeoil intermediate embedded as a capsule, 2026-08-26.** Exported to DER (`capsules/pki/snakeoil-intermediate.der`, 418 bytes) and dropped under `capsules/` — confirmed the font-capsule precedent needed zero new infrastructure: `mkcapsule`'s `process_file()` embeds any non-`.4th` file verbatim already. Shows up as capsule `pki:snakeoil-intermediate.der` in the generated capsule table (38 capsules total, up from 37) — retrieve via `capsule_find_by_name()` + `capsule_get_payload()`, never `capsule_exec_payload()` (it's a passive data blob, not executable capsule code). - [x] **Minimal DER/X.509 parser written and independently verified, 2026-08-26.** New `include/starkernel/x509_ed25519.h` + `src/starkernel/crypto/x509_ed25519.c`: `x509_extract_ed25519_pubkey()`, a from-scratch, narrow DER walker (not a general ASN.1/ X.509 parser, per this milestone's own design decision) — walks `Certificate → TBSCertificate → SubjectPublicKeyInfo`, handles the optional `[0] EXPLICIT Version` field (present on v3 certs), verifies the `AlgorithmIdentifier` OID is exactly `1.3.101.112` (RFC 8410 Ed25519) rather than assuming, and extracts the raw 32-byte key from the trailing `BIT STRING`. Handles both short-form and long-form DER lengths (a real cert with v3 extensions routinely exceeds the 127-byte short-form limit). Every step bounds-checked against the buffer end — refuses malformed input, never faults. **Verified against ground truth, not self-consistency:** run against the real embedded `snakeoil-intermediate.der`, the extracted 32-byte key matched `openssl pkey -pubin -text`'s own reported public key byte-for-byte. Refusal path verified too: truncated input, 10 random garbage bytes, an empty file, and a real RSA certificate (algorithm-mismatch case, not just structural malformation) all correctly return failure rather than misreading or crashing. Compiles clean on all three kernel architectures (no `__STARKERNEL__` guard needed — same freestanding-safe shape as the other crypto files). **Still open:** the signing step in `mkcapsule` (needs a new `sig[64]` field on `CapsuleEntry` and a parallel emitted array in `capsule_generated.c`, since `CapsuleDesc` itself has no spare bytes — confirmed exactly 64 bytes, every field used), wiring `ed25519_verify()` into the three `capsule_validate()` call sites in `capsule_birth.c` (**decided: land as WARN-only first, prove correct on all three architectures against both a valid and a deliberately-corrupted capsule, then flip to hard-refuse in a separate step** — a bug here has a much larger blast radius than anything else in Phase 8, since a false refusal on Mama's own capsule means no `ok>` at all, on any architecture), and the signature-status column on `capsules/BLOCK_MAP.md` (confirmed the real, live manifest target — `capsules/MANIFEST_AUTO.md` is stale/dead, not regenerated since 2026-07-05, flag as docs drift rather than a real target). ### From FABRIC-2.md §X, Milestone 7 — Contributor capsules / trust tiers - [ ] Create the `capsules/contrib/` directory (mechanically trivial, matches existing subdirectory convention — the directory itself is not the work). - [ ] Add a `FLAG_CONTRIB` bit to `mkcapsule.c`'s flag system, assigned by path match (`contrib/` prefix), same pattern as how `init.4th` already gets `FLAG_MAMA_INIT`. - [ ] Decide and implement one of the four spitballed trust-tier directions (signature- authority tiers / block-namespace sandboxing / QEMU-vs-real-hardware conditional enforcement) — none chosen yet, this is a real decision point, not just an implementation task. - [ ] If block-namespace sandboxing is chosen: extend `mkcapsule`'s existing conflict- detection logic to also reject a `contrib/`-path capsule claiming blocks outside its reserved range. ### From FABRIC-2.md §X, Milestone 8 — Bare-metal boot from physical USB - [ ] Build a fresh `starkernel.iso` via `make -f Makefile.starkernel ARCH=amd64 clean` + the ISO-build step. - [ ] Identify the exact block device path for the target USB drive on the host doing the flashing (`lsblk`/`dmesg` after insertion — care needed, wrong device = data loss). - [ ] `dd if=build/amd64/kernel/starkernel.iso of=/dev/sdX bs=4M status=progress` (or equivalent) — confirm `dd` is the right tool for an El Torito ISO vs. needing `isohybrid` first (open question, not yet verified). - [ ] Physically boot the real machine from the flashed drive (BIOS/UEFI boot-order menu, Secure Boot may need disabling — unknown until tried). - [ ] Capture what happens with no serial-socket log available (real hardware has no `qemu-serial-*.sock` to `socat` into) — decide the observation method. - [ ] Confirm POST reaches the same 1012/0/0 result on real hardware as every QEMU acceptance run. - [ ] Confirm `ok>` prompt is reachable and a basic command (e.g. `HEARTBEAT-TICKS@ .`) works identically to QEMU. - [ ] Document the result (pass/fail, and if fail, what diverged from QEMU) — first real external validation this project has ever had outside QEMU TCG emulation. ### From FABRIC-2.md §X, Milestone 9 — Networking / capsule distribution server - [ ] (Deferred) Revisit and punch-list this milestone once Milestone 7 closes, not before. --- ## B. Stadium unification — words/VMs/blocks/messages on the same engine Raised 2026-08-25: "words are stadium patrons, VMs are patrons, blocks are patrons, messages are stadium patrons, all should be operated on by THE SAME ENGINE." Investigated before designing anything — the real state is more nuanced than "everything's a stub," verified via direct reads and `git log`, not assumed: **`FABRIC.md` §18.3 already decided the mapping** (not invented here): blocks → `MIGRATE`, messages → `DELIVER`, ACLs → `EXPIRE`, words and VMs both → `COOL`. `stadium_evict()` (`src/starkernel/vm/stadium.c`) — real, tested infrastructure: bitmap tracking, pin/`contains` refusal, the Hera-patron-zero panic guard, heat-conservation back to the owner's reservoir on every reap — calls `stadium_dispatch()` for the actual payload action when a patron departs. **Per-behaviour status, as of 2026-08-25:** - **`MIGRATE` (blocks)** — zero consumer, genuinely stub (`stadium_dispatch()`'s case prints `"MIGRATE (stub)"` and returns). This session already built the real mechanical primitive it needs: `blk_subsys_relocate_block()`/`RELOCATE-BLOCK` (`FABRIC-2.md`, commit `36d832f`), live-verified (redirect + content survive an abrupt kill and reboot) but never wired to `stadium_dispatch()` — it's a separate, parallel, already-working mechanism today, not routed through Stadium at all. - **`COOL` (words *and* VMs, same tag)** — half real. **Words are fully live**, but via a *separate, bespoke* mechanism, `stadium_word_dispatch()` (`stadium_words.c`, item 4.1), wired directly into the real VM word-execution hot path (`vm_core.c:690,885,896`) — it does **not** go through the generic `stadium_dispatch()` switch at all. `ONTOLOGY.md` §IX claiming words are "not yet migrated" is itself stale documentation drift (same class of bug as the "glibc" misattribution corrected earlier this session — flagged as a small, separate fix below, not blocking). **VM cooling has no evidence of ever being wired anywhere** — still genuinely stub. - **`DELIVER` (Hermes messages)** — `FABRIC.md` (~line 3452) records this explicitly as "Open, surfaced not resolved": Hermes's message/channel heat already integrates with Stadium's reservoir accounting (`STADIUM-HEAT@`, `STADIUM-RES-PULL/PUSH`), but which Hermes lifecycle event maps to `DELIVER` vs. `EXPIRE` was never decided, let alone wired. Real, substantial, Hermes-specific integration work. - **`EXPIRE` (ACL TTL expiry)** — no evidence of any wiring anywhere; `ACL.4th`/ `acl_recheck()` has zero Stadium involvement today. Also substantial, separate work. **Why `DELIVER`/`EXPIRE` aren't being resolved in the same pass as `MIGRATE`:** each is a full subsystem integration (Hermes lifecycle mapping; ACL-to-Stadium wiring where none has ever existed) in its own right — attempting all four stubs at once risks exactly the rushed, shipped-but-incomplete outcome the no-stubs rule (below) exists to prevent. `MIGRATE` gets resolved for real because this session already has a tested primitive underneath it; the other three become honest, explicit punch-list items instead of being touched speculatively. **Punch list:** - [x] **Investigated (2026-08-25): block-patron admission does not exist yet, but is architecturally straightforward, not blocked.** Confirmed via `stadium_admit()`'s own doc (`stadium.h`) that it REFUSES any candidate with `mass != 1`, and that a `mass > 1` multi-cell patron would need a continuation chain nothing has ever designed — this looked at first like a hard blocker for a 1024-byte block. It isn't: confirmed via `stadium_word_dispatch()`'s real candidate construction (`stadium_words.c:245-252`) that Stadium cells carry pure identity/heat/bookkeeping only — `candidate.identity = word_id`, `payload[32]` unused — the actual word content stays in the dictionary; Stadium never holds it. By the same pattern, a block patron's cell would carry `identity = LBN`, `mass = 1`, `payload` unused — the actual 1024 bytes of block content stays exactly where it already lives (block cache / disk via `block_subsystem.c`), unmoved. So `mass != 1` is a non-issue; the real gap is just that nothing has ever built the LBN→cell_index residency map (the block-patron analogue of `stadium_word_dispatch()`'s `resolve_resident_cell()`) or the touch-on-access hook (the analogue of `vm_core.c`'s three `stadium_word_dispatch()` call sites). Not yet built — this is real, scoped, buildable work, not a stub-around candidate. **Still open**, plan to be presented before implementation per the no-stubs/no-early-coding conventions. - [x] **Resolved (2026-08-25): real block-patron admission + real `MIGRATE` dispatch, both live.** New `stadium_blocks.h`/`stadium_blocks.c` mirror `stadium_words.c`'s shape (Option B starter-grant admission, redirected Loop #3 cooling, self-healing stale-entry detection) but key residency by `(quota_slot, lbn)` in a fixed-capacity open-addressing hash table sized off `stadium_cell_count()` (tombstone-based deletion, since LBN space isn't densely bounded like `word_id`), not a dense array. Wired into `block_word_block()`/`block_word_buffer()`/ `block_word_update()` (`block_words.c`), `#ifdef __STARKERNEL__`-guarded. `stadium_dispatch()`'s `MIGRATE` case now calls `blk_flush(lbn)` for real (confirmed `blk_flush()`, not `blk_subsys_relocate_block()`, is the right primitive — the latter is for compudynamics-driven relocation to a *different* LBN mid-residency, not ordinary reap write-back). Three new Kconfig tuning constants (`STADIUM_BLOCK_HEAT_QUANTUM`/`STADIUM_BLOCK_COOL_RATE_Q48`/ `STADIUM_BLOCK_TRACK_CAP_MULT`) mirror the word-patron ones exactly, same three-layer wiring. **Verified:** clean compile, zero warnings, on all three architectures; clean boot to `zuse)ok>`/`ok>` REPL on all three, conservation (`resident_sum + reservoir == Q48_ONE`) intact identically across all three; `BLOCK`/`BUFFER` touches exercised live from the REPL on amd64 and riscv64 with no crash; a 22,000-distinct-block flood loop (amd64, artificially shrunk to a 20,971-cell Stadium via a one-off smaller `-m` to make quota pressure reachable) ran clean under heavy admission-path load with no corruption. **Live `MIGRATE` fire confirmed (2026-08-26):** interactive flooding alone never triggered it — Hera's reservoir was already sitting exactly at the `Q48_ONE / 3` floor from the boot-time self-tests, so every block-touch candidate pulled 0 heat, and a 0-heat candidate can never be *strictly denser* than an existing resident, so `stadium_admit()`'s eviction fallback correctly refuses rather than evicts once the free list is exhausted (a pre-existing reservoir-floor/density-eviction interaction, applies equally to word patrons, not introduced by this pass). Closed deterministically instead with a temporary boot probe (`kernel_main.c`, inserted into the existing Artemis 4.6 self-test block, reverted immediately after capture — no code left behind): `100 65536 0 STADIUM-ADMIT ... STADIUM-EVICT` against the live Artemis VM. Captured live: `Stadium: dispatch cell=63257 behaviour=MIGRATE lbn=100`, followed by `DBG err after MIGRATE probe=0` — `blk_flush(100)` fired for real, with the correct LBN threaded through from the departing patron's `identity` field exactly as designed. (Artemis's own reservoir went briefly out of Q48_ONE-balance during the probe — raw `STADIUM-ADMIT` doesn't debit the reservoir on its own, by its own doc, so a heat value handed to it directly is invented, not pulled; harmless here since Artemis is killed and her whole economy discarded immediately after, and Hera's own conservation was independently confirmed back at the normal 43691/21845/65536 baseline afterward.) - [x] **Bug found, reported, then fixed on request (2026-08-25/26): `capsules/lib.4th:13-14` shadowed the C primitives `USE`/`RUN`.** While chasing the live-MIGRATE test above, `S" Artemis" USE` (meant to redirect the REPL into Artemis's own vocabulary, `mama_word_use()`) instead printed `EXEC: failed: Artemis`. Root cause: `capsules/lib.4th:13` defined `: USE ( addr u -- ) EXEC ;` — a FORTH word with the same name but a completely different meaning ("load/exec a capsule"), which shadowed the C-registered `USE` in dictionary search order since `lib.4th` loads after primitive registration. `lib.4th:14` did the identical thing to `RUN`, which CLAUDE.md also names as an untouchable C primitive ("BIRTH, RUN, USE are primitives — registered in C exactly like DUP, BYE, EXEC"). Same bug class as the K-PUSH dictionary-shadowing issue (`docs/working/architecture/K-PUSH-DICTIONARY-SHADOWING-BUG-20260704.md`). Reported first per CLAUDE.md's rule against unrequested fixes; Captain Bob then explicitly asked for the fix. **Fix:** traced every real caller before touching anything — `RUN`'s alias was dead (never called anywhere as bare `RUN`); `USE`'s alias had exactly one real caller, `capsules/hermes/init.4th:397` (`S" common:msg.4th" USE`, intentionally exploiting the shadow to load that capsule right after `lib.4th` itself loaded). Both aliases were pure `EXEC` wrappers with zero added behavior, so the fix deleted both definitions from `lib.4th` outright and changed the one real call site plus its matching doc comment (`capsules/common/msg.4th:4`) to call `EXEC` directly — no new names invented, the C primitives untouched, `mkcapsule --lint` clean (31/31 pass). **Verified live:** rebuilt and booted amd64 — Hermes still births and her `COMMON-CH`-eviction self-test (which depends on `common:msg.4th` having loaded) still passes exactly as before; interactively, `S" Artemis" USE` now correctly prints `USE: now using Artemis` and switches the REPL's console-name coloring, confirming the C primitive runs unshadowed. Clean compile and clean boot with conservation intact on all three architectures (amd64/aarch64/riscv64). - [x] **Resolved (2026-08-26): real VM-patron admission + real explicit-KILL eviction, both live.** Re-scoped on request: confirmed `capsule_vm_kill()` had zero Stadium involvement (`vm_cleanup()`/`sf_free()` only) and child-VM birth only ever called `stadium_grant_quota()` (a resource pool for the VM's *own* future word/block patrons) — never `stadium_admit()` for the VM *itself*. The only precedent, `stadium_birth_hera()`, admits Hera into her own quota as a permanently pinned cell 0, which can never reach `stadium_evict()` — not a working example of `COOL` firing for a VM. On closer look this turned out NOT to be entangled with the still-iterating Tripod/Zuse/messaging vision after all (§D) — birth and kill already funnel through two single choke points, so the earlier 2026-08-25 deferral was overcautious. **Design:** added `size_t stadium_patron_cell` to `VMRegistryEntry` (`capsule_run.h`). At birth, right after the existing `stadium_grant_quota()` call (`capsule_birth.c`), admit a candidate into the new VM's own quota mirroring `stadium_birth_hera()`'s shape (`identity=0`, `mass=1`, `behaviour=COOL`) but deliberately **unpinned** — pinning would need a new "unpin" primitive (none exists) to ever evict it later, and adding a pin-bypass to `stadium_evict()`'s refusal logic isn't something to do casually; unpinned costs nothing since nothing wires `COOL`'s dispatch body to actually kill anything, so the worst case of an unrelated natural eviction is `stadium_patron_cell` going stale, which is tolerated the same way `stadium_word_forget()` already tolerates staleness. At `capsule_vm_kill()` and `capsule_vm_kill_all_nonmama()`: `stadium_evict()` the tracked cell if still resident, silently tolerating refusal (already gone). `stadium_dispatch()`'s `COOL` case needed no new payload body — same as it already is for words, where `COOL` has no defined extra action beyond `stadium_evict()`'s own universal reservoir credit; the missing piece was admission and a genuine trigger, not dispatch-body logic. **Verified live:** a second, new `Stadium: dispatch cell=... behaviour=COOL` now fires immediately before every `PARITY:KILL` line, for both Hermes and Artemis, confirmed on amd64 (distinct from the pre-existing `COMMON-CH` word-eviction self-test's own COOL print). Conservation (`resident_sum + reservoir == Q48_ONE`) intact throughout. Clean zero-warning compile and clean boot on all three architectures (amd64/aarch64/riscv64). - [x] **Re-scoped and resolved (2026-08-26): `DELIVER` (Hermes) was never actually a gap — the earlier "zero consumer, needs substantial Hermes lifecycle mapping" framing above was wrong, carried over unverified from `FABRIC.md`'s old "open, not resolved" note about *which* Hermes event maps to `DELIVER` vs. `EXPIRE`. Item 4.2 already answered that in code (messages → `DELIVER`, channels → `COOL`) without the prose ever catching up — same documentation-drift class as the stale `ONTOLOGY.md` words note and the earlier glibc misattribution. Confirmed live: `capsules/hermes/init.4th`'s `MSG-ALLOC` already admits every message with `SB-DELIVER`, and `MSG-FREE-NODE` (called from both `MSG-ACK-LAST` and heat-driven `MSG-REAP`) already evicts it — `behaviour=DELIVER (stub)` has been printing on boot logs since at least 2026-08-05. Checked whether the dispatch body needed a real payload action the way `MIGRATE` did: `MSG-DELIVER` (the FORTH word) already runs the actual delivery (`VM-EXEC` of the payload) *before* eviction, decoupled from Stadium reap — so by dispatch time delivery is already done, same shape as `COOL`, which needs no extra action beyond `stadium_evict()`'s own universal reservoir credit. **Fix:** `stadium_dispatch()`'s `DELIVER` case now prints the departing message's real identity (`DELIVER msg_idx=N`, same shape as `MIGRATE`'s `lbn=` print) instead of a misleading `(stub)` label — confirmed live via a forced `MSG-SEND`/`MSG-DELIVER-ALL`/`MSG-ACK-LAST` sequence from Hermes's own REPL context (`Stadium: dispatch cell=73653 behaviour=DELIVER msg_idx=1`). `COOL`'s case was in the identical situation (real for both words and VMs, no extra action needed) and, on request, got the same fix (2026-08-26): now prints `COOL identity=N` (word_id for a word, 0 — the patron-zero convention — for a VM) instead of `(stub)`. Confirmed live: both shapes fired correctly on the same boot — `COOL identity=0` at Hermes's/Artemis's own explicit channel-eviction self-test and again at their VM-patron eviction at `PARITY:KILL`, `COOL identity=1` at a second channel eviction — conservation intact throughout. Clean zero-warning compile and clean boot on all three architectures for both fixes. - [ ] `EXPIRE` (ACL) — confirmed genuinely unscoped (2026-08-26), not a case of stale documentation like `DELIVER` turned out to be. Two findings below, then four open questions — **decided directly on request (2026-08-26)**, decision recorded after the questions, no code written (a decision isn't a green light to build, per this session's own convention): **Finding 1 — `acl_ttl` and Stadium's `ttl` are different things wearing the same name.** `DictEntry.acl_ttl` (`vm_core.c:756`) is a per-word countdown that batches how often `ACL-RECHECK` runs — when it hits 0, `acl_recheck()` calls the FORTH word `ACL-RECHECK` (`ACL.4th:50`), which **always renews**: STRICT mode sets `allow=1, ttl=0` (recheck every time); TTL mode computes a fresh heat-based TTL and sets `allow=1`. Denial isn't a live path anywhere in current policy. This is a *renewal* cycle, not a *residency-ending* event — nothing about it resembles "leaving the Stadium floor." **Finding 2 — `StadiumPatronHeader.ttl` is completely inert.** Every candidate constructor across the whole codebase (`stadium.c`, `stadium_words.c`, `stadium_blocks.c`, Hermes's `MSG-ALLOC`/`CH-ALLOC`) sets `candidate.ttl = 0`, and nothing anywhere ever reads, decrements, or reaps on it. The generic TTL-expiry *mechanism* `EXPIRE` would need to fire from doesn't exist in Stadium's own engine at all — a gap one level deeper than "ACL isn't wired to Stadium." **Open questions:** 1. Is "ACL patron" even the right model, or was `FABRIC.md`'s original "ACLs → EXPIRE" mapping a category mismatch from the start — conflating `acl_ttl`'s recheck-amortization counter with Stadium's residency `ttl`? 2. If it is the right model: what gets admitted as a patron? One per ACL-guarded word would be redundant with the word's own patron cell item 4.1 already tracks. A different unit (e.g. per zuse session) might fit better once PKI/session auth lands (Phase 8, still open per `.claude/CLAUDE.md`'s ACL section). 3. Building this for real means building Stadium's generic ttl-decrement/reap-on-zero mechanism first — nothing to hook `EXPIRE` into today. Is that in scope here, or its own separate item? 4. What should the reap action actually *do*, given current ACL policy never revokes — would `EXPIRE` force an `ACL-RECHECK`, or something else entirely? **Decision:** 1. Not at the per-word level. `acl_ttl`-hits-zero always renews, never revokes — forcing `EXPIRE`'s residency-ending tag onto it would misuse the tag. But "ACLs → EXPIRE" isn't wrong in spirit, just aimed at the wrong unit: the one place in this system where something ACL-related genuinely has a lifetime and should be revoked is a **zuse superuser session** (Phase 8, not yet built) — authenticate, hold elevated privilege for a bounded time, then actually drop back to non-zuse. That's a real residency-ending event; per-word recheck isn't. 2. A zuse session, not a per-word ACL entry — a session is the thing with a genuine start/lifetime/end. A word already has its own patron via item 4.1; a second one for ACL purposes would be redundant bookkeeping, not a new concept. 3. **Not in scope now.** This is the decision that actually resolves the item: Phase 8 doesn't exist yet, so there is no session to admit as a patron regardless of any other choice made here. Building the generic ttl-decrement/reap mechanism now, with nothing real to feed it, would be speculative infrastructure ahead of its only consumer — close in spirit to what the no-stubs standard exists to prevent, just inverted (a real mechanism with no real caller, instead of a fake mechanism with a real caller). 4. Revoke the session's elevated privilege and drop the console back to its non-zuse state — a real action, unlike `DELIVER`/`COOL`, which needed none. **`EXPIRE` stays explicitly deferred until Phase 8 (PKI/zuse session minting) lands** — not abandoned, not left ambiguous: revisit it as part of that work, admitting the session itself (not a word) as the patron, once there is something real for it to represent. - [x] Fixed (2026-08-26): `ONTOLOGY.md` §IX's "words (dictionary, warehouse-resident today, not yet migrated)" line was stale — corrected to state words are fully migrated and live via `stadium_word_dispatch()`. Doc-only, no build/boot verification needed. --- ## C. Standing rule: no stubs or TODOs, ever Stated directly, 2026-08-25, after the `stadium_dispatch()` stub investigation above: **"I've never allowed stubs before."** Saved as a persistent memory (`feedback_no_stubs_or_todos.md`) so this applies across sessions, not just this one. Full statement: no stub function that prints a placeholder and returns, no `TODO`-and-move-on comment in place of real logic, in any language, ever committed as if it were finished work. Small, honest increments are still fine and encouraged — each increment just has to be a complete, real implementation of whatever slice it covers, never a placeholder for a later slice. A pre-existing stub found while working nearby (as here) gets flagged and resolved, not built on top of or left in place. --- ## D. Tripod final shape — minting, one-time Zuse fuse, messaging-only (vision, not yet scoped) Stated directly by Captain Bob, 2026-08-25: "we're going to have to iterate because I know what the final shape of the tripod will be." Recorded here as a forward-looking vision, not yet broken into implementable items — captured so near-term Stadium/VM work (§B above) is made against the right end-state rather than in ignorance of it. Full detail also saved as memory `project_tripod_final_shape_vision.md`. - **Thumbdrive presentation → legality check → mint.** A newly presented thumbdrive is checked for legality (against the CA-root-derived identity/certificate scheme already decided earlier this session, `FABRIC-2.md`). If not legal, it is "minted" — formatted for system use — which (1) spins a new user VM and (2) attaches the console VM to it. - **One-time Zuse mint + fuse-blow on first install.** A brand-new system instance ("Install"/"Try It") mints exactly one Zuse superuser, then irreversibly "blows a fuse": direct quote, "we mint one and only one Zuse user and blow a fuse. The only way around is a new system." Post-fuse, no further Zuse can ever be minted on that instance — but the system is NOT bricked: the existing Zuse superuser keeps working, and ordinary users can still "thumb in" via regular thumbdrives. - **Messaging-only once Tripod is fully live.** All inter-VM interaction becomes Hermes messaging, not direct calls/shared state — a stated end-state, not the current implementation. - **Polymorphic block-boundary behavior for user VMs.** Stated in one sentence, not yet elaborated: a user VM needs to behave polymorphically when crossing block boundaries. Needs a dedicated follow-up conversation before this is actionable. - **`ClaudeEXPORT/`** (repo root, new as of 2026-08-25) — a prior Claude data export the user pointed at for "concepts and thoughts as guidelines" on this vision, explicitly flagged as possibly containing superseded/conflicting ideas, not authoritative. `conversations.json` is ~66MB; mine selectively (grep or a subagent) rather than reading wholesale. See memory `project_claude_export_archive.md`. **Why this affects §B:** VM-`COOL` (and any inter-VM Stadium wiring) sits close to this still-iterating design. Building it now risks conflicting with or being thrown away by the Zuse/messaging shape once that gets its own pass — hence §B's punch list defers VM-`COOL` explicitly rather than resolving it in this pass. Block-patron `MIGRATE` work does not obviously depend on this vision and is not deferred for this reason. **Next step:** no implementation here yet. When this gets its own design pass, start by clarifying "polymorphic block-boundary behavior" and the exact legality-check mechanism (presumably the CA-root certificate scheme), then scope minting as its own capsule-birth-style protocol (mirroring the existing capsule birth writeup's rigor) before writing any code.