Files
Robert Allan James d2a0305703 Record driving deadlines for #1 (provisional expiry) and #2 (December deliverable)
Clarified 2026-08-29 with the author: "no hurry" is wrong for #1 and #2.

  - #1 (Hosted StarForth): the provisional will expire before conversion. Its
    priority claim is time-boxed (provisional filed ~Dec 2025, window ~Dec
    2026); if the non-provisional isn't filed claiming benefit before the
    window lapses the priority is lost and the same subject matter can't be
    re-staked by refiling thereafter. Hard clock regardless of December scope.
  - #2 (full StarshipOS): the December deliverable that keeps the flagship
    covered; pending counsel confirmation it may also be the conversion
    vehicle for #1's provisional.

Open point to resolve: is #2 (or the December delivery) the conversion vehicle
for #1, or is #1 converted by a filing separate from #2? Decision record only;
no filing made. Authoritative form in ROADMAP.md; FABRIC-3.md tracks it.
2026-08-29 10:36:19 -04:00

256 KiB
Raw Permalink Blame History

FABRIC-3.md — the Stadium, continued again

Status: Living working document, opened 2026-08-25 as the successor to FABRIC-2.md (now closed/archival — see its own header). This document does not repeat FABRIC-2.md's design argument or history; it restates only outcomes, with pointers back to the section that derived them. Read FABRIC.md for the original "why," FABRIC-2.md for everything derived through 2026-08-25, this document for what's left as of that date onward.

Provenance. Everything in Section A below is a full, non-sampled carry-forward of every open (- [ ]) item in FABRIC-2.md as of 2026-08-25 — 51 items, confirmed by grep -c '^- \[ \]' FABRIC-2.md, none dropped (Section A itself holds 49: the other 2, FABRIC-2.md §F.3's own two checkbox lines, were pure summaries cross-referencing items already listed individually elsewhere — 4.4s/1.11/4.3/§17.4 and 5.1/ACL-RWT re-measurement, both of which are carried forward as their own individual items above — not distinct content, confirmed by diffing item text programmatically before writing this document, not assumed). Extracted mechanically (a script pulling each checkbox item's own text, stopping at the first blank line rather than the next checkbox, to avoid pulling in unrelated already-resolved narrative that happened to sit between two open items in the source document) and spot-checked against the original. Item numbers/labels are carried forward unchanged, for traceability — this is not a renumbering or a re-prioritization. Section groupings match FABRIC-2.md's own (documentation debt, xHCI WRITE(10), Milestone 39 punch lists, etc.) — items are relocated, not reorganized.

How to use this document going forward. New findings, new punch-list items, and new decisions get added here, not to FABRIC-2.md. Follow the same discipline FABRIC-2.md §(intro) established for how work gets picked up, closed, and recorded.


A. Carried forward from FABRIC-2.md (51 items, all still open as of 2026-08-25)

From FABRIC-2.md §A — Blocked or scoped, not started

  • 1.11 — Dirty-event granularity. Leaning region-based. Blocked on item 4.3 — settled as part of the console migration, not speculatively before it. Refs (FABRIC.md): §17.5, §23.2, §23.4 #1.

  • 4.3 — Console. Umbrella item; settles 1.11 as part of the work. Nearly everything under it (4.3.14.3.7f, 4.44.4ac) is done — the parent stays open only because 4.4s below is still blocked and nothing has formally closed the umbrella. Refs (FABRIC.md): §17.5, §27.

  • 4.4s — (user) prompt segment. Scoped, blocked, not started. Extends 4.4's prompt format. Refs (FABRIC.md): §27.8, 4.4.

  • 5.1 — Re-run the DoE on the new substrate. A green POST suite is not evidence that determinism holds under the Stadium migration — needs its own campaign. Not started.

From FABRIC-2.md §D — Design questions still genuinely open

  • §17.4 — the framebuffer utility's internal heat/decay dynamics are undesigned. Explicitly "Open, deferred": not a Stadium patron, but what physics (if any) governs it internally was never designed. Not blocking anything. Blocked on item 1.11 specifically (dirty-event granularity), not "the framebuffer work" in general — see FABRIC-2.md §D's own 2026-08-13 refinement of this item before assuming it's ripe.

From FABRIC-2.md §E — Documentation debt

  • Already done — stale carry-forward, closed 2026-08-26. This item was never marked complete when the actual work concluded. FABRIC-2.md Sections M through T (2026-08-19 to 2026-08-21) already re-ran this measurement — and in the process found the original "ACL-RWT" name itself was wrong: the Rolling-Window-of-Truth mechanism it was named after was dead code, removed 2026-07-08 (ACL-RECHECK-RW was never reachable — acl_recheck() only ever looks up the 11-char ACL-RECHECK, never the 14-char -RW variant). What the original campaign actually measured was the live ACL-TTL mechanism under a misleading name. Final, accepted result (FABRIC-2.md §T): +0.0603% ACL-TTL enforcement overhead, architecture-independent (identical across amd64/aarch64/riscv64), fully deterministic (CV=0.000%) — see also memory project_acl_ttl_overhead_final.md. No new campaign run here; this entry corrects the bookkeeping, not the measurement.

From FABRIC-2.md §J — Maintainability sweep (2026-08-18)

  • Fixed 2026-08-26 (commit cbe7b49), stale carry-forward closed 2026-08-28. docs/lithosananke/ROADMAP.md and M7.1.md — stale Branch: lithosananke (no such branch exists post-split), M7.1.md's "Status: Design Complete" (shipped and live, not just designed), ROADMAP.md's self-contradiction (M8 marked OBSOLETE in one place, still a live success criterion in another), and its stale "AHCI driver" claim for M9 (real implementation is virtio_blk.c) — all corrected by cbe7b49; ROADMAP.md also gained corrected BLOCK/BUFFER/UPDATE/FLUSH and block-device-abstraction [x] marks.

  • Fixed 2026-08-26 (commit cbe7b49), stale carry-forward closed 2026-08-28. Top-level ROADMAP.md (StarForth-era, "Phase 0 Complete... Phase 1 Starting," dated 2025-12-14) — marked OBSOLETE by cbe7b49 (Captain Bob's call: more than stale), pointing to docs/lithosananke/ROADMAP.md and FABRIC*.md for current status.

  • Fixed 2026-08-26 (commit cbe7b49), stale carry-forward closed 2026-08-28. docs/03-architecture/word-acl/DESIGN.md says ACL Phase 7 (LithosAnanke kernel parity) is still "remaining" — direct contradiction with .claude/CLAUDE.md, which states Phase 7 is independently verified complete. Phase 7 restated as independently-verified complete and the stale lithosananke-branch-parity framing removed by cbe7b49.

  • Confirmed accurate, not stale (2026-08-26): VM-FLEET-ATTRACTOR-DESIGN-20260705.md's claim that doe-campaign.4th is "broken and being superseded." Live-ran SMOKE-CAMPAIGN from the current capsule (amd64) — it completes without error and correctly conserves fleet heat (VM-PHYSICS: conserved=CONSERVED, fleet_heat_sum=65536). Initially read that as contradicting the doc's claim; Captain Bob corrected this directly — it's broken. "Doesn't crash" and "runs" are not the same claim: the doc's actual argument is that the capsule has no real controlled-experimental-factor mechanism (no manual heat-injection point under the current design, so it cannot drive the fleet through controlled scenarios the way a DoE campaign needs to), a methodological gap a clean execution trace doesn't surface or disprove. Doc's claim stands; not touched.

  • Tracked (2026-08-26) — real proof-modeling gap, not just stale prose, so not fixed here. Confirmed by reading StarForth_Loop4_Pipeline.thy's own comment (lines 127137): pipeline_metrics_state's pm_last_accuracy_num/pm_last_accuracy_den fraction pair doesn't correspond to anything in the real C struct — include/vm.h's PipelineGlobalMetrics has a single double last_checked_accuracy field, no num/den pair anywhere. The .thy file's own comment already scopes the real fix correctly: "a full field-level pass over pipeline_metrics_state is its own separate task" — matches this project's standing caution that each remaining Isabelle gap needs its own subsystem model, not a documentation-sprint patch. The punch-list ask was tracking this outside the buried .thy comment, which it now has here — the actual re-model stays unattempted, on purpose.

From FABRIC-2.md §X, Milestone 2 — USB hardware stack

  • CLOSED 2026-08-27 (FABRIC-3.md §F.3). Decide and implement where the hotplug event surfaces to the rest of the kernel. Re-checked directly against sk_repl_idle() and the bot_msc_attach_pending/bot_msc_detach_pending doc comments in xhci_driver.h: Milestone 2h's flag-set-by-xHCI/flag-consumed-by-sk_repl_idle() pattern isn't literally a registered callback, but its own doc comment confirms it was built deliberately to achieve the same goal — "keeps xhci.c decoupled from block_subsystem.c" — and that decoupling is real, live on all three arches. Closed as written; no further design work on this node itself.

  • BUILT + VERIFIED 2026-08-28 (FABRIC-3.md §F.17). CBW/data/CSW for SCSI WRITE(10), a direct mirror of READ(10) with the data direction flipped. Verified live via a genuine cold-reboot round-trip on all three architectures (amd64/aarch64/riscv64), not just a same-session read. blkio_usb.c is read-write now; read_only flipped to 0. Real hardware (Milestone 8) remains untested — QEMU-only so far.

  • Implement basic error/stall recovery (CSW failure status, endpoint stall clear) — at minimum enough to not wedge the controller on a single bad transfer. SCOPED 2026-08-27 (FABRIC-3.md §F.14): full BOT-spec recovery designed — new STALL_ERROR completion code, new xHCI Reset Endpoint + Set TR Dequeue Pointer commands, CLEAR_FEATURE(ENDPOINT_HALT) via existing control-transfer plumbing, escalating to a full Bulk-Only Mass Storage Reset on a second stall, bounded via a new retry counter mirroring bot_tur_retries's own precedent. Real gap closed today: CSW status handling (PASS/FAILED/PHASE ERROR) already existed; nothing distinguished a STALL specifically, and no recovery of any kind existed before this pass.

From FABRIC-2.md §X, Milestone 3 — Block subsystem extensions

  • Implement the CA-signed-cert verification path (Milestone 6 dependency — the cert chain validator doesn't exist yet either). Now confirmed a hard prerequisite of WIREBIND too (FABRIC-3.md §F.5), not just an M3 item in isolation. SCOPED 2026-08-27 (§F.7): format and trust root decided (Zuse-signed X.509/DER cert, not the offline capsule-PKI chain); the actual DER signature-extraction code is new work, not yet built.

  • Implement the first-touch allocation function: given a verified identity pubkey and a requested block count, either read an existing range from the drive's map or claim a new one at g.total_user_lbn and write it back. (Single-block relocation itself — the mechanism this would allocate ranges for — is done: blk_subsys_relocate_block()/ RELOCATE-BLOCK, FABRIC-2.md, commit 36d832f. This item is about the identity→range allocation that decides what to relocate blocks into*, still unbuilt.)* Scope clarified 2026-08-27 (FABRIC-3.md §F.6, decision 3): a user's pool is their entire thumbdrive by default (never contested) plus any system-resident device blocks they additionally claim, first-come-first-served — this function only governs that second, system-side extension, not the drive itself. SCOPED 2026-08-27 (FABRIC-3.md §F.11): the "claim at g.total_user_lbn" phrasing above is stale, predating BMAPFMT's distributed-ownership decision — claims are found by scanning Artemis's already-attached device's own blk_meta_t records for unowned devblocks, not by extending a counter. Claims may be a scattered chain (reusing blk_meta_t's existing prev_block/next_block/chain_length fields), found via a full linear scan each time, no cached index.

  • SCOPED 2026-08-27 (FABRIC-3.md §F.4). Design the on-drive block-map format (Section U item 4). Resolved as: no separate table — repurpose the existing, fully-wired-but- zero-callers blk_meta_t/blk_get_meta()/blk_set_meta() (block_subsystem.c), replacing its stale POSIX-flavored ownership fields with a pubkey-fingerprint owner + a fast-deny ACL bit + deliberate reserved slack. Field design done; the actual code edit is not.

  • CLOSED 2026-08-27 (FABRIC-3.md §F.12) — already built, not a separate item. blk_set_meta()'s dirty-tracking already flows through the exact same cache_writeback()/ blk_flush() path ordinary block data uses; the whole devblock (data + metadata) already gets written to the device, real and unstubbed. Blocked only on WRITE(10) itself, already modeled in the graph — nothing else stands between this and working.

  • CLOSED 2026-08-27 (FABRIC-3.md §F.13) — already built, not a separate item. cache_load_devblock() already unpacks and validates every devblock's metadata on load. Tracing "on insertion" surfaced a real, separate, previously-unflagged conflict instead: homeblocks_sig_t and the block-subsystem's own generic 'STFR'/v2 header both wanted devblock 0, which would leave a home-blocks drive permanently write-refused or overwrite its own identity header. Fixed by relocating homeblocks_sig_t to devblock 1; MINT must also run the ordinary format-confirm path at devblock 0.

  • Design the migration state machine (Section U item 5) — states, transition triggers. Session direction, 2026-08-25: ACL manages when to relocate (capacity pressure, or a compudynamics heat/cold signal); migration itself is expected to be rare, not routine. The physics_hotwords_cache.c-reuse question is settled differently than originally framed — see this document's new §B below (Stadium unification), which reframes block/word placement as a compudynamics.c-driven decision generically, not a physics_hotwords_cache.c (DictEntry*-hardcoded) reuse question specifically.

  • CLOSED 2026-08-27 (FABRIC-3.md §F.10), scoped alongside DETACH. Unclean removal runs the exact same VM-kill path EJECT (M5) uses, minus the flush step — skipped outright since the device is confirmed gone by the time the existing bot_msc_detach_pending signal fires, not attempted-and-ignored. Design done; the actual wiring into sk_repl_idle()'s existing detach handling is not yet written.

From FABRIC-2.md §X, Milestone 4 — Drive/credential security

  • Designed (2026-08-26), not yet implemented — Phase 8 kickoff. Mirrors two existing precedents exactly: CAPSULE_MAGIC_PACK's bit-packed magic (include/starkernel/capsule.h) and blk_volume_meta_t's magic+version+fields+pad-to-4096 structural convention (include/block_subsystem.h). Lives at the first 4KiB devblock of the GPT metadata partition (the ~1GB partition decided 2026-08-22) — the header format doesn't depend on the still-missing GPT parser; it's just what gets written starting at that partition's first devblock once something can locate it. Deliberately narrow: identifies/authenticates the drive only, does not invent the block-map or credential/cert formats (both separate, still-open items below) — reserves offset/size pointers to where they'll live instead of embedding them.

    #define HOMEBLOCKS_SIG_MAGIC   0x4248414CULL /* 'LAHB' -- LithosAnanke Home Blocks,
                                                    * same little-endian ASCII packing as
                                                    * CAPSULE_DESC_MAGIC's 'CAPS' */
    #define HOMEBLOCKS_SIG_VERSION_0 0
    
    #define HOMEBLOCKS_SIG_PACK(ver) \
        (HOMEBLOCKS_SIG_MAGIC | ((uint64_t)(ver) << 32))
    #define HOMEBLOCKS_SIG_GET_MAGIC(m)   ((uint32_t)((m) & 0xFFFFFFFFULL))
    #define HOMEBLOCKS_SIG_GET_VERSION(m) ((uint8_t)(((m) >> 32) & 0xFF))
    
    typedef struct {
        uint64_t magic;              /* HOMEBLOCKS_SIG_PACK(...) */
        uint8_t  drive_uuid[16];     /* unique per-mint instance id -- Phase 8 mints multiple
                                      * distinct drives, needs something to tell them apart */
        uint64_t minted_time_ns;
        uint64_t metadata_devblocks; /* size of this GPT metadata partition, in 4KiB devblocks
                                      * -- sanity/bounds check against the GPT entry once a
                                      * parser exists */
        uint32_t cert_offset;        /* devblock offset within this partition where the
                                      * CA-signed cert blob starts; 0 = not yet minted */
        uint32_t cert_devblocks;
        uint32_t blockmap_offset;    /* devblock offset where the block-map (Milestone 3,
                                      * format still undesigned) starts */
        uint32_t blockmap_devblocks;
        uint64_t hdr_crc;            /* REAL from day one, not a placeholder like
                                      * blk_volume_meta_t's "unused yet" hdr_crc -- this
                                      * header's whole job is gating a warn-and-refuse
                                      * security check below, so the crc has to actually work */
        uint8_t  _pad[4096 - 64];    /* pad to one devblock; 64 = sum of the fields above */
    } homeblocks_sig_t;
    

    cert_offset/cert_devblocks are 0 on first design pass — the real CA hasn't been generated yet (Milestone 6), so this reserves the shape of where a cert will attach without committing to a cert format that doesn't exist. Same reasoning for blockmap_offset/blockmap_devblocks against Milestone 3's still-open block-map design. Implemented (2026-08-26): include/starkernel/homeblocks_sig.hhomeblocks_sig_t + HOMEBLOCKS_SIG_PACK/_GET_MAGIC/_GET_VERSION macros, a C99 compile-time size assertion (same discipline stadium.h's own header-size checks use), and the identical field layout shown above. Verified standalone: sizeof(homeblocks_sig_t) == 4096, compiles clean under -std=c99 -Wall -Wextra -Werror. Not yet consumed by any code — nothing in the block subsystem or xHCI driver reads or writes it yet, so no functional kernel change and no 3-arch acceptance boot needed for this step; that starts with the signature-check implementation, the next punch-list item below.

  • Implemented (2026-08-26): the check function itself, real and complete — not yet wired to any write path. include/starkernel/homeblocks_sig.h + src/starkernel/homeblocks_sig.c: homeblocks_sig_check(dev, sig_start_fblock, out_sig) reads the 4 consecutive 1KB blkio forth-blocks the 4KB header spans, verifies magic → version → CRC-64 in order, returns one of HOMEBLOCKS_SIG_OK/_BLANK/_BAD_VERSION/ _BAD_CRC/_READ_ERROR. Reuses block_subsystem.c's existing CRC-64/ISO (compute_crc64, previously static/file-local, now exposed) rather than a second CRC implementation — same algorithm already proven via per-block checksums. Takes the header's starting block as a plain parameter rather than resolving it internally: this function verifies a signature given a location; finding that location (GPT-partition-relative, once a parser exists) stays the caller's job, not invented here.

    Verified against the actual shipped code, not a reimplementation: a standalone host test links the real homeblocks_sig.c against a fake in-memory blkio_dev and exercises all four outcomes — blank media → BLANK, a correctly-minted header → OK (round-trips drive_uuid/minted_time_ns correctly), a flipped CRC → BAD_CRC, an unrecognized version → BAD_VERSION. All four pass. A full QEMU-hotplug live test isn't proportionate yet — nothing calls this function from the live kernel path (deliberately; wiring it into the attach path is the next item below), so a live boot check has nothing to exercise. Clean zero-warning compile and clean boot on all three architectures confirms no build/link regression from exposing compute_crc64 and adding the new source file to every kernel build.

  • Implemented (2026-08-26): the "warn" half, live and wired. Wired into sk_repl_idle()'s USB hotplug attach handler (repl.c), right between blkio_usb_open_msc() succeeding and blk_subsys_attach_device() — calls homeblocks_sig_check(&usb_blk_dev, 0, &sig) and logs a distinct message per outcome (recognized / blank-or-foreign / bad-version / bad-crc / read-error).

    The "refuse" half is deliberately not implemented — there is nothing real to gate yet. blkio_usb.c has no SCSI WRITE(10) support at all (Milestone 2's biggest open item), so there is no write path today to refuse; the only thing attach currently enables is read-only access, which is also the general-purpose USB block I/O path this repo already relies on for unrelated testing, not exclusively a home-blocks identity workflow. Refusing attach on blank media would have broken that legitimate use without protecting anything real — building that gate now would be enforcement with no live consumer, the same "don't build ahead of a real caller" reasoning EXPIRE's deferral used. Refuse belongs on the write path once WRITE(10) exists to give it something to gate against.

    sig_start_fblock is hardcoded to 0 at the call site — correct for today's unpartitioned raw test/real media (no GPT parser exists yet), explicitly flagged in the code comment as the one place that will need to change to a real GPT-partition-relative lookup once that parser lands, isolated from homeblocks_sig.c's own location-agnostic check logic.

    Verified live, not just compiled: hot-attached disk/usb-thumbdrive-test.img (blank media, no LAHB magic) through a running amd64 instance's QMP socket (blockdev-add + device_add usb-storage,bus=xhci0.0) — captured exactly 4 real TUR+READ10 BOT cycles (matching the header's 4 forth-block span) followed by xhci: USB drive not recognized (blank or foreign media) -- read-only general use only, then normal attach completing successfully afterward (no regression — no MSC block-subsystem attach failed). Conservation intact, no panic. Clean zero-warning compile and clean boot on all three architectures.

  • Resolved (2026-08-26): reuse acl_pinned directly, no new flag needed. The open question assumed credential data "isn't a dictionary word" — but the design already chosen for it (ZUSE-CERT-LO/HI, ACL-CA-KEY-LO/HI) are CONSTANT words, i.e. real DictEntrys. acl_pinned's enforcement is more general than assumed: vm_create_word() (dictionary_management.c:394-404) — the single choke point every word-defining construct goes through — unconditionally refuses to let anything shadow a pinned name ("Pin is permanent: no word may shadow a pinned entry — ever, by anyone"), a real general redefinition guard, not just an ACL-mode-change lock. zuse.4th's ACL-ZUSE-BOOT already pins both cert constants today — the mechanism is already wired for this.

    Real gap surfaced along the way, not yet fixed: ACL-ZUSE-BOOT self-activates and pins ZUSE-CERT-LO/HI unconditionally on every boot — before any legitimate minting step could ever run, permanently locking in the 0 placeholder on the very first boot. The boot sequence needs to distinguish "already minted, pin it" from "not yet minted, don't pin yet" before minting can work at all.

    This directly shaped the next design pass (2026-08-26, Captain Bob): a dedicated disk/zuse.img QEMU test thumbdrive, "bleachable" back to pristine/unminted state for repeated first-boot testing; a one-time first-boot mint-Zuse flow (mint → write real cert → blow the fuse → then pin, resolving the gap above); a separate, ongoing S" name" MINT word for an authenticated Zuse session to mint additional regular users; and a Zuse recovery path, explicitly flagged as unresolved and risky if rushed — not to be designed casually, since the earlier "no software recovery, mint a new one" rule existed specifically to close a hole a careless recovery path could reopen.

    First piece implemented (2026-08-26): the zuse.img bleach mechanism. disk/zuse.img (64MB, blank, matching the existing USB-fixture convention exactly — see disk/README.md) + scripts/bleach_zuse_img.sh (idempotent reset back to blank). Verified live: hot-attached via QMP (same method as the warn-and-refuse verification above) — reads back as HOMEBLOCKS_SIG_BLANK (xhci: USB drive not recognized (blank or foreign media)), correctly simulating a genuine first boot. Deliberately flat/raw, not GPT-partitioned, matching homeblocks_sig_check()'s current sig_start_fblock=0 call site — both move to a real GPT-partition-relative offset together once a parser exists, not attempted here. No kernel code touched this step (host-side test tooling only), so no 3-arch acceptance boot needed — single live amd64 QMP-hotplug confirmation is the right verification tier.

    Correction (Captain Bob, 2026-08-26): drives are not bound to any particular size. 64MB was only ever this fixture's arbitrary test-convenience size, matching usb-thumbdrive-test.img's existing BOT-driver-testing precedent — never a real-world constraint. Audited for anywhere this might have implied otherwise: the actual format (homeblocks_sig_t) already carries its own metadata_devblocks field and hardcodes no size anywhere, confirmed clean. scripts/bleach_zuse_img.sh gained a --size-mb override so this was never a hidden assumption baked into the tooling either. The earlier project_usb_thumbdrive_gpt_layout.md memory's "16GB reference size" phrasing (already hedged as tentative, but risked reading as a target) corrected to state the point explicitly — the GPT layout's design point is the proportions (small metadata partition, everything else block storage), not any absolute size.

    Still open, not attempted: the mint-then-pin boot-sequence fix itself, the MINT word, and the Zuse recovery path.

    Correction, supersedes the "reuse acl_pinned" resolution above (2026-08-26): a pinned CONSTANT is not actually tamper-proof. ACL-PIN/acl_pinned only guards against redefinitionvm_create_word()'s pin check blocks a second : ZUSE-CERT-LO ... ;, but nothing stops ' ZUSE-CERT-LO >BODY ! from overwriting the same word's data field in place. A CONSTANT's value lives in its data field, so the earlier design left the cert mutable from FORTH despite being "pinned." Found while starting the mint-then-pin boot-sequence fix itself; fixing that gap came first since building a real mint flow on top of a tamperable store would just re-open the hole later.

    Fixed (2026-08-26): moved cert storage out of the dictionary entirely. New VM struct fields (include/vm.h): zuse_cert_lo/zuse_cert_hi (the cert value) + zuse_cert_installed (one-time fuse bit). New vm_zuse_cert_install(vm, lo, hi) (src/vm.c) — C-only, no FORTH word wraps it, returns -1 on a second call rather than silently re-installing (a second call is a caller bug, not a runtime condition to recover from). No FORTH store word exists or should exist for these fields, closing the >BODY path structurally rather than by convention. Three new read-only C primitives (src/word_source/starforth_words.c, same shape as the existing HEARTBEAT-TICKS@): ZUSE-CERT-LO@, ZUSE-CERT-HI@, ZUSE-CERT-INSTALLED?. capsules/zuse.4th's old ZUSE-CERT-LO/HI CONSTANT words (and ACL-ZUSE-BOOT's two now-pointless ACL-PIN calls on them) deleted outright rather than left as dead/insecure scaffolding — mkcapsule --lint clean (31/31) after the edit. vm_zuse_cert_install() has no caller yet: the real mint flow still doesn't exist (Milestone 6 CA + the MINT word are both still open), and calling it with a placeholder value would just be a stub wearing the shape of a fix — so this stays an honest, complete slice (storage + read accessors) with the actual mint-then- pin sequence still explicitly open, not faked. Verified: hosted make build clean, zero warnings; clean boot to ok> on all three architectures (amd64/aarch64/riscv64), Stadium conservation intact (43691/21845/65536) on all three, no panics or guest errors. ACL is opt-in (init.4th's S" ACL.4th" EXEC commented out by default) so the new words weren't exercised live from the REPL this pass — compile/lint/boot verification only.

    MINT word design, picked up 2026-08-26. Before scoping MINT itself, found a real conflict with an existing, deliberate decision: include/starkernel/ed25519.h is verify- only by design — "this kernel never signs or generates keys (no entropy source to do so safely anyway); signing happens in the host-side build tool" (FABRIC-2.md, Milestone 6, 2026-08-22). But the vision for MINT (§D below) is an interactive, on-device S" name" MINT word — an authenticated Zuse session signing a new user's cert live, at runtime. A kernel that structurally never signs can't do that as envisioned. Raised directly; decided (Captain Bob, 2026-08-26): give the kernel a real signing capability rather than reshape MINT around verify-only. This reopens the prior "no entropy source" constraint deliberately, not by accident.

    Phase A — virtio-rng, done 2026-08-26. Checked what entropy is actually available before choosing a design: include/starkernel/vm_uuid.h already found, for VM UUIDs, that amd64 has RDRAND and riscv64 has the Zkr extension, but QEMU's aarch64 CPU models (including max) expose neither RNDR nor any RNG property at all — confirmed directly against QEMU 10.2.1. That's why VM UUIDs use a deterministic PRNG uniformly instead of a per-arch split; that same choice is not safe for Ed25519 keygen — a seed drawn from a known value makes the private key predictable. Decided: add a virtio-rng device instead of a per-arch RDRAND/Zkr split with a weaker aarch64 fallback — QEMU supplies real host entropy identically on all three arches, closing the aarch64 gap directly (QEMU-only; real hardware at Milestone 8 needs a real per-arch RNG driver, a separate later problem).

    New include/starkernel/virtio_rng.h + src/starkernel/virtio/virtio_rng.c, transport plumbing (PCI capability walk, common-cfg feature negotiation, split virtqueue) mirroring the existing virtio_blk.c exactly — same device family, same quirks. Simpler shape than block: one virtqueue, one device-writable descriptor, no request header or status byte (the entropy device has none); virtio_rng_get_bytes() loops internally since the device may return fewer bytes than requested per round. -object rng-random,id=rng0,filename=/dev/urandom

    • -device virtio-rng-pci added to all three arches' QEMU invocations (Makefile.starkernel). Wired into boot (kernel_main.c, right after the existing virtio_blk_find_artemis() call site, same graceful-noop-on-absence precedent).

    Verified live, not just compiled: a temporary probe (written, run once, captured, reverted — per this project's standing probe convention) pulled 16 real bytes through the full request/notify/poll/used-ring round trip on all three architectures and printed them: amd64 be9223909b86a8ccbfff705ccae2caa6, aarch64 861487df65a6d7a26b2c9c34f5ff96ee, riscv64 fec51d80e8c169a35aad9d908eab2f39 — three different values, confirming real entropy, not a stale or repeated buffer. Probe reverted; permanent code is just the driver + init call. A second, final 3-arch acceptance boot ran against that reverted code (not the probe build) to confirm the shipped state itself is clean. Clean zero-warning compile and clean boot to ok> on all three architectures, Stadium conservation intact (43691/21845/65536), no panics or guest errors on either pass.

    Still open: Phase B (real Ed25519 keygen/signing, seeded from this entropy) and Phase C (the MINT word itself, cert format, and whether Zuse's own keypair needs to chain to the Milestone 6 offline root CA or is a self-sovereign instance-local root of trust).

    Phase B — real Ed25519 keygen/signing, done 2026-08-26. Extended include/starkernel/ed25519.h/src/starkernel/crypto/ed25519.c (previously verify-only) with ed25519_keygen(seed, pubkey_out) and ed25519_sign(seed, msg, msg_len, sig_out), per RFC 8032 §5.1.5/5.1.6, reusing every point-arithmetic primitive verify already had (scalar_mult, point_compress, the base-point constants) — no new curve code, only the seed-expansion/clamping and per-message nonce derivation verify never needed. Signing is deterministic (nonce derived from seed+message, not fresh randomness): only keygen ever touches entropy, via a caller-supplied seed (virtio_rng_get_bytes(), Phase A) — keygen itself still generates nothing and trusts the caller for randomness quality, matching this file's original design philosophy exactly.

    New scalar_muladd() (scalar25519.c/.h) for signing's S = (k*a + r) mod L step, the one piece of scalar arithmetic verify never needed (verify only ever reduced or compared, never multiplied scalars). Schoolbook 256×256-bit multiply into a u128 wide accumulator with exactly one final carry-propagation pass — deliberately the same shape as fe25519.c's existing field multiply, because that file's own history records a real bug from trying to carry mid-accumulation instead of in one final pass; structurally can't repeat that mistake this way. Reduces the result via the existing, already-proven scalar_reduce512() rather than writing new modular-reduction logic.

    Verified against an independent implementation, not self-consistency — this project's own standing lesson (two real, invisible-by-inspection bugs in the original from-scratch field arithmetic, an off-by-one-hex-digit hand-transcribed SHA-512 vector) means a passing self-check proves nothing on its own. Built a throwaway host test harness (compiled, run, discarded — the crypto files have no __STARKERNEL__ gate, so they link as an ordinary Linux binary) against Python's cryptography library (OpenSSL-backed). Six trials — five random seed/message pairs (message lengths 1, 32, 255, 1000 bytes) plus the empty-message case — every one produced a byte-for-byte identical public key and signature to the independent implementation, not just a signature this codebase's own verify accepted. Verified on-target too: clean zero-warning compile of the crypto files on all three architectures, and a full 3-arch QEMU acceptance boot (amd64/aarch64/riscv64) — all clean to ok>, Stadium conservation intact, no panics or guest errors. Nothing calls ed25519_keygen()/ed25519_sign() from the live kernel path yet (Phase C's job); this pass is compile/link/boot-regression verification for the crypto library itself.

    Still open: Phase C — the MINT word itself, cert format, and whether Zuse's own keypair needs to chain to the Milestone 6 offline root CA or is a self-sovereign instance-local root of trust.

    Phase C scoping, 2026-08-26. Three findings before any code:

    1. Resolved, not a real conflict: Zuse doesn't need Milestone 6's CA. That CA chain is specifically for capsule/code signing (root → snakeoil intermediate → per-capsule Ed25519 signatures verified at capsule-load time) — a different trust domain from user identity. The vision's own framing ("we mint one and only one Zuse user and blow a fuse ... the only way around is a new system") already implies Zuse's authority comes from being the unique first-boot mint on this instance, not from an external chain. Decided: Zuse is a self-sovereign, instance-local root of trust, keypair generated on-device from real entropy (Phase A+B). Regular users, minted later via MINT, get certs signed by Zuse's key, not the Milestone 6 CA — two independent PKI domains.
    2. A real gap in this session's own earlier work: vm_zuse_cert_install(vm, lo, hi) (the very first change this session made, before Phase A existed) only holds two uint64_t (16 bytes) — sized against the old placeholder ZUSE-CERT-LO/HI FORTH-cell design, not against what a real Ed25519 keypair needs (32-byte pubkey alone, well over 100 bytes for a full cert). Needs expanding before Phase C can store anything real.
    3. A genuine blocker, found by asking where the cert would actually live: "mint once, ever" requires surviving reboots, but Makefile.starkernel's qemu target copied a fresh, pristine OVMF_VARS.fd on every invocation (not just after clean) — so a UEFI-NVRAM-based cert (the real-hardware-compatible option, and this codebase already has a live SetVariable/GetVariable precedent via SF_VAR_REBOOT_TRIES/SF_VAR_BOOT_ARGS) would never actually persist under this project's own normal test workflow. Digging further: aarch64's qemu recipe had no persistent NVRAM store at all — a single combined -bios $AAVMF_CODE argument, no separate writable VARS pflash drive like amd64/riscv64 have. Decided (on request): fix the harness rather than switch substrates. amd64/riscv64: the VARS-template copy is now conditional on the destination not already existing, so clean (which deletes the whole build/$(ARCH)/kernel tree, OVMF_VARS.fd/RISCV_VARS.fd included) is the bleach step, and a bare make qemu now preserves NVRAM across runs — exactly matching the existing "always pass clean before qemu" acceptance convention, no new bleach script needed. aarch64: restructured to split CODE(ro)/VARS(rw) pflash drives matching the other two (host has /usr/share/AAVMF/AAVMF_VARS.fd alongside the existing AAVMF_CODE.fd), with a graceful fallback to the old single--bios mode (and a console note) on a host that only has non-split firmware packaged, so this doesn't regress environments without one. Verified live: all three architectures still boot clean to ok> with the new pflash arrangement, Stadium conservation intact, no panics or guest errors — this is infrastructure-only (no cert code yet), so a plain boot-regression check is the right verification tier.

    Still open: the actual cert struct (expanding past the 16-byte placeholder), the first-boot mint-vs-already-minted boot sequence using SetVariable/GetVariable, and the MINT word itself.

    Cert struct expanded (2026-08-26): vm_zuse_cert_install() (both src/vm.c's hosted copy and a new kernel-side duplicate in src/starkernel/vm/vm_core.c -- the kernel build's VM_EXCLUDE list drops src/vm.c entirely, same reason vm_set_base() already has two independent copies) now takes a real 32-byte seed + 32-byte pubkey instead of the old 16-byte placeholder. FORTH-side ZUSE-CERT-LO@/HI@ replaced with ZUSE-PUBKEY@ ( i -- u ) (8-byte LE chunk i, 0..3, of the public half only -- the seed has no FORTH access at all). ACL-ZUSE-BOOT now checks ZUSE-CERT-INSTALLED? before authenticating rather than authenticating unconditionally. Verified: clean compile and clean boot on all three architectures.

    NVRAM persistence attempt: crashed, root-caused, reverted -- do not retry as designed. First attempt placed the mint-or-load GetVariable/SetVariable logic right after virtio_rng_init() (before capsule_birth_mama()); it page-faulted (CR2 inside the OVMF flash MMIO window, a supervisor write to a not-present page) partway through boot. Moved the same logic to the one place in this codebase already calling SetVariable post- ExitBootServices successfully (SF_VAR_REBOOT_TRIES, much later in boot) — identical crash, same RIP and CR2 — which disproved the "too early in boot" theory outright: it isn't a timing issue.

    Localized precisely (advisor-directed, one boot, debug markers around each call): GetVariable returns fine. SetVariable with real 64-byte data never returns — that's the exact fault site. The pre-existing SF_VAR_REBOOT_TRIES call that looked like a working precedent is actually a delete of a variable that's never existed (size=0, data=NULL) — a fundamentally different, much cheaper internal path than a real data write, so it proved nothing about real persistence being safe. Root cause: this kernel's VMM never maps whatever memory region OVMF's variable service needs to actually write flash-backed variable data — a real gap in UEFI runtime-services support, not specific to Zuse. Fixing it for real means walking the UEFI memory map for the relevant regions and mapping them into the kernel's own page tables, and per Section U's own note, the flash window's location is firmware/arch-specific (OVMF's differs from AAVMF's and EDK2-riscv64's), so "walk the map and map everything" is not guaranteed 3-arch-uniform even once attempted.

    Second, independent finding (not a bug, a design flaw in the persistence choice): storing the raw 32-byte seed in NVRAM was a real defect regardless of the crash. SetVariable was called with EFI_VARIABLE_RUNTIME_ACCESS, meaning any later-loaded UEFI application or the booted OS itself could read Zuse's private key straight out of NVRAM. For an irrevocable "one and only one Zuse, ever" root of trust, that undermines the property the design exists to provide — this would have needed fixing even had the crash not happened.

    Decision needed, not yet made: given virtio-blk writes are already proven working on all three architectures in this repo (vblk_write, Artemis's own persistence across runs), a dedicated file-backed system-identity disk (mirroring disk/artemis.img's existing pattern, separate from Artemis's internal storage and separate from home-blocks USB thumbdrives) is the substrate with no open unknowns today — recommended over either fixing the UEFI flash-mapping gap (real but large, unscoped VMM work) or accepting the NVRAM approach as originally designed (has the exposed-seed defect regardless). Not decided or built yet. Reverted to a known-safe state: all Zuse mint/NVRAM code removed from kernel_main.c (only two harmless includes remain), init.4th's ACL.4th line back to its documented commented-out default. Verified clean compile and clean boot on all three architectures in this reverted state.

    Substrate corrected (Captain Bob, 2026-08-26): no files, ever — this OS's entire reason for being is anti-POSIX, anti-file. The "dedicated system-identity disk" recommendation above was framed in file/filesystem language by mistake; corrected before any code was written. The only real persistence primitives here are content-addressed capsules and raw LBN-numbered blocks (block_subsystem.c) — never a filesystem, never file paths. Saved as feedback_no_files_anti_posix.md so this isn't re-learned next session.

    Design, agreed on request: a growable metadata fence at the TOP of a device's block space, mirroring the bottom BAM reservation from the opposite end. block_subsystem.c's BAM already reserves the bottom BLK_DISK_SYS_RESERVED (32) blocks of every attached device, invisible to FORTH's BLOCK/BUFFER. Zuse's cert (and future system metadata) gets a second reservation at the top of the same device, starting at BLK_META_FENCE_INIT (128) blocks and growing downward as needed — the two reservations grow toward each other from opposite ends, never colliding, same shape as a stack/heap. Explicitly never RAM-backed (the fast-RAM/ramdrive LBN ranges are documented as volatile in this same file's own header comment — losing Zuse's identity to a RAM eviction is exactly the failure this is designed against). Reuses Artemis's own already-attached, already-proven virtio-blk device — no new device attachment. Rejected reusing BAM's own bottom-reserved zone directly: those 32 blocks are fully claimed by BAM/volume-metadata bookkeeping, not free space.

    Step 1 (field round-trip) implemented and verified 2026-08-26, allocator not yet touched. New meta_fence_blocks field in blk_volume_meta_t, appended after reloc_devblocks and carved from _pad[] — identical graceful-default technique the reloc_start/reloc_devblocks fields already established (a pre-existing formatted volume reads the field back as 0 via its zeroed former padding, not a format-breaking change). Added a compile-time _Static_assert(sizeof(blk_volume_meta_t) == 4096, ...), same discipline homeblocks_sig.h already uses — caught a real bug immediately: the hand-summed _pad[] size formula was off by 4 bytes (a compiler-inserted alignment gap before tracked_blocks that the manual byte-count missed), found via offsetof() rather than by re-deriving the arithmetic by hand again, consistent with this project's standing rule to never trust a hand-derived numeric claim in this class of code.

    Worked against disposable clones throughout, never the real disk/artemis.img (ARTDISK=... is ?=-overridable) — disk/artemis-metafence-fresh.img (blank, exercises the fresh-format path) and disk/artemis-metafence-test.img (a copy of the pre-existing artemis.img, exercises the graceful-default-on-reload path), both kept as regression fixtures per disk/README.md's existing convention (mirrors artemis-reloc-test.img exactly). Verified independently via direct byte reads of the disk image, not the kernel's own self-report (log_message(LOG_INFO, ...) turned out not to reach serial output at all in this build — an unrelated, pre-existing log-level gap, not a regression): fresh format writes meta_fence_blocks=128 at header byte offset 184; a second boot without reformatting reads it back unchanged; the pre-existing old-format image correctly reads back 0. Full 3-arch acceptance boot against the real, untouched disk/artemis.img also clean — conservation intact, no panics.

    Step 2 (allocator + read/write accessors), done 2026-08-26. Units corrected from "Forth 1 KiB blocks" to 4 KiB devblocks (matching bam_devblocks/reloc_devblocks) before anything depended on the original meaning — a clean fix, not a migration, since nothing consumed the field yet. This let the fence fold directly into compute_totals_from_B()'s existing payload4k calculation (total_devblocks - 1 - B - R - F, F = meta_fence_blocks) instead of needing a second, separate subtraction against user_blockstotal_blocks, user_blocks, and free_blocks all shrink correctly for free, in both the fresh-format and reload code paths, from this one formula change.

    New blk_meta_zone_read()/blk_meta_zone_write() (block_subsystem.c/.h) — raw, unpacked 4 KiB devblock I/O (no Forth-block packing, same shape as the header/BAM/reloc-table regions), addressed by devblock_from_top counting down from the device's last physical devblock, refusing (not silently clamping) if the index isn't within the on-disk meta_fence_blocks. No FORTH word wraps either — C-only, same discipline as vm_zuse_cert_install() itself, which will be this zone's first real tenant.

    Verified independently at every step, never trusting the kernel's own report:

    • Capacity math: read a freshly-formatted image's header bytes directly and independently recomputed the expected total_blocks in a separate Python script using the same formula — exact match (22647, down from what it would have been without the fence).
    • Accessor correctness: a temporary probe (written, run, captured, reverted) wrote a known 256-byte-repeating pattern via blk_meta_zone_write(0, ...), read it back via blk_meta_zone_read(0, ...), and compared in-memory (PASS) — then, independently, read the raw image file at the exact expected physical byte offset ((total_devblocks-1)*4096) and confirmed the pattern landed there byte-for-byte.
    • log_message(LOG_INFO, ...) still doesn't reach serial output in this build (same pre-existing gap noted in Step 1) — all verification here used console_println (which does reach serial) for the temporary probe, and direct file reads for everything else.

    Full 3-arch acceptance boot (real, untouched disk/artemis.img, probe code fully reverted) clean on all three architectures — conservation intact, no panics.

    Still open: wiring vm_zuse_cert_install()'s seed+pubkey to actually persist through these new accessors (the zone exists and works; nothing writes Zuse's cert into it yet), and the MINT word itself.

    Step 3 (Zuse's cert wired to the fence), done 2026-08-26 -- first-boot mint-then-load is real, end to end. New include/starkernel/zuse_cert_devblock.h: a small, standalone on-disk record format (zuse_cert_devblock_t -- magic + version + 32-byte seed + 32-byte pubkey + a real CRC-64/ISO from day one, same "real from day one" discipline homeblocks_sig_t already established, since this gates a real security check) occupying devblock_from_top=0 of the fence. Deliberately its own header, not inlined at the boot-time call site: the still-open MINT word will be a second consumer of this exact format later.

    kernel_main.c's mint-or-load logic moved from the crashed NVRAM approach to this: read devblock 0 of the fence, and if magic/version/CRC all check out, install the existing cert; otherwise, if virtio_rng is ready, mint a fresh one (Phase A+B) and write it. Runs right after virtio_rng_init(), well before capsule_birth_mama() -- unlike the crashed NVRAM attempt, raw block I/O against Artemis's already-proven virtio-blk device has no boot-timing risk at all, so the earlier "re-invoke ACL-ZUSE-BOOT after Mama birth" workaround is no longer needed; ACL.4th/zuse.4th's self-activating ACL-ZUSE-BOOT sees a correctly-populated cert on its one, ordinary first pass.

    Verified live, independently, across every real scenario, never trusting the kernel's own report:

    • Fresh mint (blank disk/artemis-metafence-fresh.img): boot logs Zuse: minted, fuse blown; the on-disk record at the exact expected physical offset independently decodes to magic bytes b'ZUSE', version 1, a real 32-byte seed and pubkey, and a CRC that an independent from-scratch Python re-implementation of the exact CRC-64/ISO algorithm (table generation included, not just the check) confirms byte-for-byte.
    • Reload (reboot the same now-minted image, no reformat): boot logs Zuse: cert loaded from block fence; the on-disk seed and pubkey are byte-for-byte identical to the first boot's -- genuinely "mint once, ever," not a silent re-mint.
    • Graceful refusal on a pre-fence volume (disk/artemis-metafence-test.img, meta_fence_blocks=0): blk_meta_zone_read/write both correctly refuse (no space to read or write), so the kernel mints a cert for RAM/this-boot-only use and honestly reports Zuse: minted but fence write FAILED (not persistent) -- no crash, no silent data loss, no corruption of a device with no fence at all.
    • Real disk regression check: the same graceful-refusal path exercised identically against the real, untouched disk/artemis.img (which has no fence yet either) on all three architectures -- clean boot, conservation intact, no panics, disk/artemis.img itself reverted afterward (no committed churn).

    Phase 8's core arc is now functionally complete: real entropy (Phase A) → real signing (Phase B) → real, anti-file, block-native persistence (Phase C) → a working first-boot mint that survives reboots. Still open: the ongoing S" name" MINT word for an authenticated Zuse session to mint additional regular users (needs zuse_cert_devblock_t-format certs signed by Zuse's own key, not just installed) -- the real remaining piece of the original vision.

From FABRIC-2.md §X, Milestone 5 — Console/VM key-match binding

  • Settle the still-open question: reuse ACL-PIN/acl_allow directly, or build a separate key-matching primitive — ACL-PIN gates word execution specifically and nothing today gates console-session-to-VM ownership, so this decision needs to happen before any code gets written here.

  • Design the key/lock data shape (what the console presents, what the VM carries, how they're compared).

  • Wire drive insertion (Milestone 2e's hotplug signal, post-identity-authentication) to a call into capsule_birth_baby() (confirmed a real, callable, on-demand birth path already) to spin up or re-attach that identity's VM. Traced 2026-08-27 (FABRIC-3.md §F.5): this has no mechanism behind it yet — depends on both CERTVERIFY (identity-authentication) and a new RUNCAP mechanism (§F.6, scoped) for per-identity VM content, since capsule_birth_baby() today only births fixed, build-time-baked capsules.

  • Implement the actual attach/bind step — extending sk_repl_set_active_vm() (confirmed to exist, currently an unguarded raw pointer-set) with the key-match check from above, so a console can only bind to the one VM whose lock matches its key. SCOPED 2026-08-27 (FABRIC-3.md §F.9): real gate is USE (mama_forth_words.c:430-480), the live, already-working, currently-unguarded retarget word — check re-verifies the attached drive live on every call; installed=0 VMs stay freely targetable.

  • Implement detach behavior on console disconnect or VM teardown. SCOPED 2026-08-27 (FABRIC-3.md §F.10): a new EJECT word — flush via the existing blk_vm_flush_all(), reset the console's active VM if bound to the departing VM, then capsule_vm_kill() (already complete, no new teardown mechanism needed). Needs one small new piece of state WIREBIND must own: which VMUuid is currently attached via the home-blocks USB path.

From FABRIC-2.md §X, Milestone 6 — Kernel/capsule PKI signing chain

  • Generate (offline, outside the kernel/repo entirely) the real root CA keypair — "stays unrevocable," never embedded, never loaded by any kernel code.

  • Generate the "snakeoil" intermediate certificate, signed by that real root CA (this is a real CA-signed intermediate, not a self-signed/untrusted cert despite the name — "snakeoil" names its informal/private-project status).

  • Embed the already-CA-signed snakeoil intermediate as a capsule blob at build time (mechanically proven already via the font-capsule precedent — no new embedding infrastructure needed, just a new payload). Bootstrapping resolved: no kernel-boot-time verification of a hardcoded CA public key is needed at all — trust is established once, at build time, by whoever holds the real root CA and produces the build.

  • Add a signing step to the mkcapsule build tool (or a separate signing tool) that produces a signature alongside each capsule's existing xxHash64.

  • Extend MANIFEST_AUTO.md's generation to add a signature-status column, matching the existing xxHash64 column's generation pattern.

  • Implement magic-number-based content-type detection (Section U item 14) — a shared primitive, also usable for Milestone 4's foreign-drive check.

  • Root CA + snakeoil intermediate generated 2026-08-26, entirely offline, in a sibling directory outside this repo (/home/rajames/CLionProjects/lithosananke-ca/, not tracked by git here). Ed25519, OpenSSL 3.5.5. Root: self-signed, 20-year validity (20262046), CN=LithosAnanke Root CA. Intermediate: a real CA-signed cert (not self-signed despite the name), 10-year validity, CA:TRUE, pathlen:0 (can sign capsules, can't mint further intermediates), chain verified (openssl verify returns OK). Both private keys chmod 600. The root key never touches this repo or any kernel code, per the design's own requirement.

  • Snakeoil intermediate embedded as a capsule, 2026-08-26. Exported to DER (capsules/pki/snakeoil-intermediate.der, 418 bytes) and dropped under capsules/ — confirmed the font-capsule precedent needed zero new infrastructure: mkcapsule's process_file() embeds any non-.4th file verbatim already. Shows up as capsule pki:snakeoil-intermediate.der in the generated capsule table (38 capsules total, up from 37) — retrieve via capsule_find_by_name() + capsule_get_payload(), never capsule_exec_payload() (it's a passive data blob, not executable capsule code).

  • Minimal DER/X.509 parser written and independently verified, 2026-08-26. New include/starkernel/x509_ed25519.h + src/starkernel/crypto/x509_ed25519.c: x509_extract_ed25519_pubkey(), a from-scratch, narrow DER walker (not a general ASN.1/ X.509 parser, per this milestone's own design decision) — walks Certificate → TBSCertificate → SubjectPublicKeyInfo, handles the optional [0] EXPLICIT Version field (present on v3 certs), verifies the AlgorithmIdentifier OID is exactly 1.3.101.112 (RFC 8410 Ed25519) rather than assuming, and extracts the raw 32-byte key from the trailing BIT STRING. Handles both short-form and long-form DER lengths (a real cert with v3 extensions routinely exceeds the 127-byte short-form limit). Every step bounds-checked against the buffer end — refuses malformed input, never faults.

    Verified against ground truth, not self-consistency: run against the real embedded snakeoil-intermediate.der, the extracted 32-byte key matched openssl pkey -pubin -text's own reported public key byte-for-byte. Refusal path verified too: truncated input, 10 random garbage bytes, an empty file, and a real RSA certificate (algorithm-mismatch case, not just structural malformation) all correctly return failure rather than misreading or crashing. Compiles clean on all three kernel architectures (no __STARKERNEL__ guard needed — same freestanding-safe shape as the other crypto files).

    Still open: the signing step in mkcapsule (needs a new sig[64] field on CapsuleEntry and a parallel emitted array in capsule_generated.c, since CapsuleDesc itself has no spare bytes — confirmed exactly 64 bytes, every field used), wiring ed25519_verify() into the three capsule_validate() call sites in capsule_birth.c (decided: land as WARN-only first, prove correct on all three architectures against both a valid and a deliberately-corrupted capsule, then flip to hard-refuse in a separate step — a bug here has a much larger blast radius than anything else in Phase 8, since a false refusal on Mama's own capsule means no ok> at all, on any architecture), and the signature-status column on capsules/BLOCK_MAP.md (confirmed the real, live manifest target — capsules/MANIFEST_AUTO.md is stale/dead, not regenerated since 2026-07-05, flag as docs drift rather than a real target).

    mkcapsule signing + capsule_birth.c wiring, done 2026-08-26 -- landed WARN-only exactly as decided. First attempt shelled out to openssl pkeyutl -sign (fork/execlp, not system() -- interpolating the key path into a shell string would have been a real, if narrow, injection vector). Corrected on request: no new external host binary dependency when the repo's own code can do the job — this project's standing preference, same theme as the earlier anti-file correction. Rewritten to link ed25519_sign() (src/starkernel/crypto/ed25519.c, already independently verified against OpenSSL in Phase B) directly into mkcapsule. This needed one more small piece: tools/pkcs8_ed25519.c, a narrow DER walker (same shape as x509_ed25519.c, deliberately not shared code — small enough that duplicating a few TLV-walking lines was simpler than threading a header between the kernel crypto tree and host tooling) extracting the raw 32-byte seed from the intermediate's PKCS#8 private key, plus a minimal self-written base64 decoder (PEM is openssl genpkey's default output, not raw DER — checked the repo for an existing decoder first, found none). Verified end-to-end before wiring anything in: the extracted seed, run through this project's own ed25519_keygen(), reproduces the exact pubkey embedded in the cert; a full self-contained sign+verify round-trip (zero openssl involvement) passes.

    CapsuleDesc had no spare bytes (confirmed exactly 64, every field used), so signatures live in a new parallel array, CapsuleSigEntry capsule_signatures[] ({ sig[64], has_sig }, mirroring capsule_names[]'s existing shape), emitted by mkcapsule --sign-key <path> (new optional CLI flag; omitted or the key file missing → every capsule gets has_sig=0, graceful, not a build failure -- CI has no access to the offline key). New Makefile.starkernel variable SIGN_KEY (defaults to this machine's actual key path, ?=-overridable, gated by $(wildcard ...) so a fresh checkout or CI never fails for lacking it).

    New capsule_sig.c/.h: capsule_verify_signature(), deliberately a separate, additive function, not folded into the already-tested capsule_validate() (whose signature and behavior stay untouched). Finds and parses the embedded pki:snakeoil-intermediate.der capsule's pubkey once per boot (cached — every later capsule birth reuses it, the cert can't change mid-boot), then ed25519_verify()s the requested capsule's payload against its stored signature. Wired into all three capsule_validate() call sites in capsule_birth.c (capsule_birth_mama, capsule_birth_baby, capsule_run_experiment) via log_message(LOG_WARN, ...) on anything other than CAPSULE_SIG_OK — never refuses. (Hit one unrelated build snag wiring this in: vm.h and log.h both define LOG_LINE_MAX for two completely unrelated concepts — vm.h's own is the persistent block-log line size, unconditional #define; log.h's is #ifndef-guarded. log.h has to be included after vm.h in any TU that needs both, or the build fails under -Werror; fixed and commented in capsule_birth.c so the next person touching includes there doesn't rediscover it the hard way.)

    Verified independently, both directions, live in the real kernel: a full clean build (38 real capsules, all signed) boots clean on all three architectures with zero signature warnings — the expected outcome when every signature is genuinely valid, not merely an absence of testing. Separately, hand-corrupted one byte of Mama's own init.4th capsule's stored signature (not its payload/hash, which capsule_validate() already checks independently and would have masked the test) directly in the generated capsule_generated.c, without regenerating from source, and rebuilt only the changed object: produced exactly capsule sig: init.4th: INVALID -- signature does not verify on boot, and the kernel still reached ok> — proving warn-only genuinely doesn't refuse anything yet, precisely as decided. Reverted the tampered build before the final acceptance pass. Final, untampered state: clean 3-arch acceptance boot, conservation intact, zero warnings, real disk untouched.

    Still open: flipping WARN to hard-refuse (a separate, deliberate step once this has run clean for a while, not bundled into this pass) and the BLOCK_MAP.md signature-status column.

    BLOCK_MAP.md signature-status column, done 2026-08-26. --manifest mode's file scan (manifest_file()) is a completely separate code path from build mode's process_file() — confirmed it only ever walks .4th files (never the embedded PKI cert or font capsule), and the Makefile invokes mkcapsule twice per build for two different purposes (once to emit capsule_generated.c, once for the manifest). Extended --manifest to accept the same optional --sign-key <path> prefix as build mode (factored the key-loading code, previously inline in build mode's arg parsing, into a shared load_sign_key() used by both) so the manifest can report real per-capsule signing status without touching or requiring a rebuild of capsule_generated.c. New Signed column on the capsule summary table: yes/no when --sign-key was given, n/a (with an explanatory footnote) when it wasn't — never a bare blank that could be misread as "unsigned." Makefile.starkernel's manifest-generation call site updated to pass the same $(SIGN_KEY_ARGS) the real build already uses, so capsules/BLOCK_MAP.md reflects this machine's actual signed state by default. Verified: clean compile, capsules/BLOCK_MAP.md correctly shows yes for all 31 tracked .4th capsules on a real signed build; a quick amd64 boot (no kernel code touched by this step, only host tooling) confirmed no regression, zero signature warnings, conservation intact.

    This closes every open item from FABRIC-2.md's original Milestone 6 punch list except magic-number content-type detection (Section U item 14, shared with Milestone 4, not started) and the hard-refuse flip (deliberately deferred).

    Hard-refuse flip, done 2026-08-26 -- and it landed narrower than the original plan assumed. Before flipping, found that CAPSULE_SIG_MISSING (no signature at all) is the normal state on every machine except this one -- CI and any other checkout have no access to the offline key, by design. Refusing on MISSING the same as INVALID would brick boot everywhere but here, not catch anything real. Decided (on request): enforce ONLY on CAPSULE_SIG_INVALID (a signature that IS present but doesn't verify -- unambiguous tampering/corruption evidence, safe to refuse on regardless of who's building). MISSING/NO_ROOT_KEY stay WARN-only permanently, not as a transitional state.

    All three capsule_birth.c call sites now return CAPSULE_RUN_ERR_INVALID on CAPSULE_SIG_INVALID, after logging the same WARN as before. Verified on all three architectures, both directions, exactly per the original rollout commitment: the positive case (real signed capsules) reboots clean with zero warnings on amd64/aarch64/riscv64, identical to before this flip. The negative case (the same one-byte signature corruption technique used for the WARN-only proof, on Mama's own init.4th) now genuinely refuses -- capsule_birth_mama() returns the error -- on all three architectures identically: capsule sig: init.4th: INVALID followed by Init: Mama birth FAILED. The feared "no ok> at all" blast radius didn't materialize: kernel_main.c already had graceful error handling for a failed Mama birth (log and continue, pre-existing code, not written for this feature) -- the kernel reaches a degraded ok> rather than crashing or hanging, on all three architectures. Final acceptance pass (real, correctly-signed capsules, all tampering reverted) clean on all three, zero warnings, conservation intact, real disk untouched.

    Milestone 6 is now fully closed except magic-number content-type detection (shared with Milestone 4, genuinely separate scope, not started).

From FABRIC-2.md §X, Milestone 7 — Contributor capsules / trust tiers

  • Create the capsules/contrib/ directory (mechanically trivial, matches existing subdirectory convention — the directory itself is not the work).

  • Add a FLAG_CONTRIB bit to mkcapsule.c's flag system, assigned by path match (contrib/ prefix), same pattern as how init.4th already gets FLAG_MAMA_INIT.

  • Decide and implement one of the four spitballed trust-tier directions (signature- authority tiers / block-namespace sandboxing / QEMU-vs-real-hardware conditional enforcement) — none chosen yet, this is a real decision point, not just an implementation task.

  • If block-namespace sandboxing is chosen: extend mkcapsule's existing conflict- detection logic to also reject a contrib/-path capsule claiming blocks outside its reserved range.

From FABRIC-2.md §X, Milestone 8 — Bare-metal boot from physical USB

  • Build a fresh starkernel.iso via make -f Makefile.starkernel ARCH=amd64 clean + the ISO-build step.
  • Identify the exact block device path for the target USB drive on the host doing the flashing (lsblk/dmesg after insertion — care needed, wrong device = data loss).
  • dd if=build/amd64/kernel/starkernel.iso of=/dev/sdX bs=4M status=progress (or equivalent) — confirm dd is the right tool for an El Torito ISO vs. needing isohybrid first (open question, not yet verified).
  • Physically boot the real machine from the flashed drive (BIOS/UEFI boot-order menu, Secure Boot may need disabling — unknown until tried).
  • Capture what happens with no serial-socket log available (real hardware has no qemu-serial-*.sock to socat into) — decide the observation method.
  • Confirm POST reaches the same 1012/0/0 result on real hardware as every QEMU acceptance run.
  • Confirm ok> prompt is reachable and a basic command (e.g. HEARTBEAT-TICKS@ .) works identically to QEMU.
  • Document the result (pass/fail, and if fail, what diverged from QEMU) — first real external validation this project has ever had outside QEMU TCG emulation.

From FABRIC-2.md §X, Milestone 9 — Networking / capsule distribution server

  • (Deferred) Revisit and punch-list this milestone once Milestone 7 closes, not before.

B. Stadium unification — words/VMs/blocks/messages on the same engine

Raised 2026-08-25: "words are stadium patrons, VMs are patrons, blocks are patrons, messages are stadium patrons, all should be operated on by THE SAME ENGINE." Investigated before designing anything — the real state is more nuanced than "everything's a stub," verified via direct reads and git log, not assumed:

FABRIC.md §18.3 already decided the mapping (not invented here): blocks → MIGRATE, messages → DELIVER, ACLs → EXPIRE, words and VMs both → COOL. stadium_evict() (src/starkernel/vm/stadium.c) — real, tested infrastructure: bitmap tracking, pin/contains refusal, the Hera-patron-zero panic guard, heat-conservation back to the owner's reservoir on every reap — calls stadium_dispatch() for the actual payload action when a patron departs.

Per-behaviour status, as of 2026-08-25:

  • MIGRATE (blocks) — zero consumer, genuinely stub (stadium_dispatch()'s case prints "MIGRATE (stub)" and returns). This session already built the real mechanical primitive it needs: blk_subsys_relocate_block()/RELOCATE-BLOCK (FABRIC-2.md, commit 36d832f), live-verified (redirect + content survive an abrupt kill and reboot) but never wired to stadium_dispatch() — it's a separate, parallel, already-working mechanism today, not routed through Stadium at all.
  • COOL (words and VMs, same tag) — half real. Words are fully live, but via a separate, bespoke mechanism, stadium_word_dispatch() (stadium_words.c, item 4.1), wired directly into the real VM word-execution hot path (vm_core.c:690,885,896) — it does not go through the generic stadium_dispatch() switch at all. ONTOLOGY.md §IX claiming words are "not yet migrated" is itself stale documentation drift (same class of bug as the "glibc" misattribution corrected earlier this session — flagged as a small, separate fix below, not blocking). VM cooling has no evidence of ever being wired anywhere — still genuinely stub.
  • DELIVER (Hermes messages)FABRIC.md (~line 3452) records this explicitly as "Open, surfaced not resolved": Hermes's message/channel heat already integrates with Stadium's reservoir accounting (STADIUM-HEAT@, STADIUM-RES-PULL/PUSH), but which Hermes lifecycle event maps to DELIVER vs. EXPIRE was never decided, let alone wired. Real, substantial, Hermes-specific integration work.
  • EXPIRE (ACL TTL expiry) — no evidence of any wiring anywhere; ACL.4th/ acl_recheck() has zero Stadium involvement today. Also substantial, separate work.

Why DELIVER/EXPIRE aren't being resolved in the same pass as MIGRATE: each is a full subsystem integration (Hermes lifecycle mapping; ACL-to-Stadium wiring where none has ever existed) in its own right — attempting all four stubs at once risks exactly the rushed, shipped-but-incomplete outcome the no-stubs rule (below) exists to prevent. MIGRATE gets resolved for real because this session already has a tested primitive underneath it; the other three become honest, explicit punch-list items instead of being touched speculatively.

Punch list:

  • Investigated (2026-08-25): block-patron admission does not exist yet, but is architecturally straightforward, not blocked. Confirmed via stadium_admit()'s own doc (stadium.h) that it REFUSES any candidate with mass != 1, and that a mass > 1 multi-cell patron would need a continuation chain nothing has ever designed — this looked at first like a hard blocker for a 1024-byte block. It isn't: confirmed via stadium_word_dispatch()'s real candidate construction (stadium_words.c:245-252) that Stadium cells carry pure identity/heat/bookkeeping only — candidate.identity = word_id, payload[32] unused — the actual word content stays in the dictionary; Stadium never holds it. By the same pattern, a block patron's cell would carry identity = LBN, mass = 1, payload unused — the actual 1024 bytes of block content stays exactly where it already lives (block cache / disk via block_subsystem.c), unmoved. So mass != 1 is a non-issue; the real gap is just that nothing has ever built the LBN→cell_index residency map (the block-patron analogue of stadium_word_dispatch()'s resolve_resident_cell()) or the touch-on-access hook (the analogue of vm_core.c's three stadium_word_dispatch() call sites). Not yet built — this is real, scoped, buildable work, not a stub-around candidate. Still open, plan to be presented before implementation per the no-stubs/no-early-coding conventions.

  • Resolved (2026-08-25): real block-patron admission + real MIGRATE dispatch, both live. New stadium_blocks.h/stadium_blocks.c mirror stadium_words.c's shape (Option B starter-grant admission, redirected Loop #3 cooling, self-healing stale-entry detection) but key residency by (quota_slot, lbn) in a fixed-capacity open-addressing hash table sized off stadium_cell_count() (tombstone-based deletion, since LBN space isn't densely bounded like word_id), not a dense array. Wired into block_word_block()/block_word_buffer()/ block_word_update() (block_words.c), #ifdef __STARKERNEL__-guarded. stadium_dispatch()'s MIGRATE case now calls blk_flush(lbn) for real (confirmed blk_flush(), not blk_subsys_relocate_block(), is the right primitive — the latter is for compudynamics-driven relocation to a different LBN mid-residency, not ordinary reap write-back). Three new Kconfig tuning constants (STADIUM_BLOCK_HEAT_QUANTUM/STADIUM_BLOCK_COOL_RATE_Q48/ STADIUM_BLOCK_TRACK_CAP_MULT) mirror the word-patron ones exactly, same three-layer wiring. Verified: clean compile, zero warnings, on all three architectures; clean boot to zuse)ok>/ok> REPL on all three, conservation (resident_sum + reservoir == Q48_ONE) intact identically across all three; BLOCK/BUFFER touches exercised live from the REPL on amd64 and riscv64 with no crash; a 22,000-distinct-block flood loop (amd64, artificially shrunk to a 20,971-cell Stadium via a one-off smaller -m to make quota pressure reachable) ran clean under heavy admission-path load with no corruption. Live MIGRATE fire confirmed (2026-08-26): interactive flooding alone never triggered it — Hera's reservoir was already sitting exactly at the Q48_ONE / 3 floor from the boot-time self-tests, so every block-touch candidate pulled 0 heat, and a 0-heat candidate can never be strictly denser than an existing resident, so stadium_admit()'s eviction fallback correctly refuses rather than evicts once the free list is exhausted (a pre-existing reservoir-floor/density-eviction interaction, applies equally to word patrons, not introduced by this pass). Closed deterministically instead with a temporary boot probe (kernel_main.c, inserted into the existing Artemis 4.6 self-test block, reverted immediately after capture — no code left behind): 100 65536 0 STADIUM-ADMIT ... STADIUM-EVICT against the live Artemis VM. Captured live: Stadium: dispatch cell=63257 behaviour=MIGRATE lbn=100, followed by DBG err after MIGRATE probe=0blk_flush(100) fired for real, with the correct LBN threaded through from the departing patron's identity field exactly as designed. (Artemis's own reservoir went briefly out of Q48_ONE-balance during the probe — raw STADIUM-ADMIT doesn't debit the reservoir on its own, by its own doc, so a heat value handed to it directly is invented, not pulled; harmless here since Artemis is killed and her whole economy discarded immediately after, and Hera's own conservation was independently confirmed back at the normal 43691/21845/65536 baseline afterward.)

  • Bug found, reported, then fixed on request (2026-08-25/26): capsules/lib.4th:13-14 shadowed the C primitives USE/RUN. While chasing the live-MIGRATE test above, S" Artemis" USE (meant to redirect the REPL into Artemis's own vocabulary, mama_word_use()) instead printed EXEC: failed: Artemis. Root cause: capsules/lib.4th:13 defined : USE ( addr u -- ) EXEC ; — a FORTH word with the same name but a completely different meaning ("load/exec a capsule"), which shadowed the C-registered USE in dictionary search order since lib.4th loads after primitive registration. lib.4th:14 did the identical thing to RUN, which CLAUDE.md also names as an untouchable C primitive ("BIRTH, RUN, USE are primitives — registered in C exactly like DUP, BYE, EXEC"). Same bug class as the K-PUSH dictionary-shadowing issue (docs/working/architecture/K-PUSH-DICTIONARY-SHADOWING-BUG-20260704.md). Reported first per CLAUDE.md's rule against unrequested fixes; Captain Bob then explicitly asked for the fix. Fix: traced every real caller before touching anything — RUN's alias was dead (never called anywhere as bare RUN); USE's alias had exactly one real caller, capsules/hermes/init.4th:397 (S" common:msg.4th" USE, intentionally exploiting the shadow to load that capsule right after lib.4th itself loaded). Both aliases were pure EXEC wrappers with zero added behavior, so the fix deleted both definitions from lib.4th outright and changed the one real call site plus its matching doc comment (capsules/common/msg.4th:4) to call EXEC directly — no new names invented, the C primitives untouched, mkcapsule --lint clean (31/31 pass). Verified live: rebuilt and booted amd64 — Hermes still births and her COMMON-CH-eviction self-test (which depends on common:msg.4th having loaded) still passes exactly as before; interactively, S" Artemis" USE now correctly prints USE: now using Artemis and switches the REPL's console-name coloring, confirming the C primitive runs unshadowed. Clean compile and clean boot with conservation intact on all three architectures (amd64/aarch64/riscv64).

  • Resolved (2026-08-26): real VM-patron admission + real explicit-KILL eviction, both live. Re-scoped on request: confirmed capsule_vm_kill() had zero Stadium involvement (vm_cleanup()/sf_free() only) and child-VM birth only ever called stadium_grant_quota() (a resource pool for the VM's own future word/block patrons) — never stadium_admit() for the VM itself. The only precedent, stadium_birth_hera(), admits Hera into her own quota as a permanently pinned cell 0, which can never reach stadium_evict() — not a working example of COOL firing for a VM. On closer look this turned out NOT to be entangled with the still-iterating Tripod/Zuse/messaging vision after all (§D) — birth and kill already funnel through two single choke points, so the earlier 2026-08-25 deferral was overcautious. Design: added size_t stadium_patron_cell to VMRegistryEntry (capsule_run.h). At birth, right after the existing stadium_grant_quota() call (capsule_birth.c), admit a candidate into the new VM's own quota mirroring stadium_birth_hera()'s shape (identity=0, mass=1, behaviour=COOL) but deliberately unpinned — pinning would need a new "unpin" primitive (none exists) to ever evict it later, and adding a pin-bypass to stadium_evict()'s refusal logic isn't something to do casually; unpinned costs nothing since nothing wires COOL's dispatch body to actually kill anything, so the worst case of an unrelated natural eviction is stadium_patron_cell going stale, which is tolerated the same way stadium_word_forget() already tolerates staleness. At capsule_vm_kill() and capsule_vm_kill_all_nonmama(): stadium_evict() the tracked cell if still resident, silently tolerating refusal (already gone). stadium_dispatch()'s COOL case needed no new payload body — same as it already is for words, where COOL has no defined extra action beyond stadium_evict()'s own universal reservoir credit; the missing piece was admission and a genuine trigger, not dispatch-body logic. Verified live: a second, new Stadium: dispatch cell=... behaviour=COOL now fires immediately before every PARITY:KILL line, for both Hermes and Artemis, confirmed on amd64 (distinct from the pre-existing COMMON-CH word-eviction self-test's own COOL print). Conservation (resident_sum + reservoir == Q48_ONE) intact throughout. Clean zero-warning compile and clean boot on all three architectures (amd64/aarch64/riscv64).

  • **Re-scoped and resolved (2026-08-26): DELIVER (Hermes) was never actually a gap — the earlier "zero consumer, needs substantial Hermes lifecycle mapping" framing above was wrong, carried over unverified from FABRIC.md's old "open, not resolved" note about which Hermes event maps to DELIVER vs. EXPIRE. Item 4.2 already answered that in code (messages → DELIVER, channels → COOL) without the prose ever catching up — same documentation-drift class as the stale ONTOLOGY.md words note and the earlier glibc misattribution. Confirmed live: capsules/hermes/init.4th's MSG-ALLOC already admits every message with SB-DELIVER, and MSG-FREE-NODE (called from both MSG-ACK-LAST and heat-driven MSG-REAP) already evicts it — behaviour=DELIVER (stub) has been printing on boot logs since at least 2026-08-05. Checked whether the dispatch body needed a real payload action the way MIGRATE did: MSG-DELIVER (the FORTH word) already runs the actual delivery (VM-EXEC of the payload) before eviction, decoupled from Stadium reap — so by dispatch time delivery is already done, same shape as COOL, which needs no extra action beyond stadium_evict()'s own universal reservoir credit. Fix: stadium_dispatch()'s DELIVER case now prints the departing message's real identity (DELIVER msg_idx=N, same shape as MIGRATE's lbn= print) instead of a misleading (stub) label — confirmed live via a forced MSG-SEND/MSG-DELIVER-ALL/MSG-ACK-LAST sequence from Hermes's own REPL context (Stadium: dispatch cell=73653 behaviour=DELIVER msg_idx=1). COOL's case was in the identical situation (real for both words and VMs, no extra action needed) and, on request, got the same fix (2026-08-26): now prints COOL identity=N (word_id for a word, 0 — the patron-zero convention — for a VM) instead of (stub). Confirmed live: both shapes fired correctly on the same boot — COOL identity=0 at Hermes's/Artemis's own explicit channel-eviction self-test and again at their VM-patron eviction at PARITY:KILL, COOL identity=1 at a second channel eviction — conservation intact throughout. Clean zero-warning compile and clean boot on all three architectures for both fixes.

  • EXPIRE (ACL) — confirmed genuinely unscoped (2026-08-26), not a case of stale documentation like DELIVER turned out to be. Two findings below, then four open questions — decided directly on request (2026-08-26), decision recorded after the questions, no code written (a decision isn't a green light to build, per this session's own convention):

    Finding 1 — acl_ttl and Stadium's ttl are different things wearing the same name. DictEntry.acl_ttl (vm_core.c:756) is a per-word countdown that batches how often ACL-RECHECK runs — when it hits 0, acl_recheck() calls the FORTH word ACL-RECHECK (ACL.4th:50), which always renews: STRICT mode sets allow=1, ttl=0 (recheck every time); TTL mode computes a fresh heat-based TTL and sets allow=1. Denial isn't a live path anywhere in current policy. This is a renewal cycle, not a residency-ending event — nothing about it resembles "leaving the Stadium floor."

    Finding 2 — StadiumPatronHeader.ttl is completely inert. Every candidate constructor across the whole codebase (stadium.c, stadium_words.c, stadium_blocks.c, Hermes's MSG-ALLOC/CH-ALLOC) sets candidate.ttl = 0, and nothing anywhere ever reads, decrements, or reaps on it. The generic TTL-expiry mechanism EXPIRE would need to fire from doesn't exist in Stadium's own engine at all — a gap one level deeper than "ACL isn't wired to Stadium."

    Open questions:

    1. Is "ACL patron" even the right model, or was FABRIC.md's original "ACLs → EXPIRE" mapping a category mismatch from the start — conflating acl_ttl's recheck-amortization counter with Stadium's residency ttl?
    2. If it is the right model: what gets admitted as a patron? One per ACL-guarded word would be redundant with the word's own patron cell item 4.1 already tracks. A different unit (e.g. per zuse session) might fit better once PKI/session auth lands (Phase 8, still open per .claude/CLAUDE.md's ACL section).
    3. Building this for real means building Stadium's generic ttl-decrement/reap-on-zero mechanism first — nothing to hook EXPIRE into today. Is that in scope here, or its own separate item?
    4. What should the reap action actually do, given current ACL policy never revokes — would EXPIRE force an ACL-RECHECK, or something else entirely?

    Decision:

    1. Not at the per-word level. acl_ttl-hits-zero always renews, never revokes — forcing EXPIRE's residency-ending tag onto it would misuse the tag. But "ACLs → EXPIRE" isn't wrong in spirit, just aimed at the wrong unit: the one place in this system where something ACL-related genuinely has a lifetime and should be revoked is a zuse superuser session (Phase 8, not yet built) — authenticate, hold elevated privilege for a bounded time, then actually drop back to non-zuse. That's a real residency-ending event; per-word recheck isn't.
    2. A zuse session, not a per-word ACL entry — a session is the thing with a genuine start/lifetime/end. A word already has its own patron via item 4.1; a second one for ACL purposes would be redundant bookkeeping, not a new concept.
    3. Not in scope now. This is the decision that actually resolves the item: Phase 8 doesn't exist yet, so there is no session to admit as a patron regardless of any other choice made here. Building the generic ttl-decrement/reap mechanism now, with nothing real to feed it, would be speculative infrastructure ahead of its only consumer — close in spirit to what the no-stubs standard exists to prevent, just inverted (a real mechanism with no real caller, instead of a fake mechanism with a real caller).
    4. Revoke the session's elevated privilege and drop the console back to its non-zuse state — a real action, unlike DELIVER/COOL, which needed none.

    EXPIRE stays explicitly deferred until Phase 8 (PKI/zuse session minting) lands — not abandoned, not left ambiguous: revisit it as part of that work, admitting the session itself (not a word) as the patron, once there is something real for it to represent.

  • Fixed (2026-08-26): ONTOLOGY.md §IX's "words (dictionary, warehouse-resident today, not yet migrated)" line was stale — corrected to state words are fully migrated and live via stadium_word_dispatch(). Doc-only, no build/boot verification needed.


C. Standing rule: no stubs or TODOs, ever

Stated directly, 2026-08-25, after the stadium_dispatch() stub investigation above: "I've never allowed stubs before." Saved as a persistent memory (feedback_no_stubs_or_todos.md) so this applies across sessions, not just this one. Full statement: no stub function that prints a placeholder and returns, no TODO-and-move-on comment in place of real logic, in any language, ever committed as if it were finished work. Small, honest increments are still fine and encouraged — each increment just has to be a complete, real implementation of whatever slice it covers, never a placeholder for a later slice. A pre-existing stub found while working nearby (as here) gets flagged and resolved, not built on top of or left in place.


D. Tripod final shape — minting, one-time Zuse fuse, messaging-only (vision, substantially expanded 2026-08-27)

Stated directly by Captain Bob, 2026-08-25: "we're going to have to iterate because I know what the final shape of the tripod will be." Originally captured as a forward-looking vision, not yet broken into implementable items. 2026-08-27: expanded into a much fuller picture after Phase 8/Milestone 6 closed and unblocked this design pass — captured here in full before any of it gets scoped into implementable items, per direct request ("capture EVERYTHING first then we'll build a plan"). Full detail also in memory project_tripod_final_shape_vision.md.

D.1 — Original vision (2026-08-25), still accurate

  • Thumbdrive presentation → legality check → mint. A newly presented thumbdrive is checked for legality (against the CA-root-derived identity/certificate scheme). If not legal, it is "minted" — formatted for system use — which (1) spins a new user VM and (2) attaches the console VM to it.
  • One-time Zuse mint + fuse-blow on first install. A brand-new system instance ("Install"/"Try It") mints exactly one Zuse superuser, then irreversibly "blows a fuse": direct quote, "we mint one and only one Zuse user and blow a fuse. The only way around is a new system." Post-fuse, no further Zuse can ever be minted on that instance — but the system is NOT bricked: the existing Zuse superuser keeps working, and ordinary users can still "thumb in" via regular thumbdrives. This half is now actually built — Phase 8's entropy→signing→block-fence persistence work (2026-08-26, project_phase8_milestone6_complete.md) implements exactly this, minus the physical thumbdrive itself (see D.4, WRITE(10) gap).
  • Messaging-only once Tripod is fully live. All inter-VM interaction becomes Hermes messaging, not direct calls/shared state — a stated end-state, not the current implementation. Reaffirmed and sharpened 2026-08-27: hardwired/direct-call implementations are acceptable during development only — no piece of this design is considered actually finished until its hardwired calls are replaced with real Hermes messages. This is a standing completion criterion for everything in this section, not a separate future milestone.
  • Polymorphic block-boundary behavior for user VMs. Still just one sentence, still not elaborated — genuinely unresolved since 2026-08-25, needs its own follow-up conversation. CLOSED 2026-08-27, see D.4/§F.16 — that follow-up happened; already true of the live block-subsystem code.
  • ClaudeEXPORT/ — a prior Claude data export, "concepts and thoughts as guidelines," explicitly flagged as possibly containing superseded/conflicting ideas, not authoritative. See memory project_claude_export_archive.md.

D.2 — The session model (2026-08-27): a session IS a VM

Direct correction to how this document and the earlier EXPIRE framing (§B) had been thinking about "sessions": a session is any VM client running in the Stadium fabric — Zuse is no different from any other VM. The Stadium/FABRIC metaphor restated by Captain Bob: the fabric floats inside the Stadium; VMs actually doing work are "players on the field" (system perturbances); user interaction is "fans in the stands" providing patron-driven noise. Zuse is a player like any other, not a special boolean flag on Hera.

This directly invalidates the EXPIRE-as-TTL-on-zuse_session-boolean plan this session had started scoping (§B's EXPIRE decision, and the implementation research begun 2026-08-27 before this correction landed). Today's zuse_session field on the VM struct is not what a "session" actually is in this model — a session ending is a VM detaching, far closer to the existing, real, working COOL/capsule_vm_kill() path than to a new Stadium-behaviour-tag TTL sweep. EXPIRE as originally scoped is superseded, not merely deferred — when Stadium/ACL session-lifetime work is picked up again, it should be framed as "what happens when a user-VM's owning thumbdrive detaches," reusing COOL, not as a timeout on a flag.

When Hera has no VM attached beyond fleet-servicing strata, she idles, watching services and maintaining a serviced fleet foundation — her normal steady state, not a special case.

Validated against the original Stadium foundation, not just today's intuition (mined 2026-08-27 from FABRIC.md §17.1/§20, the oldest, most foundational layer of this whole design — see D.2b below for the full mining pass). FABRIC.md §17.1's own patron/departure table already says VMs depart via heat decay (COOL) and only messages/ACLs ever used TTL — meaning the EXPIRE-as-TTL-on-a-flag plan this session had started scoping didn't just become inconvenient once "session = VM" landed, it contradicted the original design from the start: a VM was never one of the two patron kinds TTL governs. §20's outer/inner Stadium nesting (VMs as outer-Stadium patrons; words/blocks/ACLs/messages as inner-per-VM patrons) already anticipated exactly "attach = admit a VM patron" as the outer-Stadium case — no new Stadium theory is needed for this, only a new admission trigger (thumbdrive attach) feeding the existing admit/evict machinery.

D.2b — The pentagon: five VMs, minimum, to service one attached user (2026-08-27)

Stated directly: "Just to get a user, 5 VMs are needed — if we draw every connection to every node they can make a pentagon." The five: Hera, Hermes, Artemis, the user's own VM, and the Console. Drawing every pairwise connection among five nodes (a complete graph, K5) draws a pentagon with all its diagonals — visually a pentagram. Noted in passing, and worth keeping since this project's naming is already steeped in Greek mythology (Hera, Hermes, Artemis, Ananke/Lithos, Zuse for Zeus): a pentagram at the foundation of a "pagan" system fits the theme, not an accident to explain away.

The Console is a full peer node, not a passive terminal that merely "attaches to" a user VM — D.3 step 4 undersells this; the Console is architecturally one of the five, with (at minimum) its own binding relationship to the user's VM and, per the messaging end-state, eventually its own message traffic to Hera/Hermes/Artemis too, not just a one-way relay. Reconciling D.3's "console attaches to the new user VM" phrasing against this five-node, fully-connected picture is real remaining work, not done here — captured as a gap in D.4.

Why each of the five is load-bearing, best understanding as of this capture (not all edges independently confirmed — see D.4's gaps):

  • Hera — governs the fleet, births/kills the user VM (existing capsule_birth_baby()/ capsule_vm_kill() path), holds the outer-Stadium quota the new VM draws from.
  • Hermes — the messaging fabric itself; per the standing "nothing is done until it's messaging" criterion, every other edge in this pentagon is ultimately meant to route through Hermes, making Hermes both a node in the graph and (eventually) the medium connecting every other edge — worth being explicit that this is a two-level relationship (Hermes as a peer VM and as the substrate other peers communicate over), not a design tension to smooth over by treating it as an ordinary fifth node.
  • Artemis — persistent block storage; the identity/system-metadata store (Phase 8's block-fence) and, per D.3, wherever a user's own working data actually lives during a session, live here.
  • The user's own VM — where the user plays (D.3 step 5).
  • Console — the human-facing binding point (D.3 step 4, Milestone 5).

D.3 — The attach/mint flow, end to end

  1. Idle loop watches for thumbdrive connects (sk_repl_idle(), already the right hook — confirmed 2026-08-27 during the (superseded) EXPIRE research: ~1s cadence, already used for xhci_poll_events()/MSC attach).
  2. Cert-only authentication — no password, no username. A presented thumbdrive's identity is entirely the cert on it. Verification means checking the cert's fields directly (chain of trust, most likely signed by Zuse's own key — the same x509_ed25519.c/ ed25519_verify() primitives Milestone 6 already built and proved work end to end, just verifying against Zuse's pubkey instead of the snakeoil intermediate's). The system has no requirement to maintain any single identity itself — no user directory, no central registry of who exists. Each thumbdrive is a fully self-contained, portable identity record.
  3. A dedicated user VM is born from the identity's own data — "the users/agents/services metadata can reside in their capsule and [get] encoded directly into their dictionary when they attach and their init.4th executes." This is a runtime-constructed capsule from drive content, not a build-time-embedded one like pki:snakeoil-intermediate.der — a materially different mechanism from anything mkcapsule does today (gap, see D.4).
  4. The console VM attaches to the new user VM — this is exactly Milestone 5's still-open "Console/VM key-match binding" item, now with its open question resolved: "the bumps on a thumbdrive have to fit the holes on the VM... (ACL territory)" — reuse the existing ACL primitive (acl_mode/acl_allow/acl_pinned) for the key-match check, don't build a separate mechanism. Exact comparison semantics still undesigned (see D.4).
  5. The user plays entirely inside their own VM — all their FORTH activity happens in their own isolated VM context, never Hera's or anyone else's.
  6. Ongoing minting: while Zuse is attached/active, a second blank thumb presented can be minted with a new named identity (worked example given: S" CaptBob" MINT-shaped flow) and ejected — this is the still-open MINT word from Phase 8's own punch list, now with a concrete attach-flow context to design it against.
  7. Detach: working state is copied back to the user's own thumb and flushed before eject. Explicitly deferred — stated directly as "going to be a maintenance operation later," not needed now.

D.4 — Gaps identified 2026-08-27, sorted by whether they're answered

Answered (resolved during capture, not deferred):

  • Milestone 5's "reuse ACL-PIN vs. build a new key-match primitive" question — reuse ACL. Now grounded, not just intuited (mined 2026-08-27 from FABRIC-2.md §C/§D, lines ~2884-2918): acl_pinned's one-way-ratchet (include/vm.h:346, C-only mutation, no FORTH clear path) was already identified as the right-shaped precedent for exactly two things — Zuse's own one-way "burn" and console-session ownership — but in both cases only the pattern was precedented, never actually built past word-execution gating. The "still genuinely open" framing in FABRIC-2.md is accurate as of that writing; today's decision closes it, the comparison semantics remain the real gap (below).
  • EXPIRE's framing — superseded by "session end = VM detach," see D.2. Now shown to be required, not just cleaner, by FABRIC.md §17.1's own patron/departure table (see D.2's addendum above) — TTL was never a valid mechanism for a VM patron in the first place.
  • Regular-user cert chain-of-trust mechanism — signed by Zuse's own key, verified with already-built, already-proven Ed25519/X.509 primitives (Milestone 6). No new crypto work needed, only a new verification call site using existing functions. Correction, 2026-08-27 (FABRIC-3.md §F.7): partially wrong. x509_extract_ed25519_pubkey() deliberately stops at SubjectPublicKeyInfo — no signature extraction, no TBS-byte-range capture, no chain/issuer check of any kind (x509_ed25519.h's own doc comment). Verifying a signed cert (not just reading a pubkey out of one) is new DER-walking work, not a pure reuse. Trust root is also confirmed separate from the capsule-PKI chain: Zuse's own on-device key (zuse_cert_devblock_t), never the offline root CA / snakeoil intermediate.
  • Whether the system needs a central user directory — no, by design; each thumbdrive is self-contained (reinforces D.3 point 2, was implicit, now explicit).
  • The concrete target shape for message-bus migration (mined 2026-08-27 from FABRIC-2.md, cross-checked against live capsules/hermes/init.4th): a message is MSG-CELLS = 9 cells (TYPE, FROM, TO, PADDR, PLEN, STADIUM-CELL, SEQ, CH, ORIG-TYPE), 72 bytes, payload out-of-line (a pointer+length to caller-owned bytes — there was never a fixed inline payload field, an earlier open question in FABRIC-2.md §23.4 that its own design already closed as moot). Live API: MSG-ALLOC ( heat -- addr|0 ), CH-ALLOC ( heat -- addr|0 ), MSG-DELIVER ( m -- ), MSG-DELIVER-ALL ( -- ). A future "thumbdrive attach event" or "cert verify request" message type follows this exact shape — allocate, fill TYPE/FROM/TO/CH, point PADDR/PLEN at an out-of-line event struct, send. Not built, but no longer an abstract "use messaging someday" — there's a real API to target.

Real, unanswered gaps — deferred on purpose (per "answer what we can, defer what we can't"):

  • How does thumbdrive content actually become a capsule at runtime? mkcapsule only ever builds capsules at build time, baked into capsule_generated.c — nothing today constructs a capsule from data read off a live-attached device. This is new mechanism, not a variation on anything that exists. SCOPED 2026-08-27, FABRIC-3.md §F.6: no new birth machinery needed — capsule_birth_baby() is already generic; construct a self-contained heap-allocated single-entry directory (mirroring kernel_main.c's own heap-copy of the compile-time directory) and feed it straight in. Source-location-on-drive and read-path details still open.
  • Exact ACL "bumps and holes" comparison semantics. Which specific ACL fields, compared how, constitute a console-to-VM match? Direction is settled (reuse ACL), the actual comparison logic is not designed.
  • Message-bus migration scope. Which specific interactions (attach detection, cert verify, console bind) move to Hermes messages, and what do those message shapes look like? Entirely unscoped, explicitly deferred until after a hardwired version exists to migrate from. SCOPED 2026-08-27 (FABRIC-3.md §F.15): MSG-DELIVER already executes arbitrary FORTH text on the target VM (VM-EXEC), so no new dispatch mechanism is needed. Real target is a genuine Console VM (already real: serial+framebuffer+PS2), two-hop flow (Hera→Console reports attach; Console→Hera requests the privileged operation with literal arguments). Surfaced a substantial new vision detail along the way: blank-media minting is a Console- driven interactive onboarding form, not a bare word call (§D.6).
  • Polymorphic block-boundary behavior for user VMs — still just the one original sentence from 2026-08-25, never elaborated. CLOSED 2026-08-27 (FABRIC-3.md §F.16): clarified by the user as an architectural principle (the block layer should have no idea what backs it), then confirmed already true of blk_get_buffer()'s live RAM/raw/disk-cache abstraction — not a missing feature.
  • SSD identity-store scope — does the system-resident store (the block-fence built in Phase 8) ever hold anything about regular users, or is it purely Zuse/system bookkeeping with regular users staying purely thumbdrive-resident? Best read of what was said: the latter, but not stated explicitly enough to treat as settled. CLOSED 2026-08-27 (FABRIC-3.md §F.16): confirmed the latter, both by D.3's own explicit "no central registry" statement and structurally by where this session's own scoping put every piece of user identity data (§F.6-F.8, all on the user's own drive).
  • Does a user's VM/session state round-trip across multiple attaches (unplug, come back later, resume exactly where you left off), given detach explicitly copies state back to the drive? Implied yes by the detach description, not confirmed directly. CLOSED 2026-08-27 (FABRIC-3.md §F.16): confirmed as a structural consequence of decisions already made — the user's pool being their own drive plus EJECT's flush-before-kill means the data simply never leaves; no separate resume mechanism needed. Checked against FABRIC.md §24's mutation/identity-stability rules (mined 2026-08-27) — no conflict either way: identity is only pinned stable within one residency, so a thumbdrive re-attaching later and being admitted as a fresh VM patron (possibly reusing an identity value derived from its cert/pubkey) is simply a new residency each time. Nothing in the foundational theory blocks or answers round-tripping; it stays a pure implementation decision.
  • Reconcile D.2b's five-node, fully-connected pentagon against D.3 step 4's "console attaches to the user VM" phrasing. The Console is a full peer node (Hera/Hermes/Artemis/ User-VM/Console, K5), not just a passive relay that binds to one other node — D.3's wording undersells this. Which of the ten pentagon edges are real today versus aspirational once messaging lands is not worked out; only Hera↔UserVM (birth/kill) and UserVM↔Artemis-via-Hera's-quota (existing Stadium quota mechanics) are clearly already real in some form.
  • Confirmed baseline before any message-bus migration can start (mined 2026-08-27, FABRIC-2.md cross-checked against live src/starkernel/repl.c and the xHCI/USB driver): today's hotplug→MSC-attach→homeblocks_sig_check()blk_subsys_attach_device() chain is 100% hardwired, direct C calls, zero messaging at any point — confirmed empirically, not assumed. This is the exact, real starting point "hardwire now, message later" describes; there is no partial messaging implementation anywhere in this path to build on yet.
  • The single biggest hard blocker underneath all of it: SCSI WRITE(10) still doesn't exist (blkio_usb.c is read-only — Milestone 2's biggest open item, unchanged). Minting a blank thumb, and flushing state back on detach, both need real writes to physical USB media. None of D.3's flow can run on real hardware until this lands; QEMU testing would need to proxy through a writable virtio-blk stand-in the same way Phase 8's cert persistence did for Zuse's own identity.

Small, unrelated, real bug found along the way: the prompt shows zuse)ok> (a stray trailing paren, "never corrected typo") — should be (Zuse)ok>, both parens around the identity name, generalizing to any attached identity (e.g. (CaptBob)ok>). Independent of everything else in this section; safe to fix any time. CLOSED 2026-08-27 (FABRIC-3.md §F.16): traced for the literal string — it doesn't exist in any live code path today (SK_PROMPT_TEXT is unconditionally "ok> "); the typo lives only in a doc comment describing not-yet-implemented behavior. Correct intended form recorded for whenever it's actually built.

Why this affects §B: VM-COOL (and any inter-VM Stadium wiring) sits close to this design — see D.2's EXPIRE supersession above, the concrete example of exactly this risk materializing.

Next step: no implementation yet — this section is capture only, per direct request. Next session should turn D.3/D.4 into an actual scoped, ordered plan (the user's own framing: "identify gaps, answer what we can, defer what we can't... push the pile out in front of us").


E. The pile as a graph (2026-08-27)

Stated directly: "if this is all documented, scoped and correctly interwoven we can begin iterating... I see no reason to not create this as a graph, it might better handle the way I think." Auditing D.1-D.4/D.2b against the older Milestone 2/3/5 punch lists (§X in this document) confirms the instinct is right, not just stylistic — several items across different milestones are, on inspection, literally the same open question, and one single blocker (WRITE(10)) gates work in at least four unrelated-looking places. A flat punch list hides that; a graph doesn't.

graph TD
    W10["✅ WRITE(10) SCSI support<br/>BUILT + VERIFIED 2026-08-28 (§F.17), all 3 arches"]
    HOTPLUG["✅ Hotplug event surfacing<br/>CLOSED 2026-08-27, M2 (§F.3)"]
    STALL["❌ USB error/stall recovery (M2)<br/>SCOPED 2026-08-27 (§F.14) — full BOT-spec recovery designed"]

    M6["✅ Milestone 6 — capsule PKI<br/>DONE 2026-08-26"]
    PH8["✅ Phase 8 — Zuse identity<br/>+ block-fence, DONE 2026-08-26"]

    CERTVERIFY["❌ Zuse-signed user-cert verification (M3)<br/>SCOPED 2026-08-27 (§F.7) — X.509/DER, own trust root"]
    FIRSTTOUCH["❌ First-touch identity→block-range<br/>allocation (M3) — SCOPED 2026-08-27 (§F.11)"]
    BMAPFMT["❌ On-drive block-map format (M3)<br/>SCOPED 2026-08-27 (§F.4) — repurpose blk_meta_t"]
    BMAPWRITE["✅ Write block-map to drive (M3)<br/>CLOSED 2026-08-27 (§F.12) — already built, blocked only on W10"]
    BMAPREAD["✅ Read/validate block-map on insert (M3)<br/>CLOSED 2026-08-27 (§F.13) — already built"]
    MIGSM["🟡 Migration state machine (M3)<br/>partially answered: ACL decides *when*"]
    UNCLEAN["❌ Unclean-removal handling (M3)<br/>DESIGN CLOSED 2026-08-27 (§F.10), code not yet written"]

    ACLKEY["❌ ACL 'bumps and holes' comparison<br/>semantics — SCOPED 2026-08-27 (§F.2)"]
    WIREBIND["❌ Wire hotplug → capsule_birth_baby() (M5)"]
    BINDSTEP["❌ Extend sk_repl_set_active_vm()<br/>with key-match (M5) — SCOPED 2026-08-27 (§F.9)"]
    DETACH["❌ Detach behavior / new EJECT word (M5)<br/>SCOPED 2026-08-27 (§F.10)"]

    EXPIRE["✅ EXPIRE reframed: session end<br/>= VM detach via COOL (§B + D.2)"]
    RUNCAP["❌ Runtime capsule construction<br/>from thumbdrive content — SCOPED 2026-08-27 (§F.6)"]
    MINT["❌ Ongoing MINT word (Phase 8 + D.3)<br/>SCOPED 2026-08-27 (§F.8) — GPT dropped, single-device confirmed OK"]
    PENTAGON["📍 Pentagon: Hera/Hermes/Artemis/<br/>User-VM/Console, K5 (D.2b)"]
    MSGSHAPE["✅ Hermes message shape known<br/>(MSG-CELLS, MSG-ALLOC/DELIVER)"]
    MSGMIGRATE["❌ Message-bus migration of<br/>attach/verify/bind (D.4) — SCOPED 2026-08-27 (§F.15)"]

    SSDSCOPE["✅ SSD identity-store scope<br/>for regular users — CLOSED 2026-08-27 (§F.16)"]
    ROUNDTRIP["✅ Session state round-trip<br/>across attaches — CLOSED 2026-08-27 (§F.16)"]
    POLYBLOCK["✅ Polymorphic block-boundary<br/>behavior — CLOSED 2026-08-27 (§F.16), already true"]
    PROMPTBUG["🐛 (Zuse)ok> prompt typo<br/>CLOSED 2026-08-27 (§F.16) — doc-comment only, not live code"]
    EMPTYBUFBUG["🐛 EMPTY-BUFFERS mass-zeros every<br/>block on every device (§F.17) — found 2026-08-28, NOT FIXED"]

    W10 -->|unblocked, §F.17| BMAPWRITE
    W10 -->|unblocked, §F.17| MINT
    W10 -->|unblocked, §F.17| DETACH
    W10 -.->|QEMU-verified; real hardware still untested| M8["Milestone 8 — bare metal"]

    M6 -->|unblocks| CERTVERIFY
    PH8 -->|supplies crypto primitives| CERTVERIFY
    PH8 -->|supplies mint-then-pin precedent| MINT

    BMAPFMT --> BMAPWRITE
    BMAPFMT --> BMAPREAD
    BMAPFMT --> FIRSTTOUCH
    BMAPREAD -.->|found: sig must relocate off devblock 0, §F.13| MINT

    CERTVERIFY --> RUNCAP
    RUNCAP --> MINT

    HOTPLUG --> WIREBIND
    CERTVERIFY -->|"post-identity-authentication" per M5 wording, SCOPED 2026-08-27 §F.5| WIREBIND
    RUNCAP -->|no per-identity VM content until this exists, SCOPED 2026-08-27 §F.5| WIREBIND
    ACLKEY --> BINDSTEP
    WIREBIND --> BINDSTEP
    BINDSTEP --> DETACH
    UNCLEAN -.->|same kill path, no flush| DETACH
    DETACH -->|triggers| EXPIRE

    MSGSHAPE --> MSGMIGRATE
    WIREBIND -.->|hardwired version must exist first| MSGMIGRATE
    BINDSTEP -.->|hardwired version must exist first| MSGMIGRATE
    CERTVERIFY -.->|hardwired version must exist first| MSGMIGRATE
    PENTAGON -->|which of the 10 edges get messaged| MSGMIGRATE

    classDef done fill:#1a3,stroke:#333,color:#fff
    classDef blocked fill:#a33,stroke:#333,color:#fff
    classDef unblocked fill:#a83,stroke:#333,color:#fff
    classDef open fill:#666,stroke:#333,color:#fff
    classDef partial fill:#883,stroke:#333,color:#fff
    class M6,PH8,EXPIRE,MSGSHAPE,HOTPLUG,BMAPWRITE,BMAPREAD,SSDSCOPE,ROUNDTRIP,POLYBLOCK,W10 done
    class STALL,FIRSTTOUCH,WIREBIND,BINDSTEP,DETACH,MINT,BMAPFMT,CERTVERIFY,RUNCAP,UNCLEAN,MSGMIGRATE,ACLKEY blocked
    class MIGSM partial

Legend: done · 🔓 unblocked-but-not-built · not started, no open design question left (just needs building) · real open design question, not just missing implementation · 🟡 partially addressed · 📍 topology/context node, not a work item itself · 🐛 trivial, independent bug. Solid arrows = hard dependency (A must exist before B can be built or finished). Dashed arrows = softer "gates/informs" relationships.

What the graph makes visible that the flat lists didn't:

  • WRITE(10) is a true hub, not just "the M2 punch list's biggest item" — it independently gates the block-map write path, MINT, detach/flush-back, and all real-hardware testing. Landing it once unblocks four separate-looking fronts at once.
  • ACLKEY (the exact ACL comparison semantics) is one node, not two — Milestone 5's "key and lock data shape" and D.4's "bumps and holes" gap are the literal same open question, asked in two different sessions/documents. Answering it once closes both.
  • RUNCAP (runtime capsule construction) sits on the critical path to MINT — a gap that read as "D.4 leftover detail" is actually load-bearing for the ongoing-minting flow, not a side note.
  • MSGMIGRATE has three separate hardwired prerequisites, all currently unbuilt — the message-bus migration can't even start until WIREBIND/BINDSTEP/CERTVERIFY exist to migrate from, confirming this is correctly the last thing to tackle in this whole area, not a parallel track.
  • Only SSDSCOPE, ROUNDTRIP, POLYBLOCK, and PROMPTBUG are genuinely standalone — everything else in the pile connects to at least one other node.
  • WIREBIND was missing two real incoming edges (found scoping it, §F.5): CERTVERIFY and RUNCAP both gate it, not just HOTPLUG. Without a verified identity there's nothing to bind to, and without RUNCAP there's no per-identity VM content to birth — today's capsule_birth_baby() only runs fixed, build-time-baked capsules by name, so "spin up that identity's VM" had no real mechanism behind it at all until this pass traced it.
  • CERTVERIFY and the capsule-PKI chain are two separate trust roots, not one (§F.7) — easy to conflate since both use Ed25519/X.509, but user certs are signed by Zuse's own on-device key while capsules are signed by the offline root CA's snakeoil intermediate. Neither chain validates the other; a future audit pass should keep this distinction explicit rather than let "PKI" get treated as one undifferentiated mechanism.
  • MINT closed out the GPT question for the whole arc, not just itself (§F.8) — GPT was never actually load-bearing anywhere in this design; every node that referenced it (CERTVERIFY, WIREBIND, RUNCAP, HOTPLUG) was already treating raw devblock-0 as the working format. Dropping GPT converts all of those from "interim, revisit later" to simply "the design," retroactively.
  • BINDSTEP already has a live target to gate, not new plumbing (§F.9) — USE (mama_forth_words.c) is a real, working, currently-unguarded retarget word. The only genuinely new code this node needs is a small cross-file accessor for the currently-attached device (mirroring sk_repl_get_active_vm()'s own precedent) plus the comparison itself.
  • UNCLEAN (M3) and DETACH (M5) were the same node wearing two names, in a different way than ACLKEY was (§F.10) — not identical questions asked twice, but one shared teardown path (capsule_vm_kill(), already complete) forked by whether a flush is physically possible. Scoping them together closed a long-open M3 item as a side effect of an M5 pass.
  • Found a real, pre-existing, independent bug while tracing DETACH (§F.10): plain KILL today never resets g_repl_active_vm, so killing the VM a console is currently USE'd onto leaves a dangling pointer. Reported, not fixed — the new EJECT/hot-unplug call sites handle this correctly, but the existing KILL word's own call sites do not.
  • FIRSTTOUCH's own punch-list wording was stale (§F.11) — "claim at g.total_user_lbn" predates BMAPFMT's distributed-ownership decision and doesn't describe a workable mechanism against the current design. A second, unrelated discovery in the same pass — blk_meta_t's existing chain fields — turned what looked like a fragmentation problem into a non-issue for free.
  • BMAPWRITE was already done (§F.12) — BMAPFMT's decision to repurpose the existing blk_meta_t accessors instead of a new table meant "write the map" and "write a block" were never actually two operations. The only reason this hasn't been seen working is WRITE(10) itself, already the graph's known hub.
  • BMAPREAD closed the same way, but tracing it surfaced a real, load-bearing bug that had nothing to do with reading at all (§F.13) — homeblocks_sig_t and the block-subsystem's own generic disk header physically collide at devblock 0. Every earlier node that touched homeblocks_sig_t (CERTVERIFY, WIREBIND, RUNCAP, MINT) inherited this without anyone noticing, because each only traced homeblocks_sig_check() itself, never the ordinary attach path running alongside it. Fixed now (relocate to devblock 1) rather than left to surface later as a real, confusing runtime failure.
  • STALL closes out Milestone 2 entirely (§F.14) — the last item on that punch list. Unlike most nodes this session, this one designs genuinely new protocol machinery rather than finding existing infrastructure already covers it: no completion-code distinction and no recovery of any kind existed before this pass, only a bounded-timeout safety net.
  • MSGMIGRATE needed no new mechanism, only a routing decision (§F.15) — MSG-DELIVER already executes arbitrary FORTH text on the target VM. Scoping it also surfaced a real, substantial vision expansion (§D.6): blank-media minting is meant to be an interactive, Console-driven onboarding form, connecting forward into both RUNCAP's deferred "default personality content" question and MINT's own scope.
  • WRITE(10) (§F.17), the graph's original highest-leverage node, is built and verified — the whole arc's single hard blocker is gone. BMAPWRITE, MINT, and DETACH/EJECT all move from "scoped, blocked" to "scoped, genuinely buildable." Validating it surfaced a real, separate, serious bug (EMPTYBUFBUG) — a stark reminder that this pile still has more surprises in it than the scoping pass alone can find; the "full codebase audit pass" already flagged as deferred (§D.5) has real work waiting for it.

Not yet done: an ordered plan (which node to attack first, given the graph). Per Captain Bob's own framing, that's the next pass — "start asking and answering questions iteratively until we've planned as much as possible."

F. Scoped plans (iterative Q&A pass, started 2026-08-27)

One node at a time: surface the concrete decisions the node's own doc comments/code don't already answer, get a ruling, record it here. This section grows as the pass continues — it's the output of "the punch list of punch lists," not a duplicate of §E.

F.1 — WRITE(10) (the graph's highest-leverage node)

Traced against the live READ(10) implementation in src/starkernel/usb/xhci.c (xhci_bot_read_block() / xhci_bot_send_read10() / the XHCI_XFER_*/BOT_CMD_*/ XHCI_NEXT_ACTION_* deferred-dispatch state machine) before asking anything — the mechanism already exists for READ10 and WRITE10 is a direct mirror of it, not new protocol work:

  • Data direction flips. READ10's Data-In stage rides bulk_in_ring; WRITE10's Data-Out stage must ride bulk_out_ring instead (new XHCI_XFER_BOT_DATA_OUT / XHCI_NEXT_ACTION_BOT_DATA_OUT, reusing the existing xhci_bulk_out_enqueue_and_ring() against bot_data_buf rather than the CBW). CBW's bmCBWFlags clears the DATA_IN bit (host→device) instead of setting it.
  • New SCSI opcodeSCSI_CMD_WRITE10 (0x2A) — plus BOT_CMD_WRITE10 / BOT_TUR_CHAIN_WRITE10, mirroring READ10's TUR-preflight chain exactly (same first-command UNIT ATTENTION handling on a fresh target).
  • blkio_usb.c's usb_blk_write() stops being the BLKIO_ENOSUP stub: memcpy the caller's src into xdev->bot_data_buf, then call the new xhci_bot_write_block() — same shape as usb_blk_read().

Decisions made 2026-08-27 (iterative Q&A pass):

  1. read_only flag in blkio_info() — flips to 0 immediately once xhci_bot_write_block() is proven in QEMU. No separate "bring-up mode" gate to design or later retire.
  2. Validation target — a new disposable QEMU-only test image (e.g. disk/usbwrite-test.img, following the same disposable-clone convention as disk/artemis-metafence-*.img), attached via -drive + -device usb-storage,drive=.... Never disk/artemis.img itself, matching project_disk_images_in_repo's standing rule and how Phase 8's block-fence work was validated (project_metadata_fence_design's verification method).
  3. Scope boundaryWRITE(10) alone is the unit of work. MINT's write-the-identity step and DETACH's flush-back-to-thumb step are separate graph nodes; each gets its own scoping pass once WRITE(10) lands, rather than being pulled into this one preemptively.

Not yet scoped (deferred within this node): exact CSW failure/residue handling for a partial write (READ10's PASS/FAILED logic likely transfers unchanged, not independently verified); whether a write needs a follow-up TEST UNIT READY or cache-flush SCSI command before the transfer is considered durable on real media (QEMU's usb-storage backend may mask this — a real-hardware-only question, revisit at Milestone 8).

F.2 — ACLKEY (the "bumps and holes" comparison semantics)

Traced against the actual code before asking anything: acl_mode/acl_allow/acl_pinned (include/vm.h:343-346, src/word_source/acl_words.c) are DictEntry fields — per word, not per VM. There is no VM-level ACL structure anywhere today. So D.3/D.4's "reuse the existing ACL primitive" resolution can only ever have meant reuse the pattern (a cached decision plus a C-enforced one-way ratchet — if (e->acl_pinned) return; at every mutating primitive) — the actual data has to live somewhere new. sk_repl_set_active_vm() (starkernel/repl.c:64) is confirmed to be exactly what its own doc comment says: an unguarded raw pointer swap, no credential concept at all.

Decisions made 2026-08-27 (iterative Q&A pass):

  1. Where the lock lives. A new dedicated type in its own header — include/starkernel/vm_identity.h, a VMIdentity { uint8_t owner_pubkey[32]; uint8_t installed; } — embedded by value as a field on struct VM, exactly the way VMUuid stadium_vm_id (vm_uuid.h, its own header, embedded field) already does. Explicit standing instruction from Captain Bob: don't grow struct VM/vm.h feature-by-feature with more inline fields — when a feature's data has real shape, it gets its own header and gets integrated as a field, not inlined ad hoc. VMUuid was already the right precedent; VMIdentity follows it. Deliberately holds only the public key, not a seed — unlike Zuse's own zuse_cert_seed/zuse_cert_pubkey pair, a regular VM's lock never needs to sign anything itself, so no private material belongs on it at all.
  2. Zuse override. Zuse keeps an always-allowed override to attach to any VM (consistent with the existing emergency_console/zuse_session bypass precedent, and avoids the bricking risk of Zuse ever being locked out of a VM by its own cert state) — but the override must be a distinct, explicit call path (e.g. a separate function/word, not a silently-taken branch inside the normal bind check), so invoking it requires a deliberate acknowledgment rather than happening as a side effect of an ordinary bind attempt. Exact acknowledgment UX (a confirm prompt? a two-step word?) is not designed yet — the architectural commitment is only that it can't be silent.
  3. Binding pin. No pin. sk_repl_set_active_vm() stays freely retargetable — matches today's Tripod-fleet USE-word behavior, and is needed for real cases where a single console legitimately moves between VMs mid-session (Zuse doing maintenance work in another service VM; a user's console retargeting to update something in a second VM they own). The credential check runs at each bind call, not just once at first attach.

Not yet scoped (deferred within this node): the exact acknowledgment UX for the Zuse override path; whether VMIdentity.installed needs its own one-way fuse semantics (probably yes — the lock itself being immutable once set at a VM's birth is a different axis from "binding stays retargetable," and shouldn't be conflated with decision 3 above) — worth a dedicated look when this is actually built, not decided in this pass.

F.4 — BMAPFMT (on-drive block-map format)

Worked through conversationally, one step at a time, rather than dumped as a finished design — each step below only followed once the previous one was grounded in real code, not assumption.

Step 1 — does a map cell need a state field, not just ownership? Yes, confirmed against existing precedent before answering: the internal Artemis BAM entry (blk_bam_entry_t{allocated, dirty}, block_subsystem.h:257-261) already conflates ownership and state in the simplest map this codebase has. More importantly, MIGSM and UNCLEAN (two already-identified, currently-blocked graph nodes) have nowhere else to record "this range is mid-migration" or "this range's flush was interrupted" without a state field — this node was quietly upstream of both.

Step 2 — should a cell also carry an ACL? Yes, and it can't reuse acl_mode/acl_allow/ acl_pinned directly for the same reason ACLKEY (F.2) couldn't — those are DictEntry fields. Decided: the owner field and any future ACL grant should share the same identity representation VMIdentity already established (a pubkey), not a second encoding — keeps "identity is one common primitive" (D.5) actually true in practice, not just in name.

Step 3 — ACL check ordering. "ACL denial is the fast exit path" — grounded directly against vm.c:611-624's live word-execution ACL check: a cached acl_allow bit is checked first, before any other work, and denial short-circuits immediately. A block cell's ACL needs the identical shape at its top: one cheap cached allow bit, checked before range/ state/owner logic runs.

Step 4 — buffering/granularity, and a major discovery. Confirmed: this kernel already buffers a whole 4 KiB devblock on any block touch (LOAD and friends), and BLK_META_PER_BLOCK (block_subsystem.h:83, 341u /* 341×3 ~= 1023, padded to 1024 */) already packs exactly 3×1 KiB Forth-block metadata slices into that same devblock's spare 1 KiB — the "3-block cluster + 1 KiB metadata" shape volunteered in conversation already exists as blk_meta_t, with real, wired, non-stub accessors (blk_get_meta()/ blk_set_meta(), cached, dirty-tracked, block_subsystem.c:1175-1207). Grepped and confirmed it has zero callers anywhere in the codebase — fully built, fully unused. Its existing "Security & ownership" 40 bytes (owner_id/permissions/acl_block/signature[2]) predate and directly conflict with both the anti-POSIX principle and VMIdentity's pubkey-based model — flagged, not silently reused.

Decisions made 2026-08-27:

  1. BMAPFMT is not a new structure — it's repurposing blk_meta_t. No separate on-drive block-map table gets built. Ownership/ACL/state travel with the block itself (distributed), not in a centralized table (the shape Milestone 3's original wording pictured). One I/O gets payload and map info together; nothing separate to keep in sync.
  2. Consequence flagged, not silently absorbed: homeblocks_sig_t's reserved blockmap_offset/blockmap_devblocks fields (homeblocks_sig.h:93-95) become unnecessary under this decision — that header reserved space for a centralized table this design no longer needs. Needs a comment update in that file once this is actually built, not left silently stale.
  3. Replacement for the 40-byte "Security & ownership" block (same budget the old owner_id/permissions/acl_block/signature[2] occupied; nothing else in blk_meta_t moves):
    uint8_t  owner_fp[8];      /* truncated fingerprint of owner's VMIdentity pubkey, not the
                                   full 32 bytes -- keeps this cheap per-block; full pubkey
                                   resolves via the drive's own identity record. */
    uint8_t  acl_allow;        /* cached fast-deny bit, checked first -- vm.c:611-624's exact
                                   pattern, applied to a block instead of a word. */
    uint8_t  acl_reserved[7];  /* explicitly undecided -- deliberate slack per "flexibility
                                   until we understand the recipe," not a placeholder to fill
                                   reflexively. */
    uint64_t reserved_future;  /* untouched budget, same reasoning. */
    
    flags (already existing, already generic) does double duty as the state field from Step 1 — no new field, just future-defined bit values (CLAIMED/MIGRATING/STALE/etc.). Everything else in blk_meta_t (checksum, timestamps, content_type, hash, chain links, app_data[15]) is untouched.
  4. Allocation granularity: claims quantize to whole devblocks (3 Forth blocks), matching the existing packing — a cell never needs to describe partial-devblock ranges.

Not yet scoped (deferred within this node): the actual allow-list/grant shape beyond the single fast-deny bit (lands in acl_reserved, once designed); the specific flags bit values for each state; whether blk_get_meta()/blk_set_meta() need new FORTH word wrappers or stay C-only like zuse_cert_seed's "no FORTH access" precedent; the actual repurposing edit to block_subsystem.h/.c itself (this pass produced the field design, not the code change).

F.5 — WIREBIND (breadcrumb only — followed into RUNCAP instead, 2026-08-27)

Traced against the live code before asking anything: capsule_birth_baby() (capsule_birth.c:473-614) only ever births a fixed, build-time-baked (p) capsule looked up by name — BIRTH's own mapping is literally S" Artemis" → "artemis:init.4th" (mama_forth_words.c:216). The PERSONALITY stage — where a baby would load its own per-identity content — is a confirmed no-op today: dispatch_init_forth(new_vm) (capsule_birth.c:605) does nothing, with its own comment stating "per-VM block storage is M9 scope; no-op until then."

Finding: the M5 punch-list wording — "spin up or re-attach that identity's VM" — has no mechanism behind it yet. Two real prerequisites were missing from the §E graph and are now added:

  • CERTVERIFY — M5's own wording calls this step "post-identity-authentication"; there's nothing to bind an identity to without a verified identity first. Decision: required, not optional — matches the literal M5 wording and avoids birthing VMs for unauthenticated drives before any lock/key concept (ACLKEY/BINDSTEP) exists to protect them.
  • RUNCAP — without it, "that identity's VM" can only ever mean the one generic, same-for-everyone capsule that happens to exist at build time. Real per-identity content is blocked on RUNCAP (or M9's per-VM block storage, a larger and further-out mechanism).

Working assumption recorded for whenever this node is actually scoped: a WIREBIND- triggered birth always creates a fresh VM (matches capsule_birth_baby()'s existing shape). Recognizing a returning identity and re-attaching to a prior VM/session is ROUNDTRIP's question, not WIREBIND's — deliberately not folded in here.

Not scoped further — by design. Rather than decide WIREBIND's generic-capsule shape in the abstract, the pass followed the real blocker down into RUNCAP itself (§F.6). Revisit WIREBIND once CERTVERIFY and RUNCAP both exist.

F.6 — RUNCAP (runtime capsule construction from thumbdrive content)

Traced against capsule.h/capsule_birth.h before assuming new birth machinery was needed. Two things, read directly from the code, changed the shape of this node:

  1. capsule_birth_baby() is already generic — it takes dir/descs/names/arena as plain parameters; it is not wired to the compile-time capsule_descriptors[] global in any way. mkcapsule sets .desc_capacity = capsule_count always (mkcapsule.c:955) — despite capsule.h's own comment calling this "fixed at compile time for Phase A" (implying headroom for a later phase), there is zero spare capacity reserved today, and no runtime-append mechanism exists into the baked-in directory. Not a blocker, though —
  2. ...because RUNCAP doesn't need to touch the baked-in directory at all. It can construct its own self-contained, heap-allocated CapsuleDirHeader + one CapsuleDesc + one CapsuleNameEntry + a small arena — the exact same shape kernel_main.c:729-750 already builds when it heap-copies the compile-time directory before Mama's own birth, just sized for a single entry instead of the whole set — fill the arena with raw FORTH source text read off the thumbdrive, and hand that straight to the existing, unmodified capsule_birth_baby(). No directory-merge logic, no changes to the birth path itself.

Decisions made 2026-08-27:

  1. Mechanism: heap-built single-entry capsule directory, not baked-in-directory extension. Confirmed safe and minimal per the trace above.
  2. Location of the source content on the drive. homeblocks_sig_t's blockmap_offset/ blockmap_devblocks fields are already dead weight per BMAPFMT's decision (§F.4) that ownership/ACL/state travel with blk_meta_t, not a centralized on-drive table. Repurpose those two fields (rename to identity_src_offset/identity_src_devblocks) to point at this identity's raw FORTH init/personality source instead of adding new fields — reuses already-reserved header space, keeps homeblocks_sig_t's byte budget untouched. (Comment update in homeblocks_sig.h still owed when this is actually built, same as F.4 already flagged.)
  3. Scope of a user's pool — clarified 2026-08-27, corrects the framing this node started from: a user's pool is their entire thumbdrive by default, plus any system-resident device blocks they additionally claim, first-come-first-served, plus the required metadata. The thumbdrive itself is never contested — it's already theirs, no claiming logic needed there. FIRSTTOUCH's "claim a new range at g.total_user_lbn" logic (M3) applies only to the system-device extension, not the drive itself. This simplifies the personality-source question in decision 2 above: no allocation contention to resolve, just a pointer into the drive's own metadata partition.

Not yet scoped (deferred within this node): the exact read path (which forth-blocks within the metadata partition, how identity_src_devblocks bounds a variable-length FORTH source); whether the source is plain FORTH text authored once at mint time or something a user can edit and have re-read on next attach; the interaction with CERTVERIFY (does verification gate reading this content, or only gate the resulting birth). Partially settled by MINT's own scoping pass (§F.8, decision 2): this region's first devblock holds a new user_identity_seed_t record, and the FORTH source itself occupies the remaining devblocks — narrows "the exact read path" question above, though the read-side code is still unwritten.

BUILT + VERIFIED 2026-08-28 (§F.18).

F.7 — CERTVERIFY (Zuse-signed user-cert verification)

Traced against the actual crypto/cert code before assuming D.4's "no new crypto work needed" framing was still accurate. Two files settle the trust-root question outright: x509_ed25519.h's own doc comment ("no signature verification, no chain validation, no extension parsing — this only answers what public key does this cert claim to hold") and zuse_cert_devblock_h's design (Zuse's own on-device Ed25519 keypair, persisted in the system block-fence, FABRIC-3.md §C). Confirmed: a regular user's cert has a completely separate trust root from the capsule-PKI chain — it's signed by Zuse's own locally-resident key, never the offline root CA / snakeoil intermediate. No chain walk is needed at verify time; it's one signature check against a pubkey the kernel already holds.

Decisions made 2026-08-27 (iterative Q&A pass):

  1. Format: X.509/DER, not a simple custom struct (the zuse_cert_devblock_t precedent would have been simpler, but explicitly not chosen). The cert's SubjectPublicKeyInfo holds the owner's Ed25519 pubkey (same field x509_extract_ed25519_pubkey() already reads).
  2. Drive binding: the cert's serialNumber field holds the 16-byte drive_uuid (matching homeblocks_sig_t.drive_uuid) rather than adding a new X.509v3 extension — serialNumber is an early, shallow SEQUENCE member (actually simpler to reach than SubjectPublicKeyInfo, which the existing parser already walks past), so this avoids needing any extension-parsing support at all. Binding to the drive means a valid cert's bytes alone, copied onto a different physical drive, will not verify there.
  3. Revocation: explicitly deferred. Matches this project's incremental-scope precedent elsewhere (capsules have a REVOKED flag; user-cert revocation is separate, unbuilt surface). CERTVERIFY v1 only checks magic/structure, the embedded algorithm OID (Ed25519, rejecting anything else — same check x509_extract_ed25519_pubkey() already does for the pubkey field), the serialNumber-vs-drive_uuid match, and the signature itself.
  4. Real added scope, flagged honestly rather than assumed away: verifying (not just reading) this cert requires genuinely new DER-walking code beyond x509_extract_ed25519_pubkey() — that function stops at SubjectPublicKeyInfo by design. CERTVERIFY additionally needs: the raw TBSCertificate byte range (the exact bytes Zuse's signature covers), the outer signatureValue BIT STRING contents, and the serialNumber INTEGER bytes. Verification itself is then one call: ed25519_verify(zuse_pubkey, tbs_bytes, tbs_len, signature), where zuse_pubkey comes from this system's own zuse_cert_devblock_t, not from anything read off the user's drive.
  5. Location on drive: homeblocks_sig_t.cert_offset/cert_devblocks — already reserved for exactly this ("the CA-signed cert blob") since that header was first designed. No header changes needed here, unlike RUNCAP's repurposing of the block-map fields.

Not yet scoped (deferred within this node): the actual new DER TBS/signature-extraction function's design (this pass decided what it must extract, not its implementation); whether x509_ed25519.c grows this capability in place or a sibling file holds it; how MINT (which shares this exact cert format on the signing side) constructs the DER bytes in the first place — DER encoding has never been needed in this kernel before, only decoding.

F.8 — MINT (the last piece of the original Tripod vision)

Two real dependencies surfaced while scoping this node that weren't in the §E graph, both resolved before any design decision was needed:

  1. GPT. The earlier-decided drive layout (project_usb_thumbdrive_gpt_layout: ~1GB GPT metadata partition + remainder) has no parser or writer anywhere in kernel code. Dropped entirely, per direct instruction (2026-08-27) — not deferred, decided against. The raw homeblocks_sig_t-at-devblock-0 layout every other node in this arc (CERTVERIFY, WIREBIND, RUNCAP, HOTPLUG's own sig_start_fblock=0) already treats as an interim stand-in is now the permanent format. The "small metadata region + remainder is user storage" proportion from the old GPT design still holds — it's just a plain devblock-count convention within one raw device now (metadata = devblocks 0..N covering the sig header + cert + identity-source regions; everything past that is the user's own general block-storage pool directly, consistent with FIRSTTOUCH's pool clarification in §F.6). No GPT-parser node needed in the graph at all — this closes that thread rather than deferring it. Correction, 2026-08-27 (§F.13): homeblocks_sig_t itself moved off devblock 0 to devblock 1 (forth-block 4) once a real collision with the block-subsystem's own generic disk header was found — the "devblocks 0..N" metadata-region framing above still holds, it just now starts one devblock later than stated here.
  2. Single-device USB driver. xhci_dev_t's BOT/MSC state (device_ctx, bulk_in_ring, bulk_out_ring, bot_data_buf) are singular fields, not per-slot arrays — only one USB MSC device is usable at a time today. This looked like it might conflict with the vision's "while Zuse is active, mint a second blank thumb" (implying two drives attached at once). Confirmed no conflict: Zuse's own identity is system-resident, loaded from Artemis's own top-of-device block-fence at boot (kernel_main.c:634-687), never from a removable thumb. MINT only ever needs the one target blank drive attached — today's single-device path is sufficient, no driver work needed for this node.

Decisions made 2026-08-27:

  1. The minted identity gets a real keypair, not just a pubkey badge. virtio_rng + ed25519_keygen() generate a fresh seed/pubkey pair for the new identity, stored on their thumbdrive (never system-resident — matches SSDSCOPE's "regular users stay thumbdrive-resident" read). Keeps the door open for a user VM to sign something itself later, per D.5's "identity is common to every VM" direction, without redesigning the cert format again.
  2. Storage layout (direct extension of RUNCAP's already-decided identity_src_offset/ identity_src_devblocks region, §F.6 — no further header growth needed): that region's first devblock holds a new user_identity_seed_t record (magic/version/seed[32]/ pubkey[32]/crc — same shape as zuse_cert_devblock_t, a new type rather than reusing that Zuse-specific one, per this project's own "dedicated header per real-shaped feature" convention); the remaining devblocks in that region hold the raw FORTH personality/init source RUNCAP reads at birth.
  3. drive_uuid: a fresh, independent 16 random bytes from virtio_rng at mint time (not derived from the identity seed — two separate random draws, matching homeblocks_sig_t's own field being logically "which physical drive," not "whose identity").
  4. Cert construction: builds the exact CERTVERIFY (§F.7) format — SubjectPublicKeyInfo = the new pubkey, serialNumber = drive_uuid, signed with Zuse's own seed via ed25519_sign() (real signing capability already exists, Phase 8). This is the first confirmed caller that needs DER encoding, not just the decoding x509_ed25519.c already has — flagged in F.7 already, restated here since MINT is where it's actually needed.
  5. Header assembly: hdr_crc via the existing homeblocks_sig_compute_crc() — no new checksum logic, this function already does exactly what's needed.

Not yet scoped (deferred within this node): the default personality/init FORTH content a freshly minted identity starts with (shared open question with RUNCAP, §F.6: mint-time-fixed vs. user-editable); a monotonic-ns source for minted_time_ns — checked, and there is no working precedent anywhere in this codebase today (CapsuleDesc.created_ns is hardcoded to 0 at generation time, not a real example to follow); the S" name" MINT FORTH word's own stack/argument shape; whether MINT refuses to overwrite a drive that already reads as a recognized home-blocks drive (HOMEBLOCKS_SIG_OK), mirroring WRITE(10)'s refuse-on-non-blank-media posture — reasonable by analogy, not decided here. The hard blocker is unchanged and already modeled in the graph: WRITE(10) still doesn't exist, so MINT cannot write anything to real or QEMU-emulated USB media yet regardless of the above.

BUILT + VERIFIED 2026-08-28 (§F.19) — every "not yet scoped" item above resolved: default personality is a real, working minimal capsule (not a stub); minted_time_ns stays 0, no precedent found or invented; MINT takes no name argument at all (see §F.19 for why); MINT does refuse an already-recognized drive, decided rather than left as "reasonable by analogy."

F.9 — BINDSTEP (extending sk_repl_set_active_vm() with the ACLKEY check)

Traced against the live retarget path before assuming new plumbing was needed: USE (mama_forth_words.c:430-480) is the real, already-working console-retarget word — S" name" USE looks a VM up by name (case-insensitive) and calls sk_repl_set_active_vm() directly, completely unguarded today. BINDSTEP is concretely "add the ACLKEY (§F.2) comparison to this one call site," not a new mechanism.

Real gap found, not just a policy question: the "re-verify live" approach (decided below) needs access to whatever drive is currently attached, but usb_blk_dev/xdev (repl.c:107-108) are function-static inside sk_repl_idle() — invisible to USE, which lives in mama_forth_words.c. BINDSTEP needs a small new accessor exposing "the currently attached home-blocks device," mirroring the existing sk_repl_get_active_vm() precedent (repl.h:67-69) for exactly this kind of cross-file REPL-state access.

Decisions made 2026-08-27 (iterative Q&A pass):

  1. Check freshness: re-verify live, not cached. Every USE call re-runs the drive-check/CERTVERIFY path against whatever is currently attached, rather than trusting a pubkey cached at original attach time. USE is a rare, human-triggered, interactive operation — the repeated DER-parse + ed25519_verify() cost is a non-issue, and this avoids any staleness question if a drive was swapped without a clean detach.
  2. VMs with no VMIdentity installed yet (installed=0) — Hera/Hermes/Artemis today, until D.5's per-VM-identity work lands — stay freely targetable, no check applied. Preserves USE's current working behavior exactly; refusing on installed=0 would be a real regression today for VMs that were never in scope for this lock in the first place.
  3. For an installed=1 target: BINDSTEP refuses (matching USE's existing clean early-refusal style for not-found/dead-VM cases) whenever either no drive is currently attached, or the freshly re-verified cert's owner pubkey doesn't match the target's VMIdentity.owner_pubkey. No drive attached is a plain refusal, not a special case — it simply means "no identity presented," which can never match.
  4. Zuse's override UX: deferred, per direct instruction — this pass confirms the architectural constraint from F.2 still stands (a distinct, explicit call path, not a branch folded into USE's own logic) without committing to the exact word/interaction yet.

Not yet scoped (deferred within this node): the new accessor's exact name/signature; the refusal message text (matching USE's existing console-message style); whether a refused BINDSTEP attempt should produce an audit-log record (this codebase has a real precedent for this — capsule_parity_log_birth_failed() and friends, parity.c — but nothing says a security-relevant refusal here needs the same treatment; not decided).

F.10 — DETACH (the last M5 node) — and UNCLEAN closed alongside it

Traced capsule_vm_kill() (capsule_birth.c:295-364) fully before scoping, per D.2's own framing that detach should reuse it: it's real and complete — evicts the Stadium patron cell via stadium_evict(), vm_cleanup()+sf_free()s the VM, marks it DEAD, logs via capsule_parity_log_kill(), is idempotent, and refuses to kill Hera. DETACH doesn't need a new teardown mechanism, only a caller for this one plus the flush/bookkeeping around it.

Real finding: D.3's "flush before eject" only makes physical sense as a deliberate pre-removal step — once a drive is physically pulled, there's nothing left to write to. That splits this into two genuinely different paths, and per direct instruction this pass scopes both, closing the previously-separate, long-open UNCLEAN node (M3) at the same time rather than leaving it for later:

  • Graceful path (new EJECT word, no name conflict found): the drive is still physically present. Flush, then kill, then it's safe to remove.
  • Abrupt path (the existing bot_msc_detach_pending hot-unplug signal, already wired and consumed in sk_repl_idle() for blk_subsys_detach_device()): the drive is already gone. No flush is attempted — data since the last flush is lost, which is correct unclean-removal semantics, matching what every real OS does on a surprise USB yank. This is UNCLEAN's entire remaining scope, now closed as "same kill path as EJECT, minus the flush attempt."

One genuinely new, small piece of state required for both paths (not a new mechanism, a missing link): a way to know which VM corresponds to the currently home-blocks-attached drive. BINDSTEP's live re-verify-and-compare approach (§F.9) only works while the drive is still physically present and readable — useless for the abrupt path, where the device is already gone by the time the signal fires. Decided: WIREBIND (§F.5, itself not yet built) must set a small tracked VMUuid — "the VM currently attached via the home-blocks USB path" — at successful birth; both EJECT and the hot-unplug handler read and clear it. No new registry machinery needed beyond this: capsule_vm_registry_get(vm_id, &entry) (already exists) resolves the VMUuid to a live entry, and capsule_vm_kill(entry.name) (already exists, takes a name not an ID — chain the two rather than adding a kill-by-id variant) does the rest.

Decisions made 2026-08-27:

  1. EJECT (graceful): resolve the tracked attached-VM id → entry; blk_vm_flush_all(vm) (already exists, same function sk_repl_idle()'s own periodic sync already calls) while the VM is still alive; if the console's active VM (sk_repl_get_active_vm()) is this same VM, reset it to Hera (sk_repl_set_active_vm(NULL)) before teardown — necessary, not optional, see the dangling-pointer finding below; then capsule_vm_kill(entry.name); clear the tracked attached-VM state. capsule_vm_kill()'s own idempotency covers the case where the VM was already gone by some other path — no extra guard needed.
  2. Hot-unplug (UNCLEAN, closed): identical steps minus the flush — skipped outright, not attempted-and-ignored-on-failure, since the device is confirmed gone by the time this signal fires. Same active-VM reset-before-kill step applies.
  3. W10's edge to DETACH in the §E graph is specifically about the EJECT path's flush step — the hot-unplug/UNCLEAN path never touches WRITE(10) at all (nothing is written during an unclean removal), so UNCLEAN itself is not blocked on WRITE(10), only EJECT is. Worth keeping distinct now that both share one node.

Separate, already-existing bug found while tracing this, reported per standing instruction (not fixed): capsule_vm_kill() never touches g_repl_active_vm. Today, plain KILL on a VM the console is currently USE'd onto leaves g_repl_active_vm dangling — the REPL would fault on the next command dispatched through it. This is real and reachable today, independent of DETACH/EJECT; decision 1 above only fixes it for the new EJECT/hot-unplug call sites, not the pre-existing plain KILL word.

Not yet scoped (deferred within this node): the exact new tracked-state variable's name/location (logically WIREBIND's to own, but WIREBIND itself isn't built yet); whether EJECT needs a name argument or always targets "whatever's currently attached" (leaning toward the latter, given the single-USB-device constraint confirmed in F.8 means there's never more than one candidate — not committed here); the console message text distinguishing a graceful eject from an unclean one in the transcript.

F.11 — FIRSTTOUCH (identity → system-device block-range allocation)

Traced g.total_user_lbn (block_subsystem.c:634,650,655,677,688,753) and blk_meta_t's existing chain fields before assuming the M3 wording was still accurate. Finding: it isn't, fully. "Claim a new range at g.total_user_lbn" predates BMAPFMT's decision (§F.4) to drop the centralized block-map in favor of distributed per-block ownership. total_user_lbn is a global high-water mark that only grows when an entirely new device attaches (RAM, Artemis's disk, a USB drive) — it says nothing about claiming space within an already-attached, fixed-size device. FIRSTTOUCH's real job is scanning Artemis's own system-resident device's existing blk_meta_t records for unowned devblocks, not extending any counter.

Also confirmed live and unrelated to BMAPFMT's repurposing: blk_meta_t's "Link/chain support" fields (prev_block/next_block/parent_block/chain_length, block_subsystem.h:232-236) are real, general-purpose, block-number-granularity linkage — not device-chain-specific — and completely untouched by F.4's field redesign.

Decisions made 2026-08-27:

  1. A claim is a scattered chain, not a required contiguous run. Reuses the existing prev_block/next_block/chain_length fields directly — immune to fragmentation, and this mechanism already exists, unused, waiting for exactly this. The claim's identity to the caller is its chain-head devblock number; owner_fp is written to every devblock in the chain (not just the head), so ownership is directly readable from any member block without first walking the chain — consistent with BMAPFMT's whole point of making ownership locally readable per-block.
  2. Discovery: full linear scan, every time, no cached index. Matches BMAPFMT's own explicit "no centralized table, nothing extra to keep in sync" philosophy exactly. One scan pass serves both questions this node needs answered — "does this identity already own a claim" (owner_fp match) and "which devblocks are free" (owner_fp all-zero) — collected together in the same walk, not two separate scans.
  3. Failure mode: fail outright, no partial claim. If the device doesn't have enough free devblocks (scattered or not) to satisfy the request, return an error and let the caller decide — matches this codebase's existing return-code convention throughout (capsule_birth_baby() and friends), no new partial-allocation semantics invented.
  4. Allocation unit: whole devblocks, already decided in BMAPFMT (§F.4, decision 4) — FIRSTTOUCH doesn't revisit this, just inherits it.

Not yet scoped (deferred within this node): the actual function signature/call site (this pass decided the algorithm, not its C interface); who calls FIRSTTOUCH and when — most likely MIGSM (the migration state machine, M3, still 🟡 partial) deciding "this identity needs N more devblocks," but that call site doesn't exist yet either; whether a first-touch scan should skip/short-circuit once the caller's own already-attached thumbdrive still has room (matching the pool-scope clarification in §F.6: the thumb is the default pool, this node is specifically the overflow case) — implied yes by that clarification, not explicitly re-confirmed here.

F.12 — BMAPWRITE (closed — already built, blocked only on WRITE(10))

Traced blk_set_meta()/cache_writeback()/blk_flush() fully before assuming any new write path was needed. Finding: this node is already done. blk_set_meta() (block_subsystem.c:1198-1210) marks the same cache-buffer entry's meta_dirty flag that ordinary block-data writes use; cache_writeback() (block_subsystem.c:274-285) — called from both blk_flush()'s targeted and flush-all paths, the exact function blk_vm_flush_all() and SAVE-BUFFERS already exercise — packs the in-memory blk_meta_t[] array back into the devblock's metadata region (meta_to_slice()) and writes the whole 4 KiB devblock (data + metadata together, since they share one physical buffer) via write_devblock_4k(), real and unstubbed all the way down to dev->write().

Why this closes rather than just narrows: the M3 item as originally written assumed a separate block-map structure that would need its own bespoke write mechanism. BMAPFMT's decision (§F.4) to repurpose the already-existing, already-wired blk_meta_t accessors instead of building a new table means "writing the map" and "writing a block" were never actually two different operations — the infrastructure that already existed for ordinary block persistence already covers metadata identically, with zero new code. The only reason this hasn't been observed working end-to-end is the same hard blocker already modeled in the graph: WRITE(10) doesn't exist yet, so dev->write() never actually reaches physical media for a USB-attached drive today. Nothing else stands between this and working.

Handoff note for BMAPREAD's own future pass: the read side (cache_load_devblock(), block_subsystem.c:308-322) already unpacks and validates metadata (meta_from_slice(), magic-checked, resets to a clean default on a bad/missing magic) on every devblock load — very likely the same "already built, not a separate item" outcome, but not confirmed here; that node may have its own distinct question (e.g. whether a home-blocks drive needs an explicit validate-on-insert step beyond ordinary lazy per-devblock read-through-cache), left for its own pass rather than assumed away.

F.13 — BMAPREAD (closed — already built) — and a real header collision found and fixed

Confirmed the handoff note from F.12: cache_load_devblock() (block_subsystem.c:308-322) already unpacks and validates every devblock's metadata on load (meta_from_slice(), magic-checked, resets cleanly on a bad/missing magic) — the read side is exactly as "already built, not a separate item" as the write side. Same closure, no new mechanism needed.

But tracing "validate on insertion" turned up a real, previously-unflagged conflict that has nothing to do with blk_meta_t at all. blk_subsys_attach_device() — the general attach path every device goes through, USB drives included — always calls blk_format_or_load_disk() (block_subsystem.c:581-613), which reads devblock 0 looking for the block-subsystem's own 'STFR'/v2 volume header. homeblocks_sig_t's magic is 'LAHB', not 'STFR', so a minted home-blocks drive always reads as "unrecognized" here and lands in BLK_FMT_PROVISIONAL — and blk_get_buffer() (block_subsystem.c:894) refuses all writes to a PROVISIONAL slot until blk_subsys_confirm_format() is called. That function (blk_commit_format(), block_subsystem.c:546-571) writes a fresh STFR header to the same devblock 0 — which would physically overwrite homeblocks_sig_t. Neither outcome (stay PROVISIONAL forever, or get STFR-formatted and lose the identity header) works. This was invisible to CERTVERIFY/WIREBIND/MINT because each only ever traced homeblocks_sig_check() itself, never the generic attach path running alongside it.

Decision made 2026-08-27: homeblocks_sig_t relocates to devblock 1 (forth-block 4), leaving devblock 0 for the generic STFR/v2 header untouched — no changes to the already-tested block-subsystem code. homeblocks_sig_check() already takes its starting forth-block as a plain parameter (sig_start_fblock, not hardcoded), so this is a call-site change, not a format-logic change. MINT (§F.8) must also run the ordinary blk_subsys_confirm_format() path at mint time (writing a ordinary fresh STFR/v2 header at devblock 0) so the drive is immediately writable through the normal block-buffer path, in addition to writing homeblocks_sig_t/cert/identity-source starting at devblock 1. F.8's own text has been corrected in place to point here rather than restate this.

Not yet scoped (deferred within this node): whether MINT's STFR-format-then-sig-write sequence needs a specific ordering/atomicity guarantee (e.g. what state a drive is left in if power is lost between the two writes) — real question, not addressed here; the exact sig_start_fblock value call sites (CERTVERIFY, WIREBIND, RUNCAP, MINT) all need updating from their currently-recorded devblock-0-relative assumption to devblock-1, a mechanical follow-up once any of this is actually coded, not re-litigated here.

F.14 — STALL (the last M2 item) — full BOT-spec recovery designed

Traced the actual completion-code handling before designing anything. Real finding: every transfer completion, control or bulk, shares one gateif (code != XHCI_COMPLETION_CODE_SUCCESS) { console_println("xhci: control transfer failed"); break; } (xhci.c:1267-1270). A STALL isn't distinguished from any other failure today, and there is no recovery of any kind — no xHCI Reset Endpoint command, no USB CLEAR_FEATURE(ENDPOINT_HALT). There is a real safety net though: xhci_bot_wait_for_idle() (xhci.c:831-839) has a bounded iteration count and returns BOT_STATUS_TIMEOUT rather than hanging forever, since the generic bail-out never resets bot_cmd_kind. So today's failure mode is "clean timeout, endpoint left permanently wedged for everything after it" — not a crash, but not recoverable either.

Also confirmed by inspection: none of the pieces real recovery needs exist yet — XHCI_COMPLETION_CODE_STALL_ERROR isn't defined (only _SUCCESS), nor are the RESET_ENDPOINT (xHCI spec value 14) or SET_TR_DEQUEUE_POINTER (value 16) command TRB types. Per direct instruction, full recovery is designed now rather than deferred to Milestone 8.

Decisions made 2026-08-27 — the recovery sequence, mirroring the existing bounded-retry precedent bot_tur_retries/XHCI_BOT_TUR_MAX_RETRIES already establishes for TUR:

  1. Distinguish the completion code. Add XHCI_COMPLETION_CODE_STALL_ERROR (6, per xHCI spec) alongside the existing _SUCCESS (1). The shared completion-code gate branches: a plain non-success/non-stall code keeps today's behavior (log and bail); a stall specifically enters the new recovery path below instead of falling straight to the generic message.
  2. xHCI-level reset (two new command TRB types + two new command functions, mirroring the existing xhci_cmd_disable_slot()/xhci_cmd_configure_endpoint() shape exactly):
    • xhci_cmd_reset_endpoint(dev, slot_id, ep_id) — new XHCI_TRB_TYPE_RESET_ENDPOINT_CMD (14), transitions the stalled endpoint from Halted back to Stopped in the xHC's internal context. Which endpoint is "the stalled one" is already known — it's whichever of bulk_in_ep_addr/bulk_out_ep_addr the in-flight transfer purpose (BOT_DATA_IN/BOT_DATA_OUT/CBW_SENT/CSW_RECEIVE) was using, both already tracked fields.
    • xhci_cmd_set_tr_dequeue_pointer(dev, slot_id, ep_id, new_dequeue, dcs) — new XHCI_TRB_TYPE_SET_TR_DEQUEUE_POINTER_CMD (16), repositions the ring's dequeue pointer past the failed TRB so the next enqueued transfer resumes cleanly.
  3. USB-level clear (reuses the existing control-transfer infrastructure already built for GET_DESCRIPTOR/SET_CONFIGURATION — no new transfer-stage machinery, just a new request payload): send CLEAR_FEATURE(ENDPOINT_HALT) (standard request, wValue=0, wIndex= the stalled endpoint address) to clear the device's own halt condition and reset its data toggle.
  4. Escalation on a second stall (or if step 3 itself stalls): issue Bulk-Only Mass Storage Reset (BOT class request, bmRequestType=0x21, bRequest=0xFF, no data stage — again the existing control-transfer machinery, a different request payload) followed by CLEAR_FEATURE(ENDPOINT_HALT) on both bulk endpoints (BOT spec 5.3.4's full procedure), then retry the original command from scratch — the same "fresh target" restart shape the existing TUR-chain retry already uses.
  5. Bounded, not infinite: a new bot_stall_recoveries counter capped at a new XHCI_BOT_STALL_MAX_RECOVERIES, exact value TBD but same shape as XHCI_BOT_TUR_MAX_RETRIES. Exhausting it sets bot_last_status = BOT_STATUS_FAILED and bot_cmd_kind = BOT_CMD_NONE explicitly — a clean terminal failure signaled to xhci_bot_wait_for_idle()'s caller, rather than relying purely on the outer timeout the way today's un-recovered stall does.

Not yet scoped (deferred within this node): the exact XHCI_BOT_STALL_MAX_RECOVERIES value; whether a WRITE(10)-specific residual-data concern exists on a stalled Data-Out stage (a partially-written SCSI command's recovery semantics may differ subtly from a stalled Data-In read — not analyzed here, WRITE(10) itself still doesn't exist); real-hardware validation is explicitly out of reach until Milestone 8, same caveat WRITE(10)'s own CSW-residue question carries (§F.1) — this design is believed spec-correct but has only QEMU (which rarely if ever issues real stalls) to test against for now.

D.5 — Scope expansion (2026-08-27): identity is common to every VM, not just users

Surfaced while scoping ACLKEY, stated directly: "the whole object is to deliver a functional Tripod plus identities common to all VM's... this means now that we have to create certificates for Hermes, Artemis and Console along with Zuse and CaptBob's 'thumbdrives.'" This reframes VMIdentity (F.2 above) from a user-VM-specific mechanism into the general per-VM identity primitive — every node in the pentagon (Hera, Hermes, Artemis, a user's own VM, the Console) is expected to eventually carry one, not only VMs born from an attached thumbdrive. Doesn't change F.2's design, just its intended reach — worth remembering when VMIdentity actually gets built, so it isn't accidentally scoped narrower than this.

F.3 — HOTPLUG (closed)

Re-checked directly against the live code before asking anything, per Captain Bob's own standing note not to "grep-sample" — read sk_repl_idle() (starkernel/repl.c:82-190) and the bot_msc_attach_pending/bot_msc_attached/bot_msc_detach_pending doc comments in xhci_driver.h:244-265 in full.

Finding: closed. The original Milestone 2 ask was "decide and implement where the hotplug event surfaces... likely a callback registered by whatever owns the home-blocks logic, not xHCI code calling into block_subsystem.c directly." What's built is a flag-set-by-xHCI, flag-consumed-by-sk_repl_idle() pattern — not literally a registered callback, but the bot_msc_detach_pending field's own doc comment states the goal directly: "keeps xhci.c decoupled from block_subsystem.c" — and that decoupling is real and working, confirmed live on all three arches (attach: blkio_usb_open_msc()homeblocks_sig_check()blk_subsys_attach_device(); detach: blk_subsys_detach_device(), symmetric shape).

Confirmed 2026-08-27: close this node as done. No separate design work remains on HOTPLUG itself — the answer for WIREBIND (the node depending on it) is not "build a new hotplug-detection mechanism," it's "reuse this one."

Handoff note for WIREBIND's own future scoping pass: the existing attach point already calls homeblocks_sig_check(), which already distinguishes a recognized home-blocks drive (HOMEBLOCKS_SIG_OK) from blank/foreign/corrupt media (HOMEBLOCKS_SIG_BLANK/ BAD_VERSION/BAD_CRC/READ_ERROR). That result is the natural branch point for WIREBIND: a recognized drive continues into cert-verify + capsule_birth_baby(); a blank/foreign one stays plain read-only block storage, matching the reasoning already documented for the WRITE(10) refuse-on-blank-media logic. Not built yet — WIREBIND still needs its own Q&A pass — but it doesn't start from zero.

Two new standing items that came with this, both explicitly deferred (not part of the current iterative pass, recorded so they aren't lost):

  • A full codebase audit pass, using this same iterative Q&A workflow, to scavenge for more "same gap hiding in two places" opportunities the way ACLKEY turned out to be one gap wearing two names (Milestone 5's "key/lock shape" and D.4's "bumps and holes"), and the way WRITE(10) turned out to gate four separate-looking fronts at once. Not scoped, not started — a future pass of its own.
  • v2.0.0 release target: once Tripod is functional end to end, every VM has its identity mechanism, and the codebase/documentation are clean per the above audit, tag a v2.0.0 release. Stated as the destination this whole planning arc is walking toward, not an immediate next step.

F.16 — The last four standalone nodes: SSDSCOPE, ROUNDTRIP, POLYBLOCK, PROMPTBUG

Closes out every node in §E's graph except WRITE(10) itself. Three of the four close by cross-referencing decisions already made this session, not new tracing; POLYBLOCK needed the user's own clarification first, since it had never been elaborated beyond one sentence.

SSDSCOPE — closed. "Does the system-resident store (Phase 8's block-fence) ever hold anything about regular users, or is it purely Zuse/system bookkeeping?" This was already answered, just not cross-referenced: D.3 states directly — "the system has no requirement to maintain any single identity itself — no user directory, no central registry of who exists. Each thumbdrive is a fully self-contained, portable identity record." This session's own later scoping confirms it structurally rather than just by assertion: CERTVERIFY/RUNCAP/MINT (§F.6-F.8) all put a user's cert, keypair, and personality source on their own thumbdrive, never on Artemis's system-resident fence. FIRSTTOUCH (§F.11) does let a user's claim extend onto system-resident devblocks, but that's anonymous per-block ownership (blk_meta_t.owner_fp, a pubkey fingerprint) written directly on the claimed blocks themselves — not a lookup-by-name directory entry anywhere. The system-resident fence holds exactly one identity record type: Zuse's own (zuse_cert_devblock_t).

ROUNDTRIP — closed. "Does a user's session state round-trip across multiple attaches?" Implied yes by D.3's detach description, now confirmed structurally rather than left implicit: since a user's pool is their own thumbdrive by default (§F.6, decision 3) and EJECT (§F.10) flushes working state back to wherever those blocks physically live before killing the VM, the data simply never leaves the drive between attaches. Reattaching the same drive and re-birthing via WIREBIND+RUNCAP naturally resumes from wherever the blocks were left — no separate "session resume" mechanism needs building; the drive's own persistence is the round-trip mechanism, a consequence of decisions already made rather than a new one.

POLYBLOCK — closed, reframed as an existing principle, not a missing feature. Clarified directly by the user: the block abstraction at this level should have no idea what backs it — "a re-usable, polymorphic definition of a storage device." Traced against blk_get_buffer() (block_subsystem.c:866-900) before concluding anything: it already returns a uniform uint8_t* regardless of whether the backing LBN resolves to RAM, a raw in-memory device (blk_subsys_add_raw_device()), or a real disk-cache-backed device — callers (every FORTH block word, LOAD, etc.) never know or care which. The one backend-specific behavior found (BLK_FMT_PROVISIONAL write-refusal, disk-backed slots only) doesn't violate this — it's encapsulated within the uniform interface (an ordinary failed blk_get_buffer() call), not leaked as caller-visible backend-awareness. This principle is already true of the live code, confirmed rather than assumed.

PROMPTBUG — closed, and found to be less than originally described. Traced for the literal string before closing: no runtime code anywhere prints "zuse)ok> " today — SK_PROMPT_TEXT (repl.c:50) is unconditionally "ok> "; there is no zuse_session-gated prompt-text branch in the live REPL at all yet (only emergency_console gating exists, repl.c:412,459). The typo lives in a doc comment describing not-yet-implemented intended behavior (repl.c:372), not in running code. Closed as: whenever a Zuse-specific prompt is actually implemented, use "(Zuse)ok> " (both parens, generalizing to any attached identity — "(CaptBob)ok> "), matching the already-corrected form recorded in D.4. The comment itself is a trivial one-line fix whenever code in this file is next touched — not applied here, per standing instruction against unrequested code changes.

Every node in the §E graph is now scoped or closed except WRITE(10) itself — the one remaining hard blocker gating real execution of nearly everything captured across this entire §F pass.

D.6 — Console-driven interactive mint onboarding (vision capture, 2026-08-27)

Surfaced live while scoping MSGMIGRATE's message-target question (§F.15) — capture only, per this arc's own "capture first, plan second" discipline; not designed in detail here.

Stated directly: the Console VM (real hardware ownership: serial + framebuffer + PS2 keyboard) is where blank-media minting actually happens interactively, not a bare programmatic MINT call. Zuse's own thumbdrive becoming physically present is itself what starts authentication (no separate manual step). For blank media specifically, the system pulls up an interactive "user StarshipOS (LithosAnanke+StarForth) Mint" onboarding form on the Console:

Full Name:
Address 1:
Address 2:
City:
State/Province:
Country:
Metadata: (hexdump of some encrypted metadata OR QR code later)

Why this matters beyond MSGMIGRATE itself: this directly informs two already-open questions elsewhere rather than sitting alone —

  • RUNCAP's deferred "default personality content" question (§F.6): this onboarding data is very likely part of what a freshly minted identity's personality/init source encodes, not a separate concern. Not confirmed as a final answer — flagged as the likely connection.
  • MINT's own scope (§F.8): minting a new identity now has a real interactive-collection step in front of the cert/keypair/header-writing mechanics already scoped there. The Console becomes an active participant in MINT, not just Zuse triggering it standalone.

Not designed here, explicitly deferred (per the user's own "later" on the metadata field): the "hexdump of some encrypted metadata OR QR code" field's actual mechanism, encoding, and purpose; the exact validation/editing UX for the form itself (can a field be corrected before submit? what happens on a blank/skipped field?); how collected form data actually reaches MINT's execution (addressed at the mechanism level only, in MSGMIGRATE below — the message carries the values, the encoding of the metadata field itself is separate and unaddressed).

F.15 — MSGMIGRATE (the last real design question)

Traced the actual Hermes mechanism before designing anything, rather than treating "message" as an abstract placeholder. Real finding: MSG-DELIVER (capsules/hermes/init.4th:200-203) does MSG-TO@ IDX>NAME VM-EXEC — a message's out-of-line payload (PADDR/PLEN) is arbitrary FORTH source text, executed on the destination VM via the ordinary interpreter, not a structured/typed RPC call. Migrating an interaction to Hermes needs no new dispatch machinery at all — only a decision about what FORTH text goes where.

Also worth being explicit about, since it reframes why this migration is even worth doing: every "VM" in this system is a Forth VM instance living inside one kernel address space, not a separate OS process. There is no correctness reason WIREBIND/BINDSTEP/CERTVERIFY must become messages — Hera's REPL can already call their C functions directly, today, once built. The stated reason to migrate anyway is architectural discipline: routing through Hermes gives uniform heat-tracking/STADIUM-CELL participation in Compudynamics and an audit trail via MSG-SEQ, matching the standing "nothing is done until it's messaging" completion criterion — not solving an isolation problem that doesn't exist here.

Decisions made 2026-08-27:

  1. Real target: the Console VM, not a self-addressed message to Hera. There is a real Console VM today (serial + framebuffer + PS2 keyboard) — this isn't waiting on D.2b's pentagon to become real, it already is. Corrects this pass's own first framing (a self-addressed-to-Hera fallback was floated and explicitly rejected in favor of this).
  2. The flow is two hops, not one:
    • Hop 1 (Hera → Console): Hera's existing hardwired attach detection (homeblocks_sig_check(), unchanged — this is Milestone 2 hardware-driver work, not something that itself becomes a message; nothing exists to receive a message before detection happens) results in a message to the Console VM reporting the outcome — recognized identity, or blank/foreign media.
    • Console-side behavior (not itself a message): for a recognized identity, the Console proceeds toward the existing cert-verify/bind flow; for blank media, the Console runs the interactive mint-onboarding form (§D.6) using its own owned hardware (framebuffer/PS2).
    • Hop 2 (Console → Hera): once the Console has what it needs — either confirmation to proceed with a recognized identity, or the completed onboarding fields for a new one — it sends a message back to Hera to actually execute the privileged operation (MINT/WIREBIND/BINDSTEP), since Hera owns the VM registry these operations mutate.
  3. Payload shape: arguments encoded as literals directly in the payload text, not a dedicated no-argument word reading global state. E.g. the Console's hop-2 message to Hera for a mint would look like a human-typed command line with the collected fields pushed as string literals ahead of the word call (S" Robert James" S" 123 Main St" ... MINT) — more flexible than a fixed no-argument word, at the cost of the Console needing to build that text safely (not addressed here — see deferred list).

Not yet scoped (deferred within this node): the exact literal-encoding/escaping mechanism for building a multi-field FORTH command string safely on the Console side (a real concern — untrusted-ish human-typed onboarding text landing in FORTH source text merits care, not assumed away); the precise message TYPE values for each of the three migrated interactions (this pass decided the mechanism and direction, not the constant catalog); whether CERTVERIFY's own re-verify-live behavior (BINDSTEP, §F.9) changes shape once it's Console- mediated rather than a direct Hera-side check; how this reconciles with PENTAGON's still-open "which of the 10 edges are real today" question (D.4) — this node answers it for exactly the Hera↔Console edge, not the other nine.

F.17 — WRITE(10) implemented and verified (2026-08-28) — the graph's last hard blocker, closed

Implemented exactly as scoped in F.1, a direct mirror of the existing READ(10) machinery with the data direction flipped: new XHCI_XFER_BOT_DATA_OUT/XHCI_NEXT_ACTION_BOT_DATA_OUT states, xhci_bot_send_write10()/xhci_bot_write_block()/xhci_bot_write_data_out() (xhci.c), new SCSI_CMD_WRITE10 opcode and BOT_CMD_WRITE10/BOT_TUR_CHAIN_WRITE10 enum values (xhci_driver.h, kept as fully separate fields from the bot_read10_* ones rather than renamed/reused, so the already-tested READ10 path is never touched). usb_blk_write() (blkio_usb.c) is real now, no longer the BLKIO_ENOSUP stub — read_only in blkio_info() flips to 0, per F.1 decision 1, now that it's proven.

Verified live, all three architectures, with a genuine cold-reboot round-trip — not just a same-session read, the real proof this class of change needs: BLK-CONFIRM-FORMAT's BAM/reloc zero-page writes and an explicit block-content write both completed via clean WRITE10 cycles (CSW PASS every time), then a full kernel rebuild + fresh boot read the written byte back correctly — amd64=65, aarch64=170, riscv64=201, each at LBN 32734 on a disposable disk/usbwrite-test.img attached via QEMU usb-storage, matching F.1 decision 2's validation plan exactly.

One real methodology finding, not project-specific: the disposable test drive must be hotplugged via QMP after boot reaches ok>, never attached at QEMU launch — attaching before xhci_bringup()'s controller reset means the port already reads as connected when the driver starts, so no fresh Port Status Change event ever fires and hotplug detection looks broken even though it isn't (already documented in project_xhci_milestone_2d_polling, re-confirmed here the hard way on the first attempt).

A real, serious, pre-existing bug found and reported (not fixed) during validation: EMPTY-BUFFERS (empty_all_buffers(), block_words.c:148-158) does not implement standard Forth-79 semantics — it force-writes zero to every block on every attached device, rather than discarding cache-slot assignments without writing. Calling it during testing (reasonably assuming standard semantics) mass-zeroed blocks across the whole system, including disk/artemis.img — ~3.3 million bytes, essentially its entire content region. Restored from git, confirmed byte-identical to HEAD. Avoided for the rest of validation (detach/reattach via QMP used instead to force a fresh device read, which is non-destructive).

Consequence for the rest of the graph: W10 was the single hub gating BMAPWRITE (already independently confirmed working, §F.12), MINT, and DETACH/EJECT's flush step — all three are now genuinely unblocked, not just scoped. Nothing about their own design changes; they can now actually be built and tested rather than only designed. Real hardware (Milestone 8) remains untested — this verification is QEMU-only, per F.1's own acknowledged scope.

D.7 — Birth-by-message-only, universal to Zuse and user VMs (vision, 2026-08-28)

Stated directly, prompted by planning the freshly-minted-zuse.img-thumbdrive-to-(Zuse)ok> flow: "Hera gives birth only by request once messaging and storage are functional and should display no coupling outside of messages. This will include 'user' VMs too." A sharper, universal restatement of the standing "nothing is done until it's messaging" completion criterion (D.1) — not a new principle, but a concrete commitment that this specific, next flow should go straight to the message-only shape rather than pass through a hardwired interim first, since messaging (Hermes) and storage (Artemis) are both now genuinely live.

Concretely, for the mint-zuse.img-thumbdrive-to-(Zuse)ok> flow:

  1. Hera's own hardware polling (sk_repl_idle()/xhci_poll_events()) is unavoidable — she's the only VM whose C code has xHCI MMIO access. This stays.
  2. But Hera must not decide or act on what she detects. Per direct instruction: detection ownership moves to a running Artemis VM instance — a real change from what HOTPLUG (closed) and WIREBIND (§F.5, still just a breadcrumb) already built/scoped, both of which currently have Hera deciding and acting directly. Hera's role narrows to relaying the raw hardware event; Artemis owns the identity-verification decision and the birth request.
  3. Grounded, not assumed: block LBN space is one flat, global address space (block_subsystem.c's g.total_user_lbn), not per-VM — any VM's dictionary can already read any LBN once told which one, no new C-level read plumbing needed for the identity-check content itself. The gap is coordination (when, which LBN), not capability.

Real, load-bearing prerequisite surfaced by tracing, not assumed: Artemis is not persistently alive today. Every boot this session confirms she's birthed for her own 4.6 migration self-test, then explicitly killed (KILL: Artemis dead) before the REPL ever reaches ok>. For any message to reach a running Artemis instance, she has to actually stay born — a real, new requirement this direction depends on, not something HOTPLUG/WIREBIND ever needed before. CLOSED 2026-08-28: kernel_main.c's item 4.2/4.6 self-test blocks (birth → exercise diagnostics → KILL) stripped to birth-only; verified live on all three architectures — BIRTH: Hermes live/BIRTH: Artemis live, zero KILL: anywhere in any log, and an interactive USE round-trip confirmed Hermes genuinely reachable post-boot. Found and fixed a real regression along the way: Hermes's CD-INIT (arena init, and the thing that loads lib.4th into her own dictionary) was only ever invoked by the removed self-test — added an unconditional call at the end of her own init.4th so a real birth actually initializes her. Also added STARTUP-BANNER (scaffolding — prints (no identity) until CERTVERIFY/MINT exist to populate it from a thumbdrive's PKI fields), wired into each Tripod leg's own init.4th.

What this reshapes, not yet re-scoped:

  • WIREBIND (§F.5) was deliberately left a breadcrumb, not fully designed, specifically because "wire hotplug → capsule_birth_baby() directly" is exactly the hardwired coupling this principle now rules out — this direction is the real answer that breadcrumb was waiting for, not a contradiction of it.
  • MSGMIGRATE (§F.15) scoped a two-hop Hera↔Console flow (Hera reports attach, Console requests the privileged operation). Artemis is a new participant, not a replacement — identity verification/storage (Artemis's role) and human-facing onboarding UX (Console's role, §D.6) are both real and distinct; the full flow likely needs both, reconciled together rather than re-litigated as either/or.
  • Applies universally, per direct instruction — the same message-only shape governs ordinary user-VM birth too, not a Zuse-specific carve-out.

Not yet scoped: whether Artemis becoming persistent changes her own self-test's shape or timing; the exact message types/payloads for Hera→Artemis (raw attach event) and Artemis→Hera (birth request); how/whether Console fits into this specific flow versus the already-scoped MSGMIGRATE two-hop; whether CERTVERIFY (§F.7) itself should now run as Artemis-mediated logic rather than a bare function call. Capture only, per this arc's own established discipline — design happens in its own pass.

Phase C — distributed messaging capsule + idle-loop pump (CLOSED 2026-08-28)

Prerequisite for D.7's message-only flow: every VM needs its own real MSG-SEND/CH-* vocabulary and arena, not just Hermes. Extracted the generic messaging vocabulary (arenas, field accessors, MSG-ALLOC/CH-ALLOC, MSG-SEND/MSG-DELIVER/MSG-TICK, CH-ADD-MBR, MSG-BROADCAST, ~30 words total) out of capsules/hermes/init.4th into a new shared capsules/common/messaging.4th (blocks 50035037). Hermes and Artemis each EXEC it at birth and get their own private MSG-ARENA/CH-ARENA copy (CREATE runs per-VM). Hermes stays the one owner of the canonical COMMON-CH; Artemis subscribes into it via VM-EXEC at her own birth (2 COMMON-CH @ CH-ADD-MBR run inside Hermes), and Hermes proactively subscribes Hera (idx 0) herself since Hera always exists first and can't subscribe to a VM that doesn't exist yet at her own boot.

Hera does not get her own copy — a real, discovered constraint, not a design choice. First attempt loaded common:messaging.4th into Hera's own dictionary too, matching "every VM that gets born must have all these caps." Live-tested and found broken: every colon-definition that referenced a STADIUM-* primitive (MSG-HEAT@/!, CH-HEAT@/!, MSG-COOL-ALL, MSG-TICK itself) was silently missing from Hera's dictionary after boot — not a compile error, just absent. Root cause, traced via a live serial socket and isolated word-by-word: register_child_vm_words()'s own doc comment (mama_forth_words.c) explains that the eight STADIUM-* FORTH words are deliberately never added to register_mama_forth_words(), specifically to keep Hera's own dict_hash off item 4.1's baseline. This is a pre-existing, intentional constraint, not a timing bug — moving the load later in boot (tried first, also failed) didn't help, because the words are never registered for Hera at all, at any point. Reverted: Hera orchestrates via BIRTH/VM-EXEC/VM-CALL directly instead (already proven live — VM-CALL querying Hermes's MSG-K cross-VM works fine), and the idle-loop pump below explicitly skips her own registry entry.

Idle-loop pump (repl.c's existing sk_repl_idle(), ~1s cadence): added capsule_vm_registry_get_by_index() (capsule_birth.c/.h) for registry enumeration by birth-order position (no by-index accessor existed before — only by-vm_id and by-name). Each idle beat, Hera walks every live registry entry except her own and VM-EXECs MSG-TICK into it — "fully distributed, Hera pumps each VM's drain," the design confirmed before Phase A/B began.

Verified clean (zero UNKNOWN WORD/VM-EXEC: ERROR after birth, MSG-STATUS answering correctly on both Hermes and Artemis) on all three architectures. Commit 21bca31.

Still open, unblocked by this: WIREBIND/BINDSTEP rebuilt against real messages instead of direct calls; common/msg.4th's HERMES-ACK/HERMES-NACK wrappers are now obsolete (every VM has its own local MSG-ACK-LAST/MSG-NACK-LAST — no VM-EXEC indirection needed) but the file itself was left in place, unloaded, rather than deleted unprompted; capsules/MANIFEST.md's "immutable ABI" claim for block 4055 is now stale.

F.18 — RUNCAP built + verified (the §F.6 prerequisite MINT was blocked on)

Implementing RUNCAP against §F.6's own trace ("hand that straight to the existing, unmodified capsule_birth_baby()") surfaced two real gaps that trace missed, both closed before any code shipped:

Gap 1 — capsule_birth_baby()'s signature check is not actually parameterized the way §F.6 assumed. It calls capsule_verify_signature(descs, names, capsule_get_signatures(), arena, dir->desc_count, idx)descs/arena/dir are caller-supplied (fine for a heap-built directory), but capsule_get_signatures() unconditionally returns the compile-time-baked global array, indexed against the build-time capsule_descriptors[], not anything the caller controls. For a RUNCAP-built directory idx is always 0, so the check would compare RUNCAP's own arena bytes against whatever real capsule happens to sit at slot 0 of the baked array — guaranteed to fail, not a security check. Put to Captain Bob directly (three options: explicit bypass flag / a signed placeholder slot / defer); decided: explicit skip_pki_sig flag, 0 for every existing call site, 1 only for RUNCAP — "easy to change our minds later," and confirmed real enforcement for user-originated content lives downstream in the VM's own ACL (VMIdentity/acl_caps, §F.2), not in this signature check, which was never meaningful for non-build-time content in the first place. capsule_birth_baby() itself is otherwise unchanged; all 4 existing call sites (mama_forth_words.c) pass 0.

Gap 2 — capsule_birth_baby() never sets the registry entry's own .name. vm_registry_alloc() zero-initializes it, and nothing inside capsule_birth_baby() ever fills it in — every existing caller (mama_word_birth, CONNECT-HERMES/-ARTEMIS) does this itself afterward via capsule_vm_registry_set_name(). Found live, not by inspection: a RUNCAP-born VM's own registry name stayed empty, which Phase C's idle-loop pump then read as a zero-length string and refused every tick ("VM name too long or empty") — traced via a live serial-socket session, not guessed. capsule_runcap_birth() now calls capsule_vm_registry_set_name() itself, matching the existing pattern.

A third, non-blocking discovery while building the live test fixture: capsule_exec_payload() (capsule_loader.c) requires a "Block NNNN\n" header on every chunk of IDENTITY content — without one, is_block_header() never matches, the whole payload is treated as "pre-block/unrecognised" and silently skipped line-by-line, never executed (birth still reports success — there's simply nothing to run). Not a bug, but load-bearing context for whoever authors MINT's default personality content (§F.6/§F.8's still-deferred question): that content needs the same header convention every hand-authored .4th capsule already uses, not bare FORTH text.

Verified end-to-end, live in QEMU, not just build-clean: a standalone host-side generator (mirroring the CERTVERIFY host-test precedent, §F.7) wrote a real homeblocks_sig_t + synthetic identity-source content directly into a disposable thumbdrive image's raw devblocks — no capsule build, no mkcapsule, nothing baked in at compile time. Hotplugged via QMP, recognized as a home-blocks drive, then a diagnostic RUNCAP-TEST word (Mama-only, temporary — not the real call site, that's still WIREBIND) called capsule_runcap_birth() against it: a genuinely new VM (RunCapTest2) was born, and VM-EXECing a word defined only in that thumbdrive's own source printed 424242 — proof the content was actually read, compiled, and executed, not just accepted. Clean 3-architecture regression boot (no RUNCAP drive attached) confirms no side effects on normal boot. Commit e1e8392.

Still open: the real call site is WIREBIND (§F.5), still unbuilt — RUNCAP-TEST is diagnostic-only and should not survive into that flow. A RUNCAP-born VM has no common:messaging.4th vocabulary unless its own minted personality content loads it (matches §F.6/§F.8's already-deferred "default personality content" question, not a new gap — observed live as the idle pump's expected MSG-TICK refusal for RunCapTest/RunCapTest2, harmless since these are test-only VMs never present during a normal boot). MINT (§F.8) is now genuinely unblocked.

F.19 — MINT built + verified: the full Tripod identity vision works end to end

With CERTVERIFY (§F.7/§B), RUNCAP (§F.6/§F.18), and Zuse's own on-device keypair (Phase 8, zuse_cert_devblock.h) all real, MINT closed the last gap: nothing could produce a drive RUNCAP/CERTVERIFY could consume. Same discipline as §F.18 — build against the real code, verify live, don't trust the scoping pass's own assumptions without checking.

New DER capability: x509_build_user_cert() (x509_ed25519.h/.c) — the encode-side counterpart to the three existing decode functions. A minimal DER TLV writer (short-form length only; every field here stays well under the 128-byte threshold except the two outer SEQUENCE wrappers, hand-encoded with one long-form length byte). issuer/validity/ subject are each an empty SEQUENCE — valid, zero-length TLVs the decode side already only ever skips by length, never reads the content of; version is omitted entirely (implicit v1, matching the decoder's own optional-version handling). Host-tested round-trip before trusting it in the kernel (same precedent as CERTVERIFY's own host test, §F.7): built two certs (one with an ordinary serial, one with a high-bit-set serial forcing the DER zero-pad-integer rule), decoded both back with the real x509_extract_ed25519_pubkey()/ x509_extract_serial()/x509_verify_signature() — pubkey and serial round-trip exactly, signature verifies against the real issuer, correctly rejects the subject's own key and a corrupted signature. 10/10 assertions pass. Cert size: 155 bytes, comfortably one devblock.

Fixed devblock layout, decided rather than left open: devblock 1 = homeblocks_sig_t (unchanged, §F.13); devblock 2 = the DER cert (cert_offset/cert_devblocks=1); devblock 3 = new user_identity_seed_t (magic/version/seed[32]/pubkey[32]/crc, same convention as zuse_cert_devblock_t — a dedicated type, not a reuse, per this project's own "give real-shaped data its own header" rule); devblock 4 = the default personality source. metadata_devblocks=5 (0..4 reserved; devblock 0 stays the block-subsystem's own generic header, untouched). homeblocks_sig_t is written last, deliberately, so a crash partway through minting never leaves a drive that reads as a complete, valid home-blocks drive with content actually missing behind it.

Real keypair, real entropy, real signer (§F.8 decisions 1/3/4, unchanged from scoping): virtio_rng_get_bytes() + ed25519_keygen() for the identity's own seed/pubkey; a second, independent virtio_rng_get_bytes() draw for drive_uuid (never derived from the identity seed); the cert signed with the caller's zuse_cert_seed (in practice always Hera's own VM struct, where Zuse's on-device key lives) via ed25519_sign().

Default personality is real, working code, not a stub — a small FORTH capsule defining and immediately calling WELCOME, which prints a real banner. §F.6/§F.8's "default personality content: mint-time-fixed vs. user-editable" question stays genuinely open — this doesn't answer it, but whatever mints today has to actually do something once RUNCAP births it, matching this project's "no stubs" discipline even where the surrounding design question is deliberately deferred. Needs the same "Block NNNN\n" header §F.18 found capsule_exec_ payload() requires — baked in directly, not left for a future pass to discover the hard way again.

One new decision, not left "reasonable by analogy" any longer: MINT refuses to write onto a drive that already reads as HOMEBLOCKS_SIG_OK (MINT_ERR_ALREADY_MINTED) — mirrors WRITE(10)'s own refuse-on-non-blank posture, now decided rather than deferred.

No name argument, deliberately deviating from F.8's own original S" name" MINT sketch: this design never binds a VM name at mint time at all — the drive holds an anonymous, self-contained identity; a name is assigned later, at RUNCAP birth, which is a WIREBIND decision, not MINT's. MINT ( -- ok? ) takes nothing and operates on whatever sk_repl_get_attached_blk_dev() (new accessor, repl.h/.c, mirrors §F.18's home-blocks-only one but doesn't gate on recognition — a blank drive by definition never sets that one) currently reports.

Verified end-to-end, live in QEMU — the full mint-to-birth flow working for the first time: a genuinely blank test image, hotplugged and confirmed "not recognized (blank or foreign media)"; MINT ran and reported success; RUNCAP-TEST on the same drive then birthed a real VM from its own freshly-minted content; VM-EXECing WELCOME printed the default personality banner. Every layer built this session — CERTVERIFY's DER decoder, RUNCAP's capsule construction, MINT's DER encoder and keypair generation — composed correctly on the first real end-to-end attempt. Clean 3-architecture regression boot (no test drive attached) confirms no side effects. Commit f6e2737.

One real workflow gap found, deliberately not fixed in this pass: a freshly-minted drive is not recognized by RUNCAP until physically detached and reattached — the home-blocks sig cache (§F.18's g_homeblocks_dev/sig) only refreshes at USB attach time, and MINT (which just wrote a valid sig it knows is now correct) doesn't update that cache itself. Worked around for this verification via a QMP detach/reattach; a real user plugging in a blank drive, minting it, and expecting it to work without unplugging would hit this. Flagged for whoever builds the real Console onboarding flow (§D.6) — that flow controls the actual mint UX and is the right place to decide whether MINT should refresh the cache itself or the onboarding flow should prompt for a reattach.

Still open: WIREBIND/BINDSTEP (§F.5/§F.9) — the real call sites for both RUNCAP and MINT; RUNCAP-TEST and MINT's own bare console word remain diagnostic scaffolding, not meant to survive into that flow unchanged. With CERTVERIFY, RUNCAP, and MINT all real and composing, WIREBIND is now the last piece standing between this arc and D.7's original message-only birth vision.

F.20 — general-purpose console vision + MINT profile fields (2026-08-28)

New direction, stated directly: get rid of the emergency CLI entirely (today's unauthenticated ok> bypass) in favor of a general-purpose console that attaches to any VM an identity has ACL access to — Tripod legs included, not just user-spawned VMs — via two complementary mechanisms, not alternatives: BINDSTEP's already-scoped ACL-gated USE redirect (cheap, direct, the admin/debug path) and a genuinely new console-VM + user-VM pair, talking over the real messaging layer (Phase C) rather than direct dispatch — the actual next step toward the message-passing-OS vision, since every line becomes a message instead of a C-level redirect. Zuse stays "no different than any other VM" except her ACLs (VM_IDENTITY_CAP_ALL); ACLs are inherited (mechanism not yet designed — real open item once the console-VM pair is scoped).

Real gap surfaced immediately, not assumed away: removing the emergency CLI removes the only currently-unauthenticated path into a session — but Zuse's own identity today comes from an unauthenticated first-boot self-mint (kernel_main.c's virtio_rnged25519_keygenvm_zuse_cert_install()zuse_cert_devblock_t block-fence write). Decided: Zuse becomes thumbdrive-resident too, minted through this same MINT flow, not system-resident via the block-fence. The block-fence flow's own retirement is scoped, not yet built — that's the next concrete step once the console-VM pair work reaches it.

Immediate, concrete, scoped ask actioned first: MINT needed to carry a real human profile — full name, username, email, phone (email/phone nullable) — before minting real test identities makes sense. Extended user_identity_seed_t to v2 with fixed full_name/ username/email/phone fields (plenty of unused pad space already there — 4016 bytes). Deliberately NOT encoded into the DER cert's Subject field — that would mean a real X.509 RDNSequence (commonName/emailAddress OIDs, PrintableString/UTF8String tagging), well past x509_ed25519.h's own stated "deliberately not a general ASN.1/X.509 [builder]" scope. This data isn't security-relevant the way pubkey/serial are (the only two fields CERTVERIFY/ BINDSTEP actually check) — it travels alongside the keypair in the plain identity record instead. capsule_mint_identity() and the MINT word both grew to take it (MINT's stack signature is now 4 string pairs, ( fname-c fname-u uname-c uname-u email-c email-u phone-c phone-u -- ok? )).

Apollo test rig: three disk fixtures for this next phase of live testing — disk/zuse.img (pre-existing from Phase 8 kickoff, blank, exactly the intended Zuse-thumbdrive target), disk/user1.img (new, 64MB, a regular-user thumbdrive), disk/apollo-ssd.img (new, 30MB, matching Artemis's own size convention, deliberately separate from the actively-used disk/artemis.img so this phase starts from a clean slate).

Verified live in QEMU: minted two real identities with real data — Zuse (full_name "Konrad Suse", username "Zuse", zuse@pantheon.org, +18885555555) onto disk/zuse.img, and a regular user (full_name "Captain Bob", username "CaptBob", capt.bob@pantheon.org, +18885555556) onto disk/user1.img. Both minted only one at a time (the single-USB-device constraint §F.8 already noted). Read the raw devblock bytes back off both images directly afterward and confirmed every field byte-exact at its correct struct offset (record starts at devblock 3 = byte 12288; full_name at +72, username at +136, email at +168, phone at +232). Clean 3-architecture regression boot confirms no side effects. Commit 6fd8792.

Not yet done, deliberately: disk/zuse.img/disk/user1.img now hold real, live-minted keypairs — not committed alongside the code this pass, pending a decision on whether these become tracked fixtures or stay local-only. The bigger swap this section opened (Zuse's first-boot mint routed through MINT onto a thumbdrive, block-fence flow retired, emergency CLI removed) is scoped in direction but not yet traced against the actual boot-sequence code the way RUNCAP/MINT themselves were before being built (§F.18/§F.19's own discipline) — that trace-first pass is the next step, not this one.

Update: disk/zuse.img/disk/user1.img committed after this section closed, per direct instruction. The bigger swap traced and built in §F.21, same pass.

F.21 — emergency CLI retired; Zuse goes thumbdrive-resident; ACL.4th activated

The trace §F.20 asked for, done before touching code (same discipline as §F.18/§F.19): vm->emergency_console is not a boot-time flag, it's recomputed on every REPL iteration (repl.c, both sk_repl_step and sk_repl_run): is_hera && !zuse_session → 1, and when set it skips ACL enforcement entirely at both the colon-word dispatch point and the outer EXECUTE path (vm_core.c). A real finding this trace surfaced: capsules/init.4th had ACL.4th's own load commented out all session — ACL enforcement was inactive, bypass or no bypass, the whole time, because the policy engine that would ever deny anything had never loaded. Removing the bypass alone would have changed nothing observable.

Design, confirmed with Captain Bob before building: system-resident storage keeps only Zuse's public key, never her seed — the fence remembers "a root identity exists" and what its pubkey is; the private key lives exclusively on her own minted thumbdrive. New zuse_genesis_marker_t (magic/version/zuse_pubkey[32]/crc) takes over the fence's devblock 0 slot zuse_cert_devblock_t used to occupy; that type is kept in the repo, header marked superseded, no longer written by any path.

Genesis can't use regular MINT — there's no existing Zuse to sign a cert with yet. capsule_mint_identity() grows a genesis mode: issuer_vm=NULL skips the cert build/write entirely (cert_offset/cert_devblocks stay 0) and skips the zuse_cert_installed check. Two new optional out-params, out_pubkey/out_seed, let the caller (genesis's only caller) install the cert immediately after minting — the seed otherwise never leaves the thumbdrive.

A real boot-ordering constraint, not assumed away: virtio_rng_init() already precedes the old Zuse block, and xhci_bringup() already precedes sk_repl_run() — no reorder was actually needed. What matters instead: USB attach detection only ever happens inside sk_repl_idle()'s own polling, which doesn't exist yet at kernel_main.c's old one-shot mint point. So the genesis/attach logic can't be a boot-time step at all, regardless of ordering — it has to live in the idle loop. New capsule_zuse_boot_try_attach() (capsule_zuse_boot.c), called from sk_repl_idle() on every fresh attach: no-ops if Zuse already has a cert this boot; no marker + HOMEBLOCKS_SIG_BLANK → genesis-mint, write the marker; marker present + HOMEBLOCKS_SIG_OK + matching pubkey → read the drive's own user_identity_seed_t, install the cert. Either path ends the same way: vm_zuse_cert_install() then vm_interpret(mama_vm, "ACL-ZUSE-BOOT") — re-running the existing self-activation word rather than adding new C-side auth logic (ACL-PIN only blocks redefinition, not re-execution, so calling the same policy word again later is the correct, already-designed way to activate zuse_session once a cert exists later than boot).

A real, previously-dormant bug, found the moment ACL.4th actually loaded for the first time: zuse.4th's ACL-ZUSE-BOOT tried ['] ACL-ZUSE-BOOT ACL-PIN from inside its own still-compiling definition — the word isn't findable via ['] at that point (mid-compile), so the whole colon-definition silently failed to compile, every boot, for as long as ACL.4th has existed in this repo — just never noticed because nothing ever loaded it. Fixed: pin after the definition closes (: ACL-ZUSE-BOOT ... ; then, separately, ['] ACL-ZUSE-BOOT ACL-PIN) — it only ever needs to run once anyway, and pinning doesn't block the re-invocation genesis/attach depends on.

The bypass itself, removed: the emergency_console = is_hera ? ... : 0 assignment deleted from both repl.c call sites. Every word run from Hera's own bare prompt now goes through ordinary ACL enforcement. emergency_console the field still exists and is still set, briefly, by the genuine C-level VM fault handler (EMERGENCY_CONSOLE_ENABLED) — a distinct, narrower mechanism this REPL no longer touches.

Verified end-to-end, live, not just build-clean: added ZUSE-SESSION? (read-only diagnostic, matches ZUSE-PUBKEY@'s own precedent) to check the chain directly rather than by inference. Fresh boot, no thumbdrive attached: ZUSE-SESSION? reads 0. Attach a genuinely blank drive via QMP: genesis mint fires automatically, no typing, console confirms it; ZUSE-SESSION? then reads -1 (true) — the complete chain (genesis mint → cert install → ACL-ZUSE-BOOT re-run → ZUSE-AUTHENTICATEzuse_session=1) verified working, not assumed. Clean 3-architecture regression: Hermes/Artemis both birth live, zero ACL denials, zero UNKNOWN WORD beyond the deliberate POST self-test cases (NOSUCH/NONEXISTENT/ __acl_deny_test__) — the first time this whole session ACL.4th has actually been enforcing anything, and nothing broke. Commit cc9521d.

Still open, deliberately not chased in this pass: the fence marker write itself still fails ("fence marker write FAILED (not persistent)") — the same pre-existing, already-flagged anomaly this session has seen with the old block-fence write path too (§ various), now manifesting through the new marker write instead. Doesn't block same-boot activation (that reads the freshly-generated data in memory, not the fence), but means the "already minted" marker doesn't actually survive a reboot in this QEMU test environment yet — the attach-authenticate path (marker present, matching drive reattached later) is designed and built but not yet live-verified across a real reboot, blocked on this separate write-path issue. A second, unrelated, general finding while testing: LOG-INFO" output from any FORTH-defined colon-word body (tested standalone, unrelated to Zuse) doesn't appear on the console at all in this build — (do-log-info) compiles and runs without error, just prints nothing visible. Not chased; didn't block verification since ZUSE-SESSION? gave a direct answer instead of needing the log line. With this closed, the console-VM + user-VM pair (§F.20's own opening vision) is unblocked.

F.22 — console-VM + user-VM pair: real async message relay

§F.20's opening vision picked up: a general-purpose console attaching to any VM with ACL access (BINDSTEP, §F.9 — unchanged, stays the cheap direct/admin path) and a genuinely new console-VM + user-VM pair talking over real messages, the actual next step toward "message-passing OS." Two shapes were on the table — synchronous VM-EXEC relay (low latency, still built on the same primitive real messages use) vs. fully async MSG-SEND/ MSG-DELIVER (architecturally purer, costs the idle-pump's cadence per round trip). Chosen: Option B, fully async — confirmed directly: messaging is a general capability every VM already has for its own reasons, not something that should get a console-specific shortcut bolted on beside it.

Traced before building (vm_interpret(), vm.c/vm_core.c): it's pure C, per-line, no per-VM hook for "redirect this whole line elsewhere" — a console that turns keystrokes into messages can't be a new vocabulary alone, it needs an actual new C-level REPL mode. Tracing MSG-SEND/MSG-DELIVER (Phase C) also confirmed the addressing model already works for this: whichever VM calls MSG-SEND both stores the message and later delivers it (via its own MSG-TICK), and VM-EXEC copies the string bytes across the VM boundary at delivery time — so a console VM messaging its paired user VM is exactly the same shape every other message in this system already uses, no new plumbing needed.

Built: new CONSOLE-CMD-EVENT message type (common:messaging.4th). New sk_repl_dispatch_line() (repl.c), replacing the direct vm_interpret() call in both sk_repl_step/sk_repl_run: if the active VM's own name has a live "<name>~user" counterpart registered, the raw line is wrapped as an S"-embedded CONSOLE-CMD-EVENT MSG-SEND and interpreted on the console VM instead of run directly — delivery happens later, on that console's own next MSG-TICK (Hera's idle pump), via VM-EXEC into the paired user VM, identically to every other message. Falls back to direct interpretation with no pairing, or if the line contains a " (can't be safely S"-embedded yet — a known v1 limitation, warned about explicitly rather than silently mishandled, not chased further this pass).

New capsule_console_birth() (capsule_console.h/.c): a bare VM whose only content is loading common:messaging.4th — parallel in shape to RUNCAP's user-VM birth but with fixed embedded source instead of a devblock read (no identity, no thumbdrive involved on the console side at all). New PAIR-TEST diagnostic word (matches RUNCAP-TEST's own precedent exactly): births both halves and registers the "<name>~user" mapping. Not the real pairing call site — that's the eventual attach/onboarding flow — exists to exercise the relay live before that flow exists.

A real, serious bug found and fixed live, not assumed away: console_set_vm_name() stored the caller's raw pointer, never copied it. mama_word_use() (USE) passes a VMRegistryEntry field living on its own stack frame — once USE returns, that pointer dangles, corrupting the console tag on every subsequent line. Observed directly as garbled "[[]" / binary-looking prefixes instead of "[CaptBob]" — not a theoretical concern, this broke the very relay mechanism being tested, since sk_repl_dispatch_line() reads the same name to find the pairing. Fixed at the source: console_set_vm_name() now copies into internal storage. That fix surfaced a second, related bug: every console_get_vm_name()- based save/restore call site in mama_forth_words.c (BIRTH, VM-STEP, VM-EXEC, CONNECT-HERMES, CONNECT-ARTEMIS) saved only a pointer into that same single internal buffer — an intervening console_set_vm_name() call silently corrupted the saved value before the restore ever ran. New console_save_vm_name() copies into caller-owned storage; every save/restore site updated to use it.

Verified end-to-end, live, not just build-clean: typed WELCOME at a paired console VM — it did not execute directly (no UNKNOWN WORD), printed ok immediately (queued, genuinely async), and on the next idle tick "[CaptBob~user] Minted identity -- default personality" appeared on its own — real delivery and execution in the paired user VM through the actual MSG-SEND/MSG-DELIVER pipeline, not a shortcut standing in for it. Console tags confirmed clean (no garbling) across all three architectures' full regression boot; Hermes/ Artemis both birth live, zero unexpected ACL denials or UNKNOWN WORD. Commit b0f1271.

Still open: the real pairing call site (PAIR-TEST is diagnostic scaffolding only, not meant to survive into a real attach flow); the " character in a relayed line (falls back to direct interpretation with a warning, not fixed this pass); the idle-pump's ~1s cadence is the console's own round-trip latency for every command now — acceptable for v1 per the async-over-synchronous decision, but a real UX question for later if it proves too slow in practice; whether a console VM should get a default subscription to COMMON-CH (skipped in v1 — direct 1:1 console↔user traffic doesn't need it, and it would require resolving an index in Hermes's own routing table for a VM that doesn't have one yet).

F.23 — WIREBIND: the real attach call site, no manual steps

Every phase closure from §F.18 through §F.22 pointed at the same gap: RUNCAP-TEST and PAIR-TEST were diagnostic scaffolding, exercising CERTVERIFY/RUNCAP/the console pair by hand, not a real attach flow. WIREBIND (§F.5, originally a breadcrumb) closes it — and turned out to be assembly, not new design: every hard decision it needed was already made in earlier scoping (§F.9's ACL-comparison rules, §F.7's cert trust model), and CERTVERIFY's own orchestration already existed (vm_identity_from_cert(), Phase A/B — verify-signature + check-serial + extract-owner-pubkey in one call) — it had just never been wired to a live attach.

The flow (capsule_wirebind_try_attach(), called from sk_repl_idle() alongside capsule_zuse_boot_try_attach(), same HOMEBLOCKS_SIG_OK hook): sig->cert_offset == 0 means this is Zuse's own genesis-mode drive (no cert region, §F.21) — that's already the other function's job, skip. Otherwise, with Zuse already authenticated this boot (nothing to verify a regular cert against otherwise — a real precondition, not an oversight), read the cert devblock(s) and call vm_identity_from_cert() against mama_vm's own zuse_cert_pubkey and the drive's drive_uuid. On success: read the drive's own user_identity_seed_t for its username, birth a console VM + RUNCAP-born user VM pair (idempotent — no-ops if that username is already live this session), install the verified VMIdentity onto the user VM, register the "<username>~user" pairing sk_repl_dispatch_line() (§F.22) looks for. Deliberately does not auto-USE the new console — that stays an explicit, BINDSTEP-gated step (§F.9), not something a bare attach should trigger silently.

Verified end-to-end, live, including a genuine negative case — not just the happy path: attached disk/user1.img (its cert signed by a different, earlier-session Zuse instance — Zuse's own keypair is freshly random every genesis mint) and got "WIREBIND: cert verification FAILED -- drive refused" — real proof the check actually verifies against the live session's own root, not a rubber stamp. Minted a fresh identity with this boot's own Zuse (S" Sam Sample" S" SamS" ... MINT), reattached, and got "WIREBIND: SamS attached and ready -- USE it to begin" printed with zero manual commands — no PAIR-TEST, no RUNCAP-TEST. Followed by a clean USE SamS and a working async WELCOME relay end to end, identical in shape to §F.22's own manual-pairing proof. Clean 3-architecture regression: Hermes/Artemis both birth live, no unexpected ACL denials or UNKNOWN WORD. Commit 6fc0ee3.

Still open: BINDSTEP itself (§F.9) — USE is still unguarded; WIREBIND installs a real VMIdentity onto the user VM but nothing re-verifies it on USE yet, so the ACL-gated re-attach check that closes this whole identity arc is the one piece left. The fence-marker persistence gap (§F.21) still means Zuse can't re-authenticate across a real reboot in this QEMU environment — worked around for this verification the same way, via a fresh genesis mint within the same boot session, not fixed here.

F.24 — BINDSTEP + the fence-persistence root cause: identity arc closed

Two items, closed together, both directly instructed: the fence-marker persistence gap (§F.21/§F.23's own "still open") and BINDSTEP (§F.9's own decisions, never wired to a real call site). Neither turned out to need new design — both were assembly and one real bug hunt.

Fence-persistence root cause, traced rather than reformatted blind: meta_fence_blocks (the field blk_meta_zone_write()'s own bounds check gates on) was carved out of what used to be unused padding in blk_volume_meta_t — the code's own comment already documented this exact risk. disk/artemis.img was formatted before that field existed, so the on-disk bytes at that offset have always read back as 0, and blk_format_or_load_disk()'s existing-volume path (magic/version match) never recomputes it — only a fresh format does. Every "fence write FAILED" message this entire session, the old zuse_cert_devblock_t flow and the new zuse_genesis_marker_t alike, traces to this one thing, not a timing bug or a QEMU quirk. Why an in-place patch was refused even though it's tempting: setting meta_fence_blocks alone without redoing the rest of the geometry risks the fence region overlapping data already allocated near the top of the volume under the old, un-fenced layout — a real, if low- probability, corruption risk for a one-field edit that looks harmless. The only provably safe fix is a genuine reformat, confirmed explicitly before doing it (discards accumulated persistent test state — regenerated fresh at next boot regardless, not real data). Verified: fence write now succeeds with no failure suffix, and the full cycle — mint once, reboot, reattach, re-authenticate — works for the first time this whole session: "Zuse: identity confirmed from attached thumbdrive", ZUSE-SESSION? goes 0-1 with no re-minting.

BINDSTEP: capsule_wirebind_verify_cert() extracted out of capsule_wirebind_try_attach() into a shared function so WIREBIND (the original attach) and BINDSTEP (every USE of an identity-locked VM) check the exact same thing the exact same way — one code path, not two that could quietly drift apart. mama_word_use() (USE) now re-verifies live, not cached (§F.9 decision 1), whenever the target has VMIdentity.installed=1: reads whatever drive is currently attached, re-verifies its cert, compares owner pubkey against the target's own installed identity, refuses on any mismatch or no drive attached. A target with installed=0 (Hera, Hermes, Artemis, any bare console VM) stays freely targetable — unchanged, per §F.9 decision 2.

Two related bugs found and fixed live while testing BINDSTEP, not assumed away: USE was Mama-only — a console-paired session (§F.22) had no way back to Hera at all, any attempt to call USE from inside a console VM hit a plain UNKNOWN WORD. Corrected directly: USE isn't console-specific, it should work VM-to-VM universally, the same way VM-EXEC already does — now registered in register_child_vm_words() too. That alone wasn't sufficient: the console relay (sk_repl_dispatch_line(), §F.22) would have captured a bare USE call and sent it to the paired user VM as a message instead of running it. Fixed with a small suffix-match guard (sk_repl_line_calls_use()) — real FORTH syntax always puts USE last, so a trailing-token check reliably recognizes it without needing a full tokenizer; it now always runs directly, never relayed.

Verified live, the full matrix, not just the happy path: USE on an unlocked VM works unconditionally; USE escaping a console back to Hera now works; USE on an identity-locked VM succeeds while its own drive is attached and is refused once detached — exact text: "USE: FinT~user refused -- no matching identity currently attached". Clean 3-architecture regression, including confirming the reformatted disk/artemis.img loads correctly as an already-recognized volume ("Artemis: LithosAnanke disk -- resuming") on aarch64 and riscv64 too, not just the amd64 boot it was reformatted under. Commit 09d78c9.

The whole identity/messaging arc opened at §F.20 is closed, end to end, for the first time: CERTVERIFYRUNCAPMINT → the console/user pair → WIREBINDBINDSTEP → persistence across a real reboot. Every phase closure from §F.18 onward flagged the next gap; none remain flagged as "still open" in this chain.

F.25 — make qemu boots with Zuse attached by default: a real xHCI driver gap, not a config gap

Directive: a plain make -f Makefile.starkernel qemu (all three arches) should always attach both the Artemis disk (already default via ARTDISK) and Zuse's own thumbdrive, landing already authenticated into the Zuse identity, unless told otherwise. Traced before touching anything: attaching -device usb-storage on the QEMU command line at launch (rather than hot-plugging post-boot via QMP, this session's established workaround since Milestone 2) looked like it should just work now that §F.21–§F.24 closed the whole attach/authenticate chain. It didn't, and the reason is a real gap in the xHCI driver, confirmed by reading xhci_poll_events() (src/starkernel/usb/xhci.c) directly: it is purely event-ring-driven, reacting only to XHCI_TRB_TYPE_PORT_STATUS_CHANGE_EVT events, with no code anywhere that proactively checks PORTSC.CCS across ports at xhci_bringup() time. A device already connected before the controller resets never generates a "change" event — nothing changed from the controller's perspective once it starts looking — so it stays invisible forever, not just late.

Fix, not workaround: extracted the existing "device connected" handling (Enable Slot kickoff, single-outstanding-connect discipline) out of xhci_poll_events()'s PORT_STATUS_CHANGE_EVT case into a shared xhci_handle_port_connected(). Added xhci_scan_ports_for_already_connected(), called once from xhci_bringup() right after the controller starts running and g_xhci_dev is latched: scans tracked ports directly for PORTSC.CCS, and drives the first connected one found through the exact same Enable Slot path a real hotplug event would — same single-device-at-a-time scope this driver already commits to elsewhere. This is the technically correct fix (a real USB host controller driver enumerates already-connected ports at bring-up, it doesn't rely solely on later hotplug events), not a QEMU-side or Makefile-side scripted delay. Verified live: "xhci: device already connected at bring-up" fires, followed by the full enable-slot → address-device → configure → BOT enumeration sequence, all before the REPL ever prints its first ok> — no post-boot QMP step needed.

Makefile wiring, and a real bug found in the first attempt: added ZUSEDISK ?= disk/zuse.img, attached unconditionally on the xhci0 bus each arch's qemu target already creates (ZUSEDISK= empties it for a bare boot). First attempt inlined $(if $(ZUSEDISK),-drive if=none,id=zusedrv,...,-device usb-storage,...,) directly at each call site — GNU make's $(if) splits its own arguments on every top-level comma, and the drive/device spec is comma-heavy, so the flags silently truncated to just -drive if=none with everything past the next comma dropped. Confirmed live via ps aux showing the broken qemu command line — no error, just a quietly wrong one. Fixed by precomputing a ZUSEDISK_QEMU_ARGS variable with a plain ifneq instead of inlining $(if).

A real consequence of the §F.24 disk/artemis.img reformat, not a new bug: with the xHCI fix and Makefile wiring both correct, the boot still didn't reach Zuse authentication. capsule_zuse_boot_try_attach() (src/starkernel/capsule/capsule_zuse_boot.c:70-71) only takes the genesis-mint path when the Artemis fence has no genesis marker and the attached drive reads HOMEBLOCKS_SIG_BLANK. disk/zuse.img still carried the real identity minted (§F.20, Apollo test rig) against the old, pre-reformat artemis.img — so it read HOMEBLOCKS_SIG_OK, not blank, while the reformatted artemis.img had no marker to confirm it against. Neither branch fired; the code correctly, silently declined to act on a mismatched pair rather than guessing. Confirmed with the user and reformatted disk/zuse.img too (same category of destructive action as the disk/artemis.img reformat — explicit confirmation obtained first), so a fresh genesis-mint can run against the current fence on next boot.

The apparent "hang" chased down, not assumed: the post-reformat genesis-mint boot appeared to sit silent for 9+ minutes at 100% CPU with no new log output — killed once on the assumption it had hung, per Captain Bob's direct pushback ("i know this machine, it's never that slow") that the slow-TCG explanation didn't hold up. Root-caused properly instead of re-guessing, via temporary console_println/log_message probes at every step (capsule_mint_identity(), capsule_zuse_boot_try_attach(), blk_meta_zone_read(), and a heartbeat_ticks()-instrumented xhci_bot_wait_for_idle()), each reverted immediately after capturing what was needed (feedback_revert_probes_after_capture.md). Two real, independent findings came out of it, neither a driver bug:

  1. capsule_mint_identity() ran a fully redundant second homeblocks_sig_check() on the same USB drive sk_repl_idle() had just checked moments earlier, doubling the wait for no reason. Fixed with a new drive_known_blank parameter (capsule_mint.h/capsule_mint.c): capsule_zuse_boot_try_attach() passes 1 (it already confirmed HOMEBLOCKS_SIG_BLANK before ever calling in); the manual MINT word (mama_forth_words.c) passes 0 since it has no pre-check to rely on and still needs the safety refusal.
  2. Dozens of orphaned tail -f logs/.../qemu-*.log processes (Makefile.starkernel's own qemu recipe backgrounds one per run and only reaps it via kill $TAILPID on a normal qemu exit) had accumulated going back to 2026-08-27, one per every kill -9 this session's own iterative testing had used to stop a run early. pkill -f "tail -n \+1 -f logs/" cleared ~60 of them. Timing before/after, same exact code path, same boot: ~200s to reach the sig-check result with the leaked processes present, 26s clean. That is the dominant real-world cost this whole investigation kept tripping over — not a kernel bug.

Residual, understood, not a code defect: even on a freshly cleaned host, run-to-run timing still varies (one clean run: 26s; a later one under load: 190s+, no error, still progressing). Traced to ordinary desktop contention on the dev machine itself — CLion (188+ accumulated CPU- minutes), its Rider.Backend, and the desktop shell were all live and competing for host CPU during the slower runs (uptime load average 1.31.6 at those moments). TCG needs steady host CPU time to advance guest virtual time at the intended rate; when the host scheduler starves QEMU in bursts, the guest's own polling/BOT-command pacing visibly stalls with no corresponding guest-side bug. Not something kernel code can or should compensate for.

Where this leaves the original directive: the xHCI initial-port-scan fix, the ZUSEDISK Makefile wiring, the redundant-check fix, and the leaked-process cleanup are all real, correct, and independently verified in isolation (each confirmed via its own live boot evidence above). A full clean end-to-end run — genesis-mint completing (MINT_OK) and landing in an authenticated zuse)ok> session — was in progress but not yet observed to completion when the session was paused to resume from a terminal; disk/zuse.img and disk/artemis.img may therefore be mid-genesis-mint rather than either cleanly blank or cleanly minted (mint writes its homeblocks_sig_t block last, specifically so an interrupted mint still reads back as blank rather than a corrupt partial identity — the next boot's own sig-check settles this safely either way, no manual disk surgery needed). Next session: rerun the full three-arch acceptance (clean qemu each) with the host otherwise idle, confirm MINT_OK and "Zuse: genesis minted onto attached thumbdrive", then decide on committing.

F.26 — the "hang" wasn't a hang: a mismatched marker/drive pair, root-caused live (2026-08-28)

Resuming §F.25: what looked like a genuine infinite loop (QEMU pegged at 100% CPU, zero new serial output for 60+ seconds) was chased down with temporary bracketing console_println/ log_message probes at every step of the attach path (capsule_zuse_boot_try_attach()capsule_mint_identity()blk_subsys_attach_device()blk_format_or_load_disk()), reverted immediately after capture per standing convention (feedback_revert_probes_after_capture.md). Two real findings, neither a code defect:

  1. log_message(LOG_INFO, ...) calls are silently swallowed once the REPL starts. kernel_main.c's own log_set_level(repl_level) defaults to LOG_WARN when no explicit --log-* kernel arg is given, and LOG_INFO > LOG_WARN fails the level check in log_message() — so every LOG_INFO-level probe placed in block_subsystem.c (a vendored, non-kernel-only file, so console.h isn't available there) printed nothing, giving a false "never reached" reading for code that was actually running fine. Confirmed by temporarily promoting the same probes to LOG_ERROR (unfiltered) — the exact same call sites all fired immediately, in order, no code changes needed. A real investigative trap, not a bug: don't trust LOG_INFO-level silence as "not reached" once the REPL is up.
  2. The actual root cause: capsule_zuse_boot_try_attach() found have_marker=1 (a zuse_genesis_marker_t already on Artemis's top-of-device fence, left over from before §F.25's mid-session reformat) but the freshly-reformatted disk/zuse.img read back HOMEBLOCKS_SIG_BLANK — a mismatched pair. Per the function's own documented design (§F.21: "Only act if THIS attach is that drive"), it correctly, silently declined rather than guessing, and returned immediately. With no re-mint and no re-authentication ever happening, the boot settled into the REPL's ordinary busy-poll idle loop (no HLT, polls console_getc() every iteration) with nothing left to log — indistinguishable from a true hang by CPU usage or serial silence alone, especially under slow TCG with no periodic heartbeat print. Traced end-to-end by bracketing probes through blk_subsys_attach_device() confirming the entire attach path, including its own final (also LOG_INFO, also filtered) summary log_message() call, completes and returns cleanly every time — the system was never stuck in kernel code at all, just correctly idling at an unauthenticated ok>.

Fix: zeroed both disk/artemis.img (30 MiB) and disk/zuse.img (64 MiB) at their existing sizes (dd if=/dev/zero + truncate to the exact original byte count) — a genuine reformat, explicitly confirmed with the user first (same destructive-action category as §F.24's), giving a matched blank pair with no stale marker and no stale volume header. Not a code change.

Verified, full three-architecture acceptance, one clean run each:

  • amd64: fresh genesis-mint — "Zuse: genesis minted onto attached thumbdrive"
  • aarch64: reload path — "Zuse: identity confirmed from attached thumbdrive" (shares the same now-minted disk images written by the amd64 run — correct, expected cross-arch behavior, not a separate mint per architecture)
  • riscv64: reload path — "Zuse: identity confirmed from attached thumbdrive", same as aarch64

Design question raised and answered, no code change needed: why not trigger a cold reboot right after MINT instead of activating live in the same boot? Two decisive reasons: every qemu target in Makefile.starkernel passes -no-reboot unconditionally, so a triggered reboot would terminate QEMU rather than cycle the kernel; and a real reboot re-runs the entire Tripod fleet birth sequence from scratch, a real cost under TCG. More fundamentally, there is no duplicate logic to reconcile: both the fresh-genesis branch and the reload branch already call the exact same install_and_activate() (capsule_zuse_boot.c:98 and :113) — they only differ in where the seed/pubkey come from (freshly generated vs. read back off disk), which is an inherent, appropriate difference, not drift risk. The live-activation path is also not optional overhead: it's the same mechanism a live (non-cold-boot) thumbdrive attach needs, so reboot-then-reload would still need it built, just as an additional path rather than the only one.

Confirmed going forward: genesis-mint is a one-time event tied to the persistent disk image files, not the kernel binary — make clean/make -f Makefile.starkernel never touches disk/*.img (grepped, confirmed: every ARTDISK/ZUSEDISK reference in the Makefile is a -drive attach flag, nothing writes to or resets either file). Rebuilding the kernel any number of times against the same, now-matched disk-image pair will keep taking the fast reload branch, not re-mint, exactly as seen on aarch64/riscv64 above.

Where this leaves the original F.25 directive: fully closed. make qemu (any arch, default invocation) now reliably attaches both disks, lands in an authenticated Zuse session, and reruns cleanly. Commit follows.

F.27 — first real second identity minted: a genuine MSG-TICK gap found and fixed (2026-08-28)

With Zuse's own genesis flow closed, minted the first real second identity end to end — a fresh disk/captain-bob.img, hot-attached via QMP (blockdev-add + device_add usb-storage,bus=xhci0.0, same precedent as §W's own hotplug testing) alongside Zuse's already- authenticated session, MINTed with a real human profile ("Captain Bob"/bob), detach/ reattach to refresh the sig cache (the known §F.19 gap, worked around the same way), recognized, WIREBIND-attached ("WIREBIND: bob attached and ready -- USE it to begin"), and switched into live with USE — full first-token verification (WELCOME printed "Minted identity -- default personality" from inside the bob session itself).

A real, previously-undiscovered bug found live, not assumed away: the instant WIREBIND made the freshly-born bob~user VM live, Hera's own per-idle-tick MSG-TICK distributed- messaging pump (repl.c's Phase C loop, "distributed messaging capsule + idle-loop pump") began erroring on it every tick, forever — UNKNOWN WORD: 'MSG-TICK' / VM-EXEC: ERROR in bob~user. Root cause: MINT's default personality content (capsule_mint.c's MINT_DEFAULT_PERSONALITY) only ever defined WELCOME — unlike capsule_console.c's own CONSOLE_IDENTITY_SRC, it never loaded common:messaging.4th, so the RUNCAP-born identity VM had no MSG-TICK word for the pump to find. Not specific to this one identity — every identity minted before this fix would hit the same infinite error-spam the moment it went live.

Fix: MINT_DEFAULT_PERSONALITY now loads common:messaging.4th and calls MSG-CD-INIT first, exactly mirroring capsule_console.c's own pattern, before defining/calling WELCOME. Verified live: re-minted the same drive after the fix (bleached first) — WIREBIND succeeded, zero MSG-TICK/UNKNOWN WORD errors over 18+ seconds of idle-loop cycling (several multiples of the pump's own cadence), USE bob and WELCOME both still work correctly. Full three-architecture regression clean (amd64 fresh Zuse reload + the fix compiled in; aarch64/ riscv64 clean Zuse reload, unaffected — the fix only touches content written at MINT time, no change to the boot path itself).


G. v2.0.0 / v2.5.0 RELEASE plan — recommendation and punch lists (2026-08-28)

Context. This is the continuation section of record for the release decision. Opening it as its own lettered section (§G) so it stands apart from the carried-forward item review (§A), the design-scope sections (§B–§E), and the built-and-verified backlog (§F). Everything before this section describes work already committed to origin/master (5689c39 and earlier). This section records the versioning split (v2.0.0 = QEMU, v2.5.0 = real hardware), the release-gate recommendation, and the punch lists to completion for each.

The versioning split, decided 2026-08-28 — now standing policy, see docs/lithosananke/ROADMAP.md "Release Versioning Policy" for the authoritative form. v2.0.0 is the QEMU release — the point where the QEMU-only story is complete and defensible on its own terms. v2.5.0 is the real-hardware release — the point where the same story transfers to bare metal on a real board. This is one instance of a general policy: X.0.0 = QEMU release, X.5.0 = hardware bare-metal release, even major numbers are LTS. The rationale for the split is practical, not arbitrary: v2.0.0 is the natural, honest cut for what exists today and is fully verifiable today (three-arch QEMU), and v2.5.0 bundles the real-hardware-only work (per-arch RNG drivers, real-board boot validation) that cannot be green-lit or verified until the boards physically exist (the author expects to have money for the SBCs — RasPi 6, Milk-V — within about a month of the v2.0.0 cut). Nothing is deferred that QEMU alone could already prove out; only what genuinely needs real silicon.

v2.0.0 was CUT 2026-08-29 (annotated tag v2.0.0, the three-arch QEMU release). With real boards arriving incrementally (Beelink SER5 in hand now; RasPi 5 + Milk-V orderable around Mon 2026-08-31), the hardware rollout is split per board in hand into v2.2.0 (amd64/SER5), v2.4.0 (aarch64/RasPi 5), v2.5.0 (all three, adds Milk-V riscv64) — each an even-minor LTS point-in-time cut on the same line. Authoritative form in ROADMAP.md "Release Versioning Policy": Board-by-board hardware rollout, decided 2026-08-29. The G.2/G.4/G.5 items below are re-mapped onto that cadence in each section's completion notes.

Everything about the state being shipped by v2.0.0 is unchanged by this versioning: the current tree is a complete, deterministic, three-architecture OS that boots UEFI under QEMU (M0M9 core milestones complete, M7.1 Capsules live, M9 Block I/O live, terminal/REPL I/O wired, real virtio-blk persistence proven, a working first-boot Zuse mint via the block fence).


v2.0.0 (QEMU) — release gates

The v2.0.0 punch list is the set of items that must be green on QEMU (all three of amd64/aarch64/riscv64) before the QEMU story is complete. Each item states its own exit criterion so closure is unambiguous and independently verifiable, per this project's standing discipline.

G.1 [v2.0.0] Stall/short-packet recovery for xHCI (scoped §F.14 — BUILT 2026-08-29)

Implemented per §F.14 exactly as designed: new STALL_ERROR completion code handling, xHCI Reset Endpoint + Set TR Dequeue Pointer commands, CLEAR_FEATURE(ENDPOINT_HALT), escalating to a full Bulk-Only Mass Storage Reset, and a recovery counter (XHCI_BOT_STALL_MAX_RECOVERIES = 2) mirroring bot_tur_retries.

The code is a purely additive recovery path gated off the non-success transfer-event branch (xhci.c transfer-event gate + xhci_handle_bulk_stall(), the two new command-completion states XHCI_CONN_AWAIT_RESET_ENDPOINT/AWAIT_SET_TR_DEQUEUE, the two new transfer purposes XHCI_XFER_CLEAR_HALT/XHCI_XFER_BOT_RESET, and the two new deferred next_action values XHCI_NEXT_ACTION_CLEAR_HALT/XHCI_NEXT_ACTION_BOT_RESET). The normal (non-stall) path is byte-for-byte unchanged — the only touched shared line was the single "control transfer failed" gate, now with an else-branch for the recoverable-STALL case.

  • Built clean on all three archsmake -f Makefile.starkernel ARCH={amd64,aarch64,riscv64} all.
  • QEMU amd64 boot regression: clean enumeration + BOT attach (READ CAPACITY10 → READ10 → home-blocks) completing with zero stalls and zero failures; normal-path xHCI trace identical to the pre-change baseline. Repeated across fresh boots (all attach=SUCCESS).
  • Exit criterion — partial, honest: the recovery path itself is not exercised live because QEMU's qemu-xhci model does not produce a STALL_ERROR for the injected bad transfers this driver can stage (no real hardware stall behavior). The design (§F.14) is reviewed and the code paths are structured and reviewed; live STALL_ERROR → recover → retry verification is deferred to hardware (v2.5.0/Artemis bare-metal), where a real bad transfer can be staged. This is the last QEMU-verifiable storage-integrity gap and the recovery logic is in place; the one thing QEMU cannot prove is the live stall injection itself.
  • Follow-up (pre-existing, NOT G.1): ROOT-CAUSED and FIXED 2026-08-29. During G.1 verification an intermittent boot-time attach race was observed (the sk_repl_idle() bot_msc_attach_pending handoff occasionally does not progress on a cold QEMU boot, independent of source, with baseline HEAD exhibiting it too). Unrelated to G.1. Root cause: xhci_poll_events()'s event-ring drain loop had no hard ceiling — ERDP is written back only when the loop exits, so the controller cannot reclaim event TRBs mid-drain, and on pathological controller behavior the head can chase the software dequeue pointer indefinitely, causing the drain (and hence the attach handoff) to livelock. Fixed by bounding the drain to one full ring (XHCI_EVT_RING_MAX_DRAIN, xhci.c xhci_poll_events()); the loop always terminates and always writes ERDP each call, and unprocessed events keep their cycle bit and are re-read next poll (nothing is dropped). Verified across six consecutive fresh boots.

G.2 [v2.0.0] Real-hardware RNG driver plumbing, QEMU-verifiable slice (rest of it lands at v2.5.0)

The full real per-arch RNG driver (rng_get_bytes() backing off virtio-rng to RDRAND / Zkr / peripheral-RNG) is a v2.5.0 item because it cannot be verified on hardware that does not exist yet. What is a v2.0.0 gate is the shape of the driver — the unified rng_get_bytes() entry point, the probe-and-refuse-loudly contract (never silently degrading to a deterministic seed — the exact failure Phase A existed to avoid), and the QEMU path staying on virtio-rng unchanged — all taken up before v2.0.0 so the single entropy entry point is stable and the v2.5.0 work is purely per-arch backends, not a redesign of the call path.

  • Exit criterion (v2.0.0 slice): rng_get_bytes() exists as the single entropy entry point, virtio-rng is its sole active backend on all three QEMU arches, the code refuses loudly (boot-time message, safe fallback to this-boot-only cert) if no backend is present, and the pre-existing Zuse mint/cert behavior is unchanged on all three arches. The per-arch backends themselves are parked, explicitly, for v2.5.0.
  • v2.0.0 slice — BUILT and VERIFIED 2026-08-29. include/starkernel/rng.h + src/starkernel/rng/rng.c provide the single rng_get_bytes() entry point; rng_init() probes the backend set (v2.0.0: virtio-rng only) and, on no backend, prints a loud boot-time warning and rng_get_bytes() returns RNG_ERR_NO_BACKEND — it never silently degrades to a deterministic seed. The backend-selection switch in rng.c is the exact seam the v2.5.0 real per-arch drivers (amd64 RDRAND, riscv64 Zkr, aarch64 peripheral RNG) plug into without touching the call path. Kernel consumers no longer touch virtio_rng_* directly: capsule_mint.c (identity seed + drive_uuid) and kernel_main.c phase 8 (rng_init()) route through the unified layer; virtio_rng.c stays as the sole backend. Built clean on all three arches (make -f Makefile.starkernel ARCH={amd64,aarch64,riscv64}). QEMU amd64 boot: POST 1012/0/0 + ok>, rng: backend = virtio-rng + entropy: ready printed by the unified layer, and Zuse attach/identity confirmed from the thumbdrive — Zuse mint/cert behavior unchanged.
  • Exit criterion (v2.5.0 completion, for reference): on each real board rng_get_bytes() returns genuinely non-deterministic bytes (two boots differ) and the Zuse mint path seeded from it produces a valid distinct cert per boot when bleached.

G.3 [v2.0.0] NVRAM persistence decision — DECIDED 2026-08-28: permanently de-scoped

The NVRAM SetVariable crash is root-caused (this kernel's VMM never maps the flash-window region OVMF's variable service needs; also the exposed-seed design defect) and, since the Zuse cert now lives on the block fence on Artemis, functionally superseded. DECIDED (Captain Bob, 2026-08-28): de-scope NVRAM-backed identity permanently. The block fence on Artemis is the substrate of record for Zuse's cert. The alternative — scoping the real VMM memory-map work to make OVMF's flash window writable — is not being scheduled: it is large, unscoped VMM work with no live consumer now that the fence carries the cert, and the NVRAM approach carried a real exposed-seed security defect regardless of the crash.

  • Exit criterion (MET): a dated decision recorded in this section (de-scope), and the ROADMAP/backlog updated to match. Decision required, not necessarily work — closed, no work required beyond this record.

v2.5.0 (real hardware) — release gates

These are the real-hardware-only items. None can be verified until the boards exist (RasPi 6, Milk-V; SER5 either in hand or not — see G.5). They are the entire v2.5.0 gate and are listed together so the v2.5.0 scope is explicit and closed.

G.4 [v2.5.0] Per-arch RNG backends behind rng_get_bytes() (the G.2 body)

The real RNG driver backends, completed behind the v2.0.0-formed entry point. The existing VM-UUID finding already maps the primitives: amd64/SER5 RDRAND, riscv64/Milk-V Zkr (RNDR), aarch64/RasPi peripheral RNG. On QEMU all three arches stay on virtio-rng; nothing changes there.

  • Exit criterion: on each real board (rng_get_bytes() falling back through the real primitive) returns genuinely non-deterministic bytes (two boots differ), the Zuse mint path seeded from it produces a valid distinct cert per boot when bleached, and QEMU behavior is unchanged.
  • v2.0.1 — amd64 RDRAND backend BUILT + QEMU-verified 2026-08-29 (SER5 slice). The first real per-arch backend lands in src/starkernel/rng/rng.c, behind the v2.0.0 unified entry point, #if-guarded to amd64: CPUID.01H:ECX[30] detection + inline-asm rdrand draws feeding rdrand_fill() (whole-byte emission, partial draw discarded — throwing away entropy is always safe). Probe order honors the policy: virtio-rng is tried first, so the QEMU path stays on virtio-rng unchanged; RDRAND is the fallback that only real hardware (which has no virtio-rng) reaches. QEMU-verified both ways on amd64: with virtio-rng present → rng: backend = virtio-rng (unchanged); with virtio-rng absent and RDRAND exposed (-cpu max) → rng: backend = rdrand + entropy: ready + Zuse attach confirmed. rdrand_fill()'s exact logic host-proven to fill 32-byte/16-byte buffers and produce differing draws run-to-run (non-deterministic). aarch64/riscv64 builds unaffected (guarded off). Still parked for their real boards: riscv64 Zkr (RNDR), aarch64 peripheral RNG. Next: prove RDRAND live on the real SER5 (v2.2.0), then add Zkr (v2.5.0/Milk-V) and the aarch64 peripheral RNG (v2.4.0/RasPi 5) as those boards reach v2.2.0-level closure.

G.5 [v2.2.0 / v2.4.0 / v2.5.0] Real-machine boot validation (SER5 / RasPi 5 / Milk-V)

Boot the generic thumbdrive image (make -f Makefile.starkernel ARCH=<arch> thumbdrive) on each real board in hand, confirming the same acceptance story QEMU keeps green: POST 1012/0/0 + ok>, block-fence Zuse load true to the already-minted Artemis image, and the corresponding G.4 RNG backend working live. Mapping onto the board-by-board cadence (decided 2026-08-29): v2.2.0 = SER5 (amd64, in hand now — two 16 GB sticks available, one to boot the image, one to mint a real Zuse identity on real hardware); v2.4.0 = RasPi 5 (aarch64, orderable around Mon 2026-08-31); v2.5.0 = adds Milk-V (riscv64), the full three-arch bare-metal cut. Each board that boots is a separate, recorded data point, and each milestone closes only when its board's items are proven live (not build-only).

  • Exit criterion (per milestone): the board cold-boots to ok> with Arena conservation (43691/21845/65536), Zuse cert loads from the fence (or mints fresh on bleached media), and its G.4 RNG backend returns non-deterministic bytes live. A milestone does not close on any of its board's G.4/G.x items being build-only.

G.6 [v2.0.1] Generic UEFI-bootable thumbdrive image (SER5 path — BUILT + VERIFIED 2026-08-29)

First concrete v2.0.1/SER5 work item, taken up because the Beelink SER5 is in hand and a "generic thumbdrive bootable OS" is the target. The new make -f Makefile.starkernel thumbdrive goal builds a generic GPT/FAT32 disk image (one "STARKERNEL" EFI System partition with EFI/BOOT/BOOT<ARCH>.EFI + startup.nsh) writable to a USB stick with dd and bootable on any real amd64 UEFI firmware as well as under QEMU. The monolithic loader embeds the entire kernel, so the ESP needs only the UEFI fallback boot path — this is what makes the image "as generic as possible."

  • Verified by attaching the image to QEMU as a USB mass-storage device (not cdrom): OVMF BDS auto-selected Boot0002 "UEFI QEMU QEMU USB HARDDRIVE", then the kernel booted normally — v2.0.0 banner, POST 1012/0/0 + ok>, rng: backend = virtio-rng + entropy: ready, Artemis ready, and the USB BOT/xHCI storage path enumerated (READ CAPACITY10 → MSC device ready → READ10 CSW PASS). This mirrors the real SER5 flow: firmware boots the thumbdrive's ESP, and the OS's own storage rides the same USB BOT/xHCI path. Works for all three arches via the arch-mapped EFI boot name.
  • Next (v2.0.1→v2.2.0): G.5 real-board proof on the SER5 (in hand) — boot the thumbdrive image, live RDRAND entropy, real Zuse mint on a second stick — then extend per arch to v2.4.0 (RasPi 5) and v2.5.0 (Milk-V). Authoritative cadence in ROADMAP.md "Board-by-board hardware rollout".

Post-release tracking (not gating either milestone)

Beyond the v2.0.0/v2.5.0 gates, the carried-forward open items (§A) and open design questions (§D) remain genuinely open. Not release-gating; tracked here so nothing silently ages and so the punch lists above explicitly delimit what is excluded from each milestone:

  • Block-map on-drive format field edit (scoped §F.4) and associated §F.6/§F.11 allocation.
  • CA-signed-cert verification path + DER signature extraction (scoped §F.7) and CERTVERIFY/BINDSTEP consumers (§F.8/§F.9).
  • Migration state machine design (§A item, §B Stadium unification reframe).
  • Console umbrella closure (§A 4.3) and (user) prompt segment (§A 4.4s) once 4.3's blocker clears.
  • DETACH/unclean-removal wiring into sk_repl_idle() (scoped §F.10).
  • Dirty-event granularity (1.11, blocked on 4.3) and the §17.4 framebuffer heat/decay design.
  • Re-run the DoE on the new substrate (§A 5.1).
  • First-touch allocation first-call-free.

Beyond v2.5.0 — Zynq FPGA, then a coloring-in period, decided 2026-08-29. The next big milestone after the three-board bare-metal cut (v2.5.0) is transferring the story to a Zynq FPGA (configurable silicon — soft/hard cores, PL fabric, non-standard memory map) — a genuinely larger step than any prior board, expected on a new even-major line. Between v2.5.0 and starting the Zynq sits a "coloring in" hardening phase: making v2.5.0's real-hardware story production-honest (USB BOT/xHCI + block robustness; live-entropy and Zuse-cert hardening on real ASICs; SMP/multi-core + IRQ routing from the HAL notes; driver breadth) so the FPGA carries a thickened, not thin, shape forward.

The FPGA creates a three-product split, decided 2026-08-29. The Zynq is the pivot that forces the project into three distinct products, each with its own host, delivery, and proof character: (1) Hosted StarForth — the existing hosted/interpreted StarForth runtime; (2) A full StarshipOS — the standalone OS on general silicon (LithosAnanke → StarshipOS); (3) Hardware steady-state machinery with sealed executions, HOL-proven — the FPGA-native product: hardware-enforced sealed executions and steady-state machinery machine-checked in a proof assistant (HOL); delivery is the bitstream + HOL proof artifacts. Product 3 is born on the FPGA and its existence is what cleanly separates 1 from 2 from 3. Full form in ROADMAP.md "Release Versioning Policy → The FPGA creates a three-product split".

IP framing, decided 2026-08-29. The three products are treated as three patent applications with a matching mark each: Patent 1 + StarForth (hosted runtime); Patent 2 + StarshipOS + LithosAnanke (full standalone OS); Patent 3 + Compudynamics (Zynq steady-state/sealed-execution HOL-proven hardware, anchored by the physics-adaptive runtime). Open for reconciliation with counsel: the repo already has a "Patent pending" USPTO provisional (filed Dec 2025) for the Compudynamics physics-adaptive runtime (docs/patent/); patent 3 and the three product applications must be positioned relative to it (continuation? additive?) — recorded as intent here, nothing filed or drafted. Driving deadlines (clarified 2026-08-29): #1's provisional will expire before conversion (filed ~Dec 2025, window ~Dec 2026 — priority lost if unconverted; the subject matter can't be re-staked by refiling after lapse), so #1 is on a hard clock; #2 (full StarshipOS) is the December deliverable that keeps the flagship covered — and, pending counsel confirmation, may be the conversion vehicle for #1. Authoritative form in ROADMAP.md "IP framing of the three products".