Files
LithosAnanake/docs/working/architecture/VM-FLEET-ATTRACTOR-DESIGN-20260705.md

2066 lines
121 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# VM Fleet Attractor Experiment — Design Doc
**Date:** 2026-07-05 (rev g)
**Branch:** `lithosananke`
**Status:** Phase 1 and Phase 2 implemented and verified three-arch. Rev f
found the "100% heat at Hera" result from revs b/d/e was never an attractor
finding — a seeding bug left the redistribution mechanism completely inert.
Rev g fixed that seed and a second bug it exposed (wrong time-source in
`vm_physics_touch`'s call sites), both verified amd64-clean, but found a
third, deeper, pre-existing bug underneath in `src/starkernel/hal/
host_services.c`'s `kernel_monotonic_ns()` — it doesn't handle the
RELATIVE-timer-mode fallback the rest of the kernel already uses when TSC
calibration fails (as it does under this QEMU/TCG environment). **Not yet
fixed** — likely affects word-level background decay too, not just VM-fleet
physics. Phase 3 stays blocked until heat can actually be observed moving.
**Author:** Captain Bob / Claude Code
---
## Research question
Does the VM fleet — Hera, Hermes, Artemis, and any future children — find an
attractor the same way individual FORTH words do?
This is not a rhetorical question. It's directly motivated by the
**L8 Attractor Map campaign** (`docs/working/experiments/campaigns/l8_attractor_map/SESSION_REPORT_2025-12-10.md`,
180 runs, 6 workloads × 30 replicates): word-level execution heat, under the
existing physics engine (Loops #1#7), was empirically observed to converge
deterministically to a steady-state "coldest" configuration — low
coefficient-of-variation, workload-independent convergence time
(23.3 ± 2.61 ticks), and a candidate universal oscillation frequency
(ω₀ ≈ 13.5 Hz, later suspected to be CPU-clock-linked rather than a true
constant). That finding, plus the independent 38,400-run 2^7 factorial DoE,
is the empirical basis for the whole "physics-grounded adaptive runtime"
claim this project is built on.
`capsule_vm_physics.c` (added this session, see
`VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md`) is a deliberate structural
mirror of that same mechanism — per-entity execution heat, a rolling
touch-history window, regression-inferred slope — applied to VMs instead of
words. The mirroring was designed in, but never tested for the property that
motivated the original mechanism: does it converge? Does the fleet's heat
distribution settle into a repeatable steady state, the way word heat did?
Nobody knows yet. This doc is step one toward finding out.
## Structural note: the shared-slope question is not a blocker
An earlier framing of this question (see prior conversation) suggested that
`fleet_transfer_slope_q48` being a single fleet-wide value, rather than
per-VM, might mean the mechanism *can't* produce genuine per-VM attractors.
That framing was checked against the word-level mechanism and doesn't hold:
`decay_slope_q48` (`include/vm.h:504`) is likewise a single global value
applied uniformly to every word's individually-varying execution heat — one
shared decay rate, many independently-converging heat values. That's exactly
the structure the L8 campaign found attractor behavior in. So the VM-fleet
mechanism has the same shape, and the question is genuinely open rather than
foreclosed by the shared slope — worth testing, not worth assuming either way.
## Design goal: transparency first
Per explicit direction: **transparency is a design goal of this system, not
an incidental nice-to-have.** Before any experiment can answer "does the
fleet find an attractor," the mechanism has to be fully observable —
individual VM state, not just an aggregate pass/fail. Today it is not:
`vm_physics_status()` (`capsule_vm_physics.c`) prints exactly four things:
fleet heat sum, conservation verdict (CONSERVED/DRIFTED), the one shared
`fleet_transfer_slope_q48`, its fit quality, and whether the touch-history
window is warm. **No per-VM heat is ever surfaced.** You cannot currently
ask "what is Hermes's heat right now" from FORTH or from a boot log — only
"what is the fleet's total" (which conservation forces to Q.1 always, so it
carries no information about *distribution*).
This is the opposite of transparent. A conservation check that always reads
"1.0" tells you nothing about whether the fleet is balanced 33/33/33 or
99/0.5/0.5 among Hera/Hermes/Artemis. If the goal is to observe attractor
behavior — or, just as importantly, to *rule it out* — the system has to
expose per-VM heat as a first-class, always-available fact, not something
inferred indirectly or reconstructed after the fact from birth/kill/touch
parity logs.
**Concretely, "transparent" means, at minimum:**
1. Any live VM's current `execution_heat_q48` is readable by ID or by name,
at any time, without needing to halt or inspect via debugger.
2. The fleet-wide state (`fleet_transfer_slope_q48`, fit quality, window
warmth) remains visible as it is today — this doc adds to that, doesn't
replace it.
3. A time-series of the above, not just an instantaneous snapshot, since
"attractor" is a claim about behavior *over ticks*, not a single reading.
4. The instrumentation itself must be a passive observer, per the existing
VM-fleet-physics design principle — recording and reporting only,
changing nothing about when or how VMs actually execute.
## Current instrumentation gap (concrete)
| What's needed | What exists today |
|---|---|
| Per-VM heat readable by ID/name | Not exposed. `VMPhysics.execution_heat_q48` lives inside the kmalloc-backed registry linked list, private to `capsule_vm_physics.c`. |
| Time-series export (per-tick, per-VM) | Nothing. The hosted VM's `--doe` CSV pipeline (`doe_metrics.c`, `heartbeat_export.c`) has no kernel-build equivalent — bare-metal experiments stream CSV-shaped rows to serial from FORTH capsules instead (see `experiments/bare_metal/`), and no such capsule exists yet for VM-fleet physics. |
| A controllable experimental factor | The mechanism is a passive observer by design — it never decides who runs. The *experiment* needs something else (a driving capsule) to vary dispatch patterns across Hera/Hermes/Artemis in a controlled way; no such capsule exists (the old `doe-campaign.4th`, which attempted something adjacent, is broken and being superseded by this doc — see below). |
| Analysis pipeline | The L8 campaign's `l8_analysis_safe.R`-style pipeline (CV, convergence time, phase-space portraits) is workload/word-shaped, not VM-fleet-shaped. Would need a new script, though the statistical *methodology* — CV over time, convergence-tick counting, ANOVA across factor levels — transfers directly. |
## Relationship to `doe-campaign.4th`
This design doc supersedes migrating `doe-campaign.4th` as originally scoped.
That capsule's model — manually setting VM heat via `VM-HEAT!` and manually
pumping ticks via `K-BUMP` — has no equivalent under the current mechanism,
which deliberately has no manual heat-injection point (heat only moves in
response to real `VM-EXEC`/`VM-CALL`/`VM-STEP` dispatch, by design, to keep
the physics a passive observer rather than a puppet). Reviving the old
capsule's *mechanics* would mean re-adding exactly the kind of synthetic
heat-injection surface the current design deliberately removed. The old
capsule's *goal* — drive the fleet through controlled scenarios and observe
the result — is exactly this experiment's goal, so its ideas carry forward
here; its code does not. `doe-campaign.4th` itself should be deleted once
this experiment has a working replacement, not before.
## Rev b: minimum step implemented, and a first real observation
Per open question #4, the smallest useful step: `vm_physics_heat_of(uint32_t
vm_id)` (`capsule_vm_physics.c`/`.h`) and an extension to the existing
`vm_physics_status()` that walks the live-VM list and prints
`vm_id=N name=X heat_q48=Q` for each, using `capsule_vm_registry_get()` for
the name. No new FORTH word needed — `VM-PHYSICS-STATUS` already existed and
now carries the per-VM breakdown automatically.
Verified by boot-injecting `VM-PHYSICS-STATUS` at the `ok>` prompt via the
serial socket (`socat - UNIX-CONNECT:<sock>`, same mechanism the Makefile's
`DOE_INJECT` uses for `EXEC-DOE`) right after a normal TRIPOD-TEST run. Output:
```
[Hera] VM-PHYSICS: fleet_heat_sum=65536
[Hera] VM-PHYSICS: conserved=CONSERVED
[Hera] VM-PHYSICS: fleet_transfer_slope_q48=0
[Hera] VM-PHYSICS: slope_fit_quality_q48=0
[Hera] VM-PHYSICS: fleet_window_warm=YES
[Hera] VM-PHYSICS: vm_id=3 name=Hermes heat_q48=0
[Hera] VM-PHYSICS: vm_id=1 name=Artemis heat_q48=0
[Hera] VM-PHYSICS: vm_id=0 name=Hera heat_q48=65536
```
**This is already a real data point, not just a smoke test.** At the end of
a normal boot (after `PASS: fleet K`, `PASS: Hermes liveness`, `PASS: Artemis
ready`, the Hermes kill/rebirth soak, and `PASS: E2E msg flow`), 100% of
fleet heat sits at Hera — the root — and both children read exactly zero.
That's one snapshot, not a trajectory, so it doesn't answer the research
question by itself, but it's suggestive: if this holds up as the fleet's
actual resting state rather than an artifact of a short boot sequence, the
"attractor" this mechanism finds may not look like the word-level one
(distributed heat settling into a stable per-word split) — it may look like
total heat drift back to the structural root every time activity quiets
down, which would itself be worth understanding (`vm_physics_retire`'s
parent-pointer walk to Hera on kill is one obvious contributor; whether
`vm_physics_touch`'s pull-toward-the-touched-VM mechanic ever produces a
*sustained* non-root distribution during activity, only draining afterward,
is exactly what a time-series would show and a single snapshot can't).
`fleet_window_warm=YES` but `fleet_transfer_slope_q48=0` in the same
snapshot is also worth noting for rev c: the touch-history window has
enough samples to be "warm," yet the inferred slope is still zero. Worth
checking whether that's expected (e.g. inference hasn't been triggered by a
`vm_physics_tick` call recently enough) or a sign the inference path needs
its own scrutiny before trusting slope-based conclusions later.
## Experiment roadmap (rev c)
Four phases, each a prerequisite for the next. Escalate only as far as
needed — do not build phase N+1 until phase N's result says it's necessary.
### Phase 1 — Readiness handshake (existing primitives only)
**Goal:** the simplest possible rendezvous, as a clean baseline before any
workload variation is introduced. No new broadcast machinery — this phase
deliberately uses only what already exists (`MSG-SEND`, `MSG-DELIVER-ALL`,
`HERMES-TICK`, `MSG-ACK-LAST`), per the finding above that real multi-member
delivery doesn't exist yet and shouldn't be built just for this.
**Mechanism:** after each VM's own `CD-INIT` completes, it sends a `READY`
message (new event-code constant, alongside the existing `SPAWN-EVENT`/
`PAUSE-EVENT`/`RESUME-EVENT`/`KILL-EVENT` in `hermes/init.4th` block 4100).
The receiving side (whichever VM hosts the message arena — needs
confirming against the actual capsule wiring when implementation starts,
not assumed here) drains the queue via repeated `HERMES-TICK`/
`MSG-DELIVER-ALL` passes and counts distinct `READY` senders. Once all
three are accounted for, it sends each an ACK back via the existing
`MSG-ACK-LAST` pattern (`common:msg.4th`'s `HERMES-ACK`).
**Observation:** snapshot `VM-PHYSICS-STATUS` immediately after all three
ACKs land. This is the cleanest possible baseline reading — no workload
variation, no DoE — and directly extends the rev b observation (which was
taken after a full TRIPOD-TEST run, not a clean handshake). Compare the two:
does heat still end up 100% at Hera after *just* a handshake, or was rev
b's observation an artifact of everything else TRIPOD-TEST does (the K
soak kill/rebirth in particular)?
**Explicitly out of scope for phase 1:** any workload, any DoE, any
multi-member broadcast. Point-to-point messages and a counter only.
### Phase 2 — Real broadcast delivery
**Goal:** implement actual multi-member channel delivery, consuming the
existing but currently-unused `CH-ADD-MBR`/`CH-MBRS`/`MBR-VM@` scaffolding
in `hermes/init.4th`. A message sent to a channel should reach every
member, not just one `to` VM.
**Why after phase 1, not before:** phase 1 proves the simpler point-to-point
handshake works and gives a clean baseline reading before adding new
delivery machinery to the mix. If phase 1's baseline is already confusing
or unexpected, that's a signal to understand *before* introducing a second
new mechanism on top of it.
**Not yet designed:** fan-out delivery semantics (does a channel message
get one arena slot copied N times, or N independent slots?), how
`MSG-DELIVER-ALL`'s scan loop changes to walk `CH-MBRS` per channel message
instead of a single `to` field, and whether `COMMON-CH`'s existing
heat-floor-only role expands or stays separate from its new delivery role.
This needs its own design pass when phase 1 is done and verified — not
specified further here.
### Phase 3 — DoE small: one config per VM per run
**Goal:** the actual attractor question, at the smallest scale that can
plausibly answer it.
**Factor:** `doe.4th`'s existing 2⁴ factorial (`CFG-ENT`/`CFG-CV`/`CFG-TMP`/
`CFG-STB`, `CURR-CFG` 015 via `APPLY-CFG`) — the "16 workloads." Confirmed
choice over the 9 separate `init-*.4th` capsule files, which are alternate
*Mama personalities* (mutually exclusive, one per boot) and structurally
wrong for "assign each of 3 co-existing VMs its own workload."
**Design:** each run assigns one of the 16 configs to each of Hera/Hermes/
Artemis independently (with or without replacement across the 3 — TBD when
implementation starts), shuffled/replicated across many runs — the same
scale and randomization-for-temporal-bias-elimination approach as the
180-run L8 campaign. **Not** the full 16³ cross product (see Phase 4).
Response variable: per-VM heat trajectory over the run (via the phase-1/
rev-b transparency primitives), watched for CV convergence, stabilization
time, anything resembling the word-level campaign's findings — exact
statistical treatment TBD, likely adapting `l8_analysis_safe.R`'s approach
(CV by group, ANOVA, phase-space portraits) rather than inventing new
methodology.
**Escalation criterion:** if Phase 3's results are inconclusive — no clear
convergence signal, too much noise to distinguish configurations, etc. —
proceed to Phase 4. If Phase 3 already shows something as clean as the
word-level campaign's findings, Phase 4 may not be needed at all.
### Phase 4 — DoE large: full 16×16×16 cross product (only if Phase 3 is inconclusive)
Every combination of (Hera's config, Hermes's config, Artemis's config),
4,096 distinct combinations. At 30 replicates each (the L8 campaign's
convention) that's ~123,000 runs — a real undertaking, not a first pass.
Explicitly gated on Phase 3 not being sufficient; not scheduled otherwise.
---
## Open questions (carried forward, narrowed by the roadmap above)
1. ~~**Driving mechanism.**~~ Resolved by the roadmap: Phase 1 uses existing
point-to-point messaging; Phase 2 adds real broadcast; Phases 3/4 drive
workload via `doe.4th`'s existing `APPLY-CFG`. Remaining detail: which
VM hosts the message arena / where `READY` gets received in Phase 1 —
deferred to implementation time, not blocking this doc.
2. **What counts as "found an attractor" for a VM fleet?** Still open.
Phase 1's clean-baseline reading and Phase 3's trajectory data are both
needed before this can be answered rather than guessed at.
3. **Where does this live?** Still open — `docs/working/experiments/campaigns/`
following the `l8_attractor_map` precedent is the likely answer, but no
campaign directory has been created yet since no phase has code yet.
4. ~~**Minimum instrumentation vs. full pipeline.**~~ Resolved in rev b:
smallest step first. The four-phase roadmap above is the same principle
applied one level up — each phase is itself the smallest step toward
the next.
---
## Rev d: Phase 1 implemented and verified (three-arch)
**Implementation.** Hermes gained `READY-EVENT`/`READY-COUNT`/`NOTE-READY`/
`READY-ALL?`/`ENQUEUE-READY`/`HERMES-ANNOUNCE-READY`/`READY-ACK` (block 4118,
previously empty). Artemis gained `ARTEMIS-ANNOUNCE-READY`/`ARTEMIS-READY-ACK`
(new block 4129, previously unassigned) — since Artemis has no message arena
of its own, its announcement is a real `MSG-SEND` enqueued via a single-level
`VM-EXEC` call into Hermes (`S" 2 ENQUEUE-READY" S" Hermes" VM-EXEC`), not a
nested command string. `init.4th` gained `READINESS-HANDSHAKE` (new block
2053, defined before its call site per file-order execution — same pattern
`BOOT-BANNER`'s block 2057 already uses) and now calls it right after
`BOOT-BANNER`, before `TRIPOD-TEST` runs.
**A sequencing constraint the doc's rev-c text didn't anticipate:** Artemis
is birthed before Hermes in the existing boot order, so it can't announce
its own readiness at the tail of its own `CD-INIT` — Hermes (the message
host) doesn't exist yet at that point. Resolved by having Hera trigger all
three announcements explicitly, once both children are alive, rather than
each VM self-announcing inline in its own `CD-INIT`.
**Verified three-arch, first try:** amd64/aarch64/riscv64 all show
`Hermes: ready-ack`, `Artemis: ready-ack`, `PASS: readiness handshake`, then
every existing TRIPOD-TEST gate (`fleet K`, `Hermes liveness`, `Artemis
ready`, `reap`, `K soak`, `E2E msg flow`) — zero `FAIL` lines, identical
`dict_hash` (`0xcc590996ecda7654`) across all three.
**The observation holds up, and gets cleaner.** The `VM-PHYSICS-STATUS`
snapshot taken immediately after the handshake — before `TRIPOD-TEST`, before
any K-soak kill/rebirth — shows the *same* pattern as rev b's post-boot
reading: 100% of fleet heat at Hera, both children at exactly zero.
`fleet_window_warm=NO` this time (vs. `YES` in rev b), consistent with this
being a much shorter, lower-activity sequence. This rules out the rev-b
hypothesis that the all-heat-at-root reading was an artifact of TRIPOD-TEST's
kill/rebirth soak — it isn't. Even the simplest possible sequence (birth two
children, exchange three point-to-point messages, done) settles heat
entirely at the structural root. Two independent readings now agree; this
looks like the mechanism's actual resting behavior, not noise.
Next: Phase 2 (real broadcast delivery), still not started.
## Rev e: Phase 2 implemented and verified (three-arch)
**Resolves the two "not yet designed" questions from rev c's Phase 2
section.** Fan-out semantics: **N independent message slots**, not one
slot copied/shared — `MSG-BROADCAST` is sugar that walks `CH-MBRS` at send
time and calls the existing single-recipient `MSG-SEND` once per member.
This means `MSG-DELIVER`/`MSG-DELIVER-ALL` needed **zero changes** — every
fanned-out message still has a normal single `to` field, so the entire
existing delivery/ack/type-state machinery works unmodified. `COMMON-CH`'s
existing heat-floor role stays untouched; `CH-ADD-MBR` only touches
`CH-MBRS`, never `CH-HEAT`.
**Implementation.** Hermes gains `MSG-BROADCAST ( type from paddr plen ch -- )`
(new block 4151 — walks the channel's member list via `CH-MBRS@`/`MBR-NEXT@`,
calling `MSG-SEND` once per `MBR-VM@`) and `REGISTER-COMMON-MEMBERS`/
`BCAST-RECV`/`SEND-BROADCAST-TEST` (new block 4152). Artemis and Hera each
get their own `BCAST-GOT`/`BCAST-RECV` pair (own block 4129 addition;
`init.4th` new block 2054) — same word names, independent per-VM state,
mirroring how `NOTE-READY`/`READY-COUNT` worked in Phase 1. `init.4th` gains
`BROADCAST-TEST`, called right after `READINESS-HANDSHAKE`: registers all
three VMs as `COMMON-CH` members, triggers one broadcast from Hermes, drains
via `HERMES-TICK`, then checks each VM's own `BCAST-GOT` counter via
`VM-CALL` to confirm genuine 3-way delivery (not just "a message was sent").
**Verified three-arch, first try:** amd64/aarch64/riscv64 all show
`PASS: broadcast reached all 3` immediately after the Phase 1 handshake,
then every existing TRIPOD-TEST gate — zero `FAIL`, identical `dict_hash`
(`0x23c8d067ec5709b6`) across all three, no crashes or hangs on any
architecture.
**Third independent reading, same result.** The `VM-PHYSICS-STATUS`
snapshot after the broadcast — now after *two* real message exchanges
(3 point-to-point readiness messages, then 3 more from the broadcast
fan-out; 6 total) — shows the identical pattern as revs b and d: 100% of
fleet heat at Hera, both children at exactly zero. Three independent
readings, three different activity levels, same resting state every time.
This is no longer "suggestive" — it's a reproducible finding: whatever
activity the fleet does, heat returns entirely to the structural root
between observations. Whether that's `vm_physics_touch`'s pull-toward-the-
touched-VM never producing a *sustained* distribution, or something else,
is squarely a question for Phase 3's trajectory data (a single post-hoc
snapshot, however many times repeated, still can't show what happens
*during* activity — only what's left after it settles).
Next: Phase 3 (DoE small — one of `doe.4th`'s 16 configs per VM per run),
not yet started.
## Rev f: the time-series gap was already closed, and it found a real bug
Rev a's instrumentation-gap table claimed no kernel-build time-series
export existed. **That was wrong.** `src/starkernel/doe_log.c` already
streams a genuine per-tick (100Hz), 15-column CSV row to serial on every
single boot — this is what has been producing the "N data rows" line at
the end of every acceptance run all session (`experiments/bare_metal/runs/
doe-*.csv`). It's word-level dictionary metrics only (hot word count, avg
word heat, window width, etc.) — no per-VM fleet heat. Extended it to 18
columns, adding `hera_heat_q48`/`hermes_heat_q48`/`artemis_heat_q48`,
looked up by name each tick via the existing `vm_physics_heat_of()`
accessor (robust to vm_id changes across kill/rebirth). Fixed to the three
known Tripod VMs — a logging-schema choice, not a change to the physics
mechanism itself, which stays VM-count-agnostic.
**This immediately surfaced why all three prior readings agreed: the
mechanism has never once activated.** All 1797 ticks across a full ~18s
boot (handshake, broadcast, TRIPOD-TEST, K-soak, E2E messaging — every
phase run so far) show the *exact same three values on every single tick*:
`hera=65536, hermes=0, artemis=0`. Not "returns to this" — never deviates.
**Root cause, traced to `capsule_vm_physics.c` (this session's own code):**
`fleet_transfer_slope_q48` is seeded at `0` (`capsule_vm_physics.c:82`).
`vm_physics_touch`'s transfer amount is `elapsed_ns * slope >> 16` — at
slope=0, every touch moves exactly zero heat, permanently. Meanwhile
`vm_physics_tick`'s regression is supposed to *infer* a better slope from
the touched VMs' heat trajectory, but those VMs' heat has been frozen at
zero by the very fact that slope=0 — and zero-heat samples are skipped
outright in the log-linear regression (`if (trajectory[i]==0) continue`).
Closed loop: zero slope → touched VMs never receive heat → regression sees
no signal to fit → slope stays zero. It cannot bootstrap itself out of the
starting condition. The word-level mechanism this was mirrored from avoids
exactly this by seeding `decay_slope_q48` at a nonzero default (2:1 =
131072, `include/vm.h:504`) — the one place this session's mirroring
wasn't faithful to the pattern it was copying.
**Every finding from rev b through rev e needs re-reading in this light.**
"100% of heat sits at Hera" was never evidence about where the fleet's
attractor is — it was evidence that the redistribution mechanism has been
inert since it was written. The `doe_log.c` extension itself is correct
and worth keeping regardless (it's what exposed this); the flat CSV data
it's been recording is the artifact to fix, not a finding to interpret.
Reported, not yet fixed — awaiting a decision on the nonzero seed value
before touching `capsule_vm_physics.c` again.
---
*Rev a posed the question and found the instrumentation gap — incorrectly,
as rev f later discovered. Rev b closed the per-VM-heat half of that gap
and got one real observation. Rev c documented the full four-phase
roadmap. Rev d implemented and verified Phase 1, getting a second reading.
Rev e implemented and verified Phase 2, getting a third. Rev f extended
the time-series export that already existed and discovered all three
readings agreed because the redistribution mechanism has never activated —
a seeding bug in this session's own `capsule_vm_physics.c`, not a finding
about attractors. Rev g fixed the seed and a second bug it exposed, then
found a third, deeper, pre-existing bug blocking both — not yet resolved.*
## Rev g: two real fixes applied, a third bug found underneath, not yet fixed
**Fix 1 (applied):** `fleet_transfer_slope_q48` reseeded from `0` to
`65536/3`, matching the word-level `decay_slope_q48`'s actual bootstrap
value (`src/vm_bootstrap.c`: `(1ULL << 16) / 3` — the `include/vm.h:504`
comment claiming "2:1 = 131072" is stale relative to the real code).
Confirmed via direct debug instrumentation that the seed takes effect
(`VM-PHYSICS-STATUS` reads `fleet_transfer_slope_q48=21845` throughout).
**On its own, insufficient** — heat still never moved.
**Fix 2 (applied):** traced the continued flatness to `vm_physics_touch`'s
three call sites (`mama_word_vm_exec`/`vm_call`/`vm_step` in
`mama_forth_words.c`) using `timer_now_ns()` — a low-level raw timer
function distinct from `vm_monotonic_ns(vm)`, the time source the rest of
the kernel's physics/heartbeat code actually uses successfully. Switched
all three to `vm_monotonic_ns(vm)`. Needed `#include "starkernel/vm/
vm_internal.h"` (qualified path required: an unrelated, unguarded-by-name
`src/vm_internal.h` for the *hosted* build exists too, and `-Isrc` searches
before `-I<KERNEL_SRC>/vm`, so a bare `#include "vm_internal.h"` silently
resolved to the wrong file and produced an implicit-declaration error with
no hint why).
**Bug 3 (found, not yet fixed — the actual remaining blocker):**
`vm_monotonic_ns(vm)` *also* returns 0, unconditionally, confirmed via
direct instrumentation showing `hz=0` on all 933,702 calls across a full
boot. Root cause is in `src/starkernel/hal/host_services.c`'s
`kernel_monotonic_ns()`: it checks `if (timer_tsc_hz() == 0) return 0`
with no further fallback. Under this QEMU/TCG environment, TSC frequency
calibration genuinely fails (confirmed in the boot log: `Timer: WARNING:
could not derive TSC frequency; PM Timer will be used for RELATIVE ns` /
`Timer: trust=1 (0=NONE,1=REL,2=ABS), TSC=0 Hz`) — and the timer subsystem
already has a correct, working fallback for exactly this case
(`timer_now_ns()`'s `calib_record.vm_mode && trust < TIMER_TRUST_ABSOLUTE`
branch, calling `timer_now_ns_vm_relative()`). `kernel_monotonic_ns()`
just never learned about it.
**This is pre-existing infrastructure, not introduced this session, and
it's bigger than VM-fleet physics.** `vm_tick_apply_background_decay(vm,
vm_monotonic_ns(vm))` — word-level heat decay, called every heartbeat
tick — goes through the identical broken function. If this environment's
TSC calibration failure is representative of real hardware this project
targets (not just this QEMU/TCG dev environment), word-level background
decay may have been silently inert under the same conditions as the
VM-fleet mechanism.
**Not yet resolved:** even `timer_now_ns()` itself returned 0 in the same
debug session, before the fix to use `vm_monotonic_ns` was applied — so
either `timer_now_ns_vm_relative()` has its own issue, or its trigger
condition wasn't being met the way expected. That second layer wasn't
traced further; scope check requested before continuing, since a
`host_services.c` fix has a much wider blast radius than the two capsule
files this whole investigation started in.
**Phase 3 stays blocked** — pending a decision on how far to take the
timer investigation, plus fixes 1-2 above still can't be verified to
actually move heat until fix 3 (or an equivalent) lands.
## Rev h: bugs 3, 4, and a fifth found underneath — all fixed, heat finally moves, three-arch verified
**Fix 3 (applied):** `kernel_monotonic_ns()` (`src/starkernel/hal/
host_services.c`) reimplemented TSC→ns conversion directly and gave up
(returned 0) whenever `timer_tsc_hz()` was 0 — which it reliably is under
QEMU/TCG. Replaced the whole body with a direct delegation to
`timer_now_ns()`, which already has the correct `TIMER_TRUST_RELATIVE`
fallback via the ACPI PM Timer (`timer_now_ns_vm_relative()`) that this
function never used.
**Bug 4 (found and fixed):** `vm_tick_inference_engine()` (word-level) and
`vm_physics_tick()` (VM-fleet) shared the same `last_inference_tick`
counter in `HeartbeatState`. The word-level call runs first in `vm_tick()`
and resets that counter every `HEARTBEAT_INFERENCE_FREQUENCY` ticks,
so the VM-fleet gate immediately after it never saw its own threshold
satisfied — starved permanently. Added a dedicated
`last_fleet_inference_tick` field to `HeartbeatState` (`include/vm.h`),
initialized in both `src/starkernel/vm/vm_bootstrap.c` and the hosted
`src/vm_bootstrap.c` (shared header). Checked `test_contracts.c` first —
its axiom snapshot only captures `tick_target_ns`, not the whole struct,
so the new field carries no contract risk.
**Bug 5 (found and fixed — the actual remaining blocker):** even with
fixes 3 and 4 in place, heat *still* never moved. Traced end to end with
temporary instrumentation (added, used, fully removed — confirmed via
`git diff --stat` showing zero net change beyond the three real fixes):
1. `vm_physics_touch()` received `now_ns=0` on every call.
2. `kernel_monotonic_ns()``timer_now_ns()` returned `0` on all ~930,000
calls across a full boot, not just early ones.
3. Inside `timer_now_ns()`: `trust=1` (RELATIVE) and `vm_mode=1` were
correctly set the whole time, so it *did* take the
`timer_now_ns_vm_relative()` branch.
4. Inside `timer_now_ns_vm_relative()`: `pmtimer_read()` returned
`0xFFFFFF` (masked all-ones) on *every* call, including the very first
one used to set `vm_pm_start` — so the tick delta was permanently 0.
An all-ones port read is the signature of an unmapped I/O port.
`PMTIMER_IO_PORT` was hardcoded to `0x408` — the legacy PIIX4/i440fx ACPI
PM Timer address — but `Makefile.starkernel` boots amd64 with `-machine
q35,accel=tcg`, and Q35's ICH9 LPC bridge puts the ACPI PM Timer at
`0x608`. The kernel had been reading a dead port the entire time; this
had nothing to do with fixes 3/4 and would have defeated any timer fix
built on top of it.
**The real fix, not a machine-specific patch:** added
`fadt_find_pm_tmr_port()` to `src/starkernel/arch/amd64/timer.c`, which
walks RSDP → XSDT → FADT ("FACP") — the same pattern already used by
`pci.c`'s MCFG lookup, independently duplicated here since the kernel has
no shared ACPI header and touching already-working PCI enumeration code
wasn't worth the risk. Prefers the ACPI 2.0+ `X_PM_TMR_BLK` Generic
Address Structure (offset 208) when it names a valid SystemIO address,
falls back to the legacy 32-bit `PM_TMR_BLK` field (offset 76), and only
falls back further to the old hardcoded `0x408` if FADT parsing fails
entirely. `pmtimer_port` is now a runtime variable, discovered once in
`timer_init()` from `boot_info->acpi_table`. This is machine-type-agnostic
by construction — it would have found `0x608` on Q35 or `0x408` on
i440fx without needing to know which one it's running on.
**A sixth bug found underneath *that*:** the FADT lookup initially failed
outright (`Timer: WARNING: FADT PM_TMR_BLK lookup failed`). Root cause:
`src/starkernel/boot/uefi_loader.c`'s ACPI table discovery loop matched
`EFI_ACPI_20_TABLE_GUID` *or* `EFI_ACPI_TABLE_GUID` and broke on
whichever came first in the firmware's configuration table — silently
handing back the legacy ACPI 1.0 RSDP (revision 0, no XSDT) even though
OVMF also publishes an ACPI 2.0 one. Any XSDT-based table lookup
(MCFG *or* FADT) fails against that pointer. Fixed with two explicit
passes: ACPI 2.0 first, ACPI 1.0 only as a fallback if 2.0 isn't found.
This is shared boot code (all three architectures), so it was subject to
full three-arch acceptance, not just amd64.
**Verification — all three architectures, clean acceptance runs:**
- amd64: `Timer: PM_TMR_BLK discovered from FADT at port 1544` (0x608,
exactly the expected Q35 address). TSC calibration also started
succeeding as a side effect (`TSC=2305485654 Hz`, previously always 0)
`calibrate_tsc_with_pmtimer()` depends on the same working PM Timer
read.
- All three (amd64/aarch64/riscv64): identical `dict_hash=
0x23c8d067ec5709b6`, `PASS: E2E msg flow`, 0 `FAIL` lines, no panics/
faults, virtio-blk/Artemis attach unaffected by the `uefi_loader.c`
change.
- **Heat now actually moves.** DOE CSV `(hera_heat_q48, hermes_heat_q48,
artemis_heat_q48)` triples show three distinct states across all three
architectures — `(65536,0,0)`, `(0,65536,0)`, `(0,0,65536)` — tracking
which VM was touched most recently, with the conservation invariant
intact (sum always exactly `65536` = `Q48_ONE`). This was a flat,
unmoving trace on every single run before this fix chain.
**Phase 3 is now unblocked.** The observation mechanism this whole
investigation was blocked on — being able to see real heat trajectories
move between VMs — is confirmed working on all three architectures.
## Rev i: Phase 3 implemented and run — three-arch, 180 runs each
**Design, as implemented:** `EXEC-FLEET-DOE ( seed n-runs -- )`, added to
Hera's `init.4th`. Each run draws `(hera_cfg, hermes_cfg, artemis_cfg)`
independently and uniformly from 015 (with replacement — consistent
with Phase 4's full 16³ model, not a permutation of distinct values),
applies each, drives one cheap touch per remote VM to exercise
`vm_physics_touch`, and emits a `FLEETDOE,run_id,hera_cfg,hermes_cfg,
artemis_cfg` marker row. The already-working `doe_log.c` per-tick heat
CSV is the response variable; no new instrumentation was needed for
that half. Not run at boot — invoked post-boot via the same serial-socat
injection pattern the existing `DOE_INJECT`/`EXEC-DOE` mechanism uses,
with `12345 180 EXEC-FLEET-DOE` as the command.
**Three more capsule/runtime bugs found and worked around, not fixed at
the root (each would require touching C subsystems well beyond this
investigation's scope, per the standing rule to report and confirm
before that kind of change):**
1. **Cross-capsule `EXEC`/`USE` doesn't reliably resolve words when
triggered by a remote `VM-EXEC` call after boot.** Confirmed
pre-existing, not introduced here: `common:msg.4th`'s `USE` already
silently fails inside Hermes's own `CD-INIT` in every prior accepted
run (`[CAPSULE][DEFER] ... UNKNOWN WORD: 'USE'`) — it just never
mattered because nothing calls `HERMES-ACK`/`HERMES-NACK`. Worked
around: Hermes and Artemis's config application bypasses `APPLY-CFG`
entirely. Hera computes the four Q48.16 factor values locally and
sends them as literals straight to `L8-UPDATE`/`L8-APPLY` (both C
primitives, present in every VM) via a `FDOE-CFG-CMD-LO`/`-HI`
lookup-table dispatcher (split across two words/blocks to fit
`mkcapsule`'s 16-content-line-per-block limit at this string length).
2. **`VARIABLE`/`CREATE` storage allocated via nested `EXEC` of a
separately-named capsule fails the VM's own `vm_addr_ok()` bounds
check on later `!`/`@`.** Found by elimination while debugging bug 1's
workaround: a shared `common:doe-cfg.4th` capsule (mirroring
`common:msg.4th`'s pattern) was tried for Hera's own `APPLY-CFG`.
`0 APPLY-CFG` typed directly at Hera's live REPL failed with a bare
`ERROR` (no message — `memory_word_store`/`memory_word_fetch` set
`vm->error=1` silently on a failed bounds check). Isolated via direct
REPL testing: `0 CFG-ENT .` and `0 0 0 0 L8-UPDATE L8-APPLY` both
worked fine standalone; `VARIABLE FOO 5 FOO ! FOO @ .` (a brand-new
variable declared *interactively*, no capsule involved) also worked
fine. Only `CURR-CFG`, defined via the nested `EXEC`/`USE` of
`common:doe-cfg.4th`, failed on store/fetch — a distinct bug from #1,
specific to `VARIABLE`/`CREATE` addressing through that load path.
Fixed by deleting the shared capsule and inlining `APPLY-CFG` (plus
its `CURR-CFG` variable) directly into Hera's own `init.4th`, where
the identical pattern works.
3. **`ART-TICK` scans all 22,998 Artemis data blocks per call** — far
too expensive as a per-run "touch" once multiplied by 180 runs under
QEMU/TCG with full execution logging; the first real attempt stalled
rather than erroring. Replaced with a new one-line no-op `ART-PING`
word added directly to Artemis's own capsule.
**A fourth finding, not a capsule bug:** the kernel's `--log-level` boot
flag is parsed (`cmdline.c`) but never wired to `log_set_level()` —
`sk_vm_bootstrap.c` hardcodes `log_set_level(LOG_TEST)` unconditionally
for Hera, so the flag has no effect on the kernel target (it does work
on the hosted build). This made the ECW/INFO per-word execution trace —
appropriate for boot-time POST visibility — impossible to suppress for
a long post-boot campaign, and a first 20-run smoke test only completed
6 runs in 600 seconds as a result. Rather than touch `sk_vm_bootstrap.c`,
`EXEC-FLEET-DOE` now brackets its run loop with the already-existing
`LOG-LEVEL!` FORTH word (`1 LOG-LEVEL!` / `3 LOG-LEVEL!`, i.e. WARN
during the loop, back to TEST after) — a capsule-only fix with zero data
impact, since `FLEETDOE` rows and the `doe_log.c` CSV are raw
`console_puts` output, not gated by log level. Cut the same 20-run smoke
test to well under two minutes.
**Verification — three-arch, 180 runs each, single boot per
architecture, seed 12345:**
| Arch | FDOE errors | `doe_log.c` rows | FLEETDOE rows extracted |
|------|-------------|------------------|--------------------------|
| amd64 | 0 | 5781 | 124 / 180 |
| aarch64 | 0 | 5781 | 124 / 180 |
| riscv64 | 0 | 5781 | 124 / 180 |
Identical row counts across all three architectures, consistent with
deterministic same-seed behavior. Config draws land roughly uniformly
across 015 (512 draws each per factor, out of 180 — expected mean
11.25, normal sampling variance).
**The missing 56/180 FLEETDOE rows per run, every time:** the
interrupt-driven 100Hz heartbeat tick's `doe_log.c` CSV emission can
preempt Hera's foreground `FLEET-DOE-ROW` print mid-string on the shared
serial line (no locking between the two console-write paths), splicing
an embedded `[HADES][DOE ] ...` fragment into the run marker. The
extraction script (`extract_fleetdoe.py`) detects any `FLEETDOE,` line
that doesn't split into exactly 5 clean fields and drops it rather than
guessing — a known, logged, non-silent loss, not a fix to the underlying
console race (which would mean adding locking to an interrupt-context
console path, out of scope here).
**Correction to an earlier claim in this rev:** this section originally
reported "5767 of 5781 distinct heat-triple states" as evidence of rich,
continuous trajectory data. That number was wrong — a measurement
error, not a finding. Rechecked directly against
`experiments/bare_metal/runs/doe-amd64-20260705-223649.csv`: there are
exactly **5** distinct `(hera_heat_q48, hermes_heat_q48,
artemis_heat_q48)` states in the entire 5781-row trace — `(65536,0,0)`,
`(0,65536,0)`, `(0,0,65536)`, and two single-row transient states
(`(59317,6219,0)` and `(63042,2494,0)`) caught mid-transfer. The
mechanism is overwhelmingly winner-take-all, exactly as the transfer
formula (`amount = elapsed_ns * slope >> 16`, clamped to
`others_total`) predicts for any touch separated by a realistic time
gap. The likely source of the earlier wrong number: counting distinct
full CSV rows (near-unique by construction, since `tick_number`/
`elapsed_ns` differ every row) rather than distinct heat-triples
specifically — a one-off shell-pipeline mistake, not a code or data
bug. Recorded here rather than silently fixed, per this document's own
transparency-first design goal.
## Rev j: statistical analysis of the 180-run campaign — a clean null result
**Response-variable reconstruction.** `FLEET-DOE-ROW` and
`doe_log.c`'s per-tick heartbeat CSV emission share one
un-synchronized serial line — the interrupt-driven 100Hz tick can
preempt Hera's foreground print mid-string, which is also what corrupts
56/180 `FLEETDOE` rows per run (Rev i). Since no explicit join key was
recorded at collection time, a join was reconstructed after the fact by
walking each redacted log in original line order and tagging every
per-tick heat sample with whichever `FLEETDOE` config triple was most
recently seen before it. This is an approximation (a sample shortly
after a config change may still reflect the previous config), not a
precise per-run isolation, and it degrades further whenever several
consecutive `FLEETDOE` markers are corrupted (the samples in between
get attributed to the last *good* marker, silently spanning more than
one real run) — good enough for point-in-time transition analysis,
not for clean per-run window statistics. Both scripts
(`join_fleetdoe.py`, `analyze_transitions.py`) live in session
scratch, not committed to the repo; numbers below are reproducible from
the committed logs.
**Cross-architecture determinism: confirmed, cleanly.** Comparing the
124 runs with usable config markers across all three architectures:
**0 config-triple mismatches.** amd64/aarch64/riscv64 drew the identical
sequence of `(hera_cfg, hermes_cfg, artemis_cfg)` triples from the same
seed. This is the one unambiguous, load-bearing result of the campaign.
**Heat dynamics during the actual 180-run campaign: sparse, winner-take-
-all, no fractional states.** Counting state changes directly from the
log (not through the run-window join, so unaffected by marker
corruption): **318 total transitions** across the whole campaign, landing
on hera/hermes/artemis 140/98/80 times respectively — a real, uneven
distribution, but see below on causal weight.
**Does any config bit predict which VM a transition lands on?** For each
of the 12 binary factors (entropy/cv/temporal/stability × hera/hermes/
artemis), a permutation test (10,000 shuffles) compared the rate at
which transitions landed on VM *X* when *X*'s own factor bit was high
vs. low, using the 306/318 transitions with a known config at that log
position. **No factor reached significance** (all p > 0.3; most
differences were 0.000.06 in either direction, well inside permutation
noise for n≈150 per group). This is a clean null, not an ambiguous one.
**Why a null result is the architecturally expected outcome, not a
surprise:** tracing `APPLY-CFG`/`L8-UPDATE` → `ssm_l8_update()` →
`ssm_apply_mode()` shows the DoE config only ever writes into
`vm->ssm_config` (word-level physics tuning — window width, decay-slope
inference). `fleet_transfer_slope_q48`
(`src/starkernel/capsule/capsule_vm_physics.c:94`) — the only knob that
controls how much heat moves per touch — is updated exclusively by
`vm_physics_tick()`'s own touch-history regression
(`capsule_vm_physics.c:347`), with no code path connecting it to L8
state at all. **There is no causal channel, direct or indirect through
the shared-slope mechanism, from the DoE's manipulated factor to the
fleet-heat transfer rate.** Any effect the campaign could possibly have
detected would have to be mediated through incidental timing side
effects (does a different L8 window-width config change how long
`HERMES-TICK`/`ART-PING` takes, shifting `elapsed_ns` between touches
enough to matter) — a confound, not the intended manipulation, and one
weak enough that 318 transitions found no trace of it.
**What this means for the escalation criterion.** The design doc's
Phase 3 → Phase 4 escalation criterion was "inconclusive → escalate to
the full 16³ campaign." This result is not inconclusive — it's a clean
null with an identified structural cause. Running Phase 4 (≈123,000
runs) against the same architecture would almost certainly reproduce
the same null at a hundredfold cost, because the missing causal channel
doesn't appear at a larger sample size. **Escalating to Phase 4 is not
recommended without first giving L8 config an actual causal pathway
into `fleet_transfer_slope_q48`** (e.g., letting `ssm_apply_mode()` or
an equivalent feed a per-VM or fleet-wide slope adjustment) — a real
design decision, not a bug fix, and squarely a call for Bob rather than
something to implement speculatively here.
## Rev k: the real defect — the fleet's self-tuning loop is starved, not just disconnected from config
Bob's instinct after seeing the rev j numbers ("looks like a coin toss
or not enough data... probably a miswired feedback loop") was right,
and more precise than rev j's framing. Rev j established that config
has no path *into* the fleet mechanism. This rev establishes something
underneath that: **the fleet mechanism's own self-tuning loop —
`vm_physics_tick()`'s regression that's supposed to periodically re-fit
`fleet_transfer_slope_q48` from the fleet's own touch history — never
successfully runs, for any campaign of realistic size, regardless of
config.** Confirmed by direct instrumentation (added, verified, fully
removed — `git diff --stat` on both touched files showed zero residual
change before the finding was committed):
**Instrumenting `vm_physics_tick()` directly: zero calls, ever.**
Across an entire 20-run smoke test — full boot, POST, readiness
handshake, broadcast test, `TRIPOD-TEST`, and the campaign itself — a
call counter placed at the top of `vm_physics_tick()` never
incremented once. Not "called but bailing on `traj_len < 2`" — never
entered at all.
**Tracing the call site (`vm_tick()`, `vm_runtime.c:137`) to why:** the
fleet-inference gate is `(VM*)sk_get_mama_vm() == vm && (tick_count -
last_fleet_inference_tick) >= HEARTBEAT_INFERENCE_FREQUENCY`.
`HEARTBEAT_INFERENCE_FREQUENCY` is `1000` (`starforth_config.h:72`) —
1000 invocations of `vm_tick()`, which itself only runs once every
`HEARTBEAT_CHECK_FREQUENCY` (`256`) word executions **on that specific
VM instance**. Instrumenting `vm_tick()` itself, filtered to calls
where `vm` is Hera: across the full 20-run campaign, **Hera's own
`vm_tick()` fired a total of 61 times** (`tick_count` and the
gate's `diff` both at 61 by the end, `last_fleet_inference_tick` still
0 throughout — never once satisfied). A parallel unfiltered sample
showed Hermes and Artemis's own tick counts climbing far faster (233,
863, 1363 within the same window) — Hera's *own* word-execution rate is
the bottleneck, not the fleet's overall activity level.
**Why Hera's own rate is so low:** the gate was inherited unchanged
from the word-level inference engine, where it's well-calibrated —
Hera's dictionary handles millions of word executions during a real
boot's POST phase, so 256,000 words between inference attempts is
nothing. But Hera, in her role as fleet orchestrator, is mostly
*dispatching* — `VM-EXEC`/`VM-CALL` calls that block while the
*target* VM (Hermes/Artemis) does the actual work in *its own*
`heartbeat.check_counter`, which never counts toward Hera's. A DoE run
loop built from `RANDOM`, stack shuffling, a `CASE` dispatch, and a
handful of `VM-EXEC` calls executes on the order of a few dozen of
Hera's *own* words per run — nowhere near the volume the constant
assumes. At the observed rate (~3 of Hera's own ticks per run), the
180-run campaign accumulates roughly 540 — confirmed directly:
`VM-PHYSICS-STATUS` read before and after the full 180-run campaign (a
separate diagnostic run, not the debug-instrumented one) showed
`fleet_transfer_slope_q48=21845` and `slope_fit_quality_q48=0`
**identically, before and after** — the seed value, untouched, despite
`fleet_window_warm=YES` (satisfied almost immediately from a handful of
early boot-time touches) the entire time. Both gates exist; only one of
them is reachable at DoE-campaign scale.
**This means Phase 3's premise was compromised in a way rev j didn't
yet capture:** it's not just that config can't reach the fleet
mechanism — the fleet mechanism's own adaptive half never engages at
all during a targeted campaign like this one. `fleet_transfer_slope_q48`
behaved as a hardcoded constant (`65536/3`, the bootstrap seed) for the
entire 180-run run, on all three architectures. Every transition
observed in Phase 3's data was governed by that one frozen number, not
by anything resembling the "regression-inferred, self-tuning" behavior
the design doc specifies. This is a stronger explanation for the null
result than "no causal channel from config" alone — even a campaign
that *did* somehow influence timing enough to matter would be pushing
against a slope that never moves.
**Not fixed here.** This is a real defect in shared kernel code
(`vm_runtime.c`'s heartbeat gate), not a capsule-side workaround
target, and warrants its own decision on the right fix — a fleet-scoped
inference frequency independent of `HEARTBEAT_INFERENCE_FREQUENCY`
(most direct), gating on fleet touch count rather than Hera's own tick
count (matches the window's own warm-up signal, which already fires
promptly), or something else — flagged for Bob rather than implemented
speculatively.
## Rev l: rev k's gate fixed with a shared counter — and a second defect found underneath
Bob's diagnosis of rev k's null result ("miswired feedback loop") led
straight to the fix, and to a spec passage that had already ruled out
the broken version: `VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md`
(line 82) specifies that `fleet_transfer_slope_q48` must be inferred
"from aggregate fleet statistics... not one per VM" — the same
principle applies to the *readiness signal* that gates computing it,
not just the regression's own input trajectory. Gating on Hera's own
`heartbeat.tick_count` was exactly the "one per VM" framing the spec
already rules out for this parameter — not an missing case, a
mis-implementation of a spec that said the opposite.
**Fix, implemented and committed:** `vm_physics_heartbeat_tick()`
(`capsule_vm_physics.c`), a new fleet-wide tick counter distinct from
any VM's own `HeartbeatState`, called from every VM's own `vm_tick()`
(not gated to Hera specifically). Removes the now-dead
`last_fleet_inference_tick` per-VM field this replaces. Builds clean
(zero warnings on touched files) on the hosted build and all three
kernel architectures.
**Confirmed working, for what it fixes:** `vm_physics_tick()` now
fires and completes — `slope_fit_quality_q48` moved from `0` to the
fit-succeeded placeholder `52429`, confirmed via `VM-PHYSICS-STATUS`
before/after a 180-run campaign. The starvation defect is real and
gone.
**A second, distinct defect found immediately while verifying the
first:** the regression computed `fleet_transfer_slope_q48 = 0`.
`vm_physics_touch()` requires `slope > 0` to transfer any heat at all
— so the moment the now-functional loop fires and lands on exactly
zero, it freezes heat movement permanently, reproducing Bug 1's
original zero-slope deadlock (`rev g`) through a different door.
Confirmed directly: **zero heat-state transitions across an entire
180-run verification campaign** — one state, `(0,0,65536)`, from
injection to completion, with the slope already at 0 before the
campaign even started (from earlier boot-time touches).
Likely cause, not yet distinguished: the regression assumes smooth
exponential decay (`heat(t) = h0·e^(-slope·t)`), but the fleet's actual
dynamics — confirmed back in the correction to rev i — are
winner-take-all step functions: one VM at full heat, the others at
zero, snapping abruptly rather than decaying smoothly. Fitting a
smooth-decay model to a step-function trajectory could plausibly
produce an exact-zero fitted slope as a genuine (if unhelpful) result,
not a bug in the fitting code. Q48.16 integer-division truncation of a
small nonzero result to 0 hasn't been ruled out either — the two would
need different fixes (a model that matches the actual dynamics, vs. a
precision fix in the existing division). **Not chased further here** —
flagged as the next open item, same standing as rev k's original
finding: a real question for Bob, not something to guess at by
implementing a fix for either hypothesis speculatively.
**Net effect on Phase 3's data:** unchanged from rev k's conclusion.
The 180-run campaign already committed ran entirely under the frozen
seed value (`65536/3`), not this newly-discovered zero — the shared-
counter fix and this second defect were both found *after* that
campaign, during verification of the fix. Re-running Phase 3 now would
just freeze at a different, worse constant (`0` instead of `65536/3`)
until the model/precision question above is resolved.
## Rev m: rev l's zero-slope defect resolved — model mismatch, confirmed empirically, not truncation
Bob's question: "let's figure out which it is, a model mismatch or a
Q48.16 truncation. I'm thinking the former." Confirmed empirically —
it's model mismatch, and not a close call.
**Method:** temporary instrumentation (`TOUCHDBG` in
`vm_physics_touch()`, `SLOPEDBG` in `vm_physics_tick()`, printing
`traj_len`, distinct-value count, `sum_log_heat`, `sum_t_log_heat`,
the pre-`abs()` signed numerator, and `denominator`), a 20-run
amd64/QEMU verification boot, then full removal — confirmed via
`git diff --stat` showing zero residual change before this write-up.
**What the boot-time touches showed:** the seeded slope (`65536/3`)
does move heat during ordinary boot activity — `TOUCHDBG` recorded a
real, nonzero transfer (`elapsed_ns=3310198 amount=1103382`) — so the
transfer arithmetic itself is not the problem, and Q48.16 precision is
in no way starving the transfer step. But the fleet's own dynamics
(rev i's winner-take-all finding) drove that transfer to *completion*
before Phase 3 even started: by the first post-boot
`VM-PHYSICS-STATUS`, one VM already held the entire `65536` and the
others held `0`.
**What `SLOPEDBG` showed at the moment the regression fired:**
```
SLOPEDBG traj_len=64 distinct=2 n=64 sum_t=2016
sum_log_heat=0 sum_t_log_heat=0
numerator_signed=0 (pos-or-zero) denominator=1397760 slope=0
```
`sum_log_heat` and `sum_t_log_heat` are exact zero, not small values
that rounded to zero — and the reason is exact, not approximate:
`vm_physics_tick()` looks up each history sample's *current*
`execution_heat_q48`, and skips zero-heat samples
(`if (trajectory[i] == 0) continue`) before ever computing a log. With
heat fully concentrated in one VM, every *contributing* sample has the
identical value `Q48_ONE` (`65536`) — which is Q48.16 for `1.0`, and
`ln(1.0) = 0` exactly. So `sum_log_heat = 0` and `sum_t_log_heat = 0`
are not underflowed approximations of something small; they are the
sum of exact zeros. `numerator_signed` is then `0 - 0 = 0`, again
exactly, before any division ever happens. There is no division step
where a small nonzero value could have been truncated away — the
truncation hypothesis doesn't have anywhere left to hide once the
inputs to the division are already exactly zero.
This holds for **any** distribution the fleet ever reaches under the
current winner-take-all dynamics, not just the one observed: whichever
single VM holds all the heat holds exactly `Q48_ONE`, whose log is
always exactly zero, and every other live VM's zero-heat samples are
filtered out before contributing to the sums at all. The regression's
input is structurally incapable of being anything but a constant series
of exact zeros once the fleet has fully concentrated — which, per rev
i, it does quickly and by design (winner-take-all is the intended
touch semantics, not a bug).
**Conclusion: this is model mismatch, confirmed, not Q48.16
truncation.** The regression assumes smooth log-linear decay
(`heat(t) = h0·e^(-slope·t)`); the fleet's actual dynamics are a step
function between exactly `0` and exactly `Q48_ONE`. Fit that model to
that data and an exact algebraic zero is the *correct* output of the
math — infinite-precision floating point would produce the identical
zero, for the identical reason (`ln(1) = 0`, and OLS on a constant
series has zero slope by construction). No fix has been applied. Per
this session's practice, this is a finding to report and confirm, not
implement speculatively — the next question, not yet asked, is what
model *should* replace log-linear decay for a fleet whose real dynamics
are step functions, and that's Bob's call.
## Rev n: two real fixes applied ("continue the best path according to the architecture") — a third, distinct root cause found and left for Bob
Bob approved rev m's diagnosis as "the best resolution for the desired
architecture" and asked to continue. Two fixes landed; a third question
is still open.
**Fix 1 (implemented, committed `9806263c`):** `VMFleetWindow.touch_history`
renamed to `heat_history`, now storing `target->physics.execution_heat_q48`
at the moment of each touch (after any transfer has settled) instead of a
`vm_id` to re-resolve at fit time. This is the rev m fix itself — it makes
the trajectory an actual historical time series rather than N lookups of
"now." Correct and necessary, confirmed by rebuild across the touched file.
**Verification after fix 1 found a second, independent defect:** slope
still landed on exactly `0`. Root cause, confirmed by instrumentation
(added, checked, removed): `vm_physics_touch()` applied
`fleet_transfer_slope_q48` against raw `elapsed_ns`, but that slope's seed
value (`65536/3`) was explicitly chosen to mirror the word-level engine's
own bootstrap constant — and the word-level engine applies its slope
against elapsed **microseconds** (`physics_metadata.c:337`,
`elapsed_us = elapsed_ns / 1000`, commented there as a deliberate
overflow-avoidance choice). Using nanoseconds directly was a straight
1000x scale error, not a design choice.
**Fix 2 (implemented, committed `5784b5ac`):** convert to `elapsed_us`
before applying the slope, exactly mirroring `physics_metadata_apply_linear_decay`'s
own convention.
**Verification after fix 2 found a third, distinct defect — not yet
fixed:** `TOUCHDBG2` instrumentation (added, checked, removed) on a fresh
boot showed exactly why. The first two real touches after boot:
```
TOUCHDBG2 vm_id=2 elapsed_us=207021 amount=69005
TOUCHDBG2 vm_id=1 elapsed_us=361602 amount=120532
```
Both `amount` values already exceed the entire conserved heat pool
(`Q48_ONE = 65536`) — these are one-time boot-sequence pauses (capsule
loading / disk I/O), 207ms and 362ms respectively, and at the corrected
per-microsecond rate the full-drain threshold is only `65536*65536/21845 ≈
196,608 us ≈ 197ms`. So the very first real touch, before the fleet window
has a chance to accumulate any pre-transition samples, still fully drains
the pool in one step. Contrast with steady-state touches captured moments
later in the same boot, once the interpreter is in its normal dispatch
rhythm:
```
TOUCHDBG2 vm_id=0 elapsed_us=9367 amount=3122
TOUCHDBG2 vm_id=0 elapsed_us=9639 amount=3212
TOUCHDBG2 vm_id=0 elapsed_us=9194 amount=3064
```
These are genuinely gradual — roughly 5% of the pool per touch, exactly
the smooth-pull behavior the model was meant to produce. The unit fix is
real and working; it's just arriving too late relative to two anomalously
long one-time pauses baked into this specific boot sequence.
**Assessment, not yet acted on:** this reconfirms rev m's core finding
from a cleaner angle. The fleet's dynamics, once correctly scaled, *are*
capable of gradual, regressable transfer — but only during steady
interpreter execution. Boot-sequence pauses (disk I/O, capsule loading)
are structurally one-time and long relative to the corrected time
constant, so they will keep pre-empting the gradual regime before the
window fills, for any fleet whose boot includes comparable pauses. Fixing
this further means either (a) a maximum per-touch transfer cap (bound
`amount` to some fraction of the pool regardless of elapsed time), (b)
excluding known one-time boot pauses from the elapsed-time calculation
(treating first-touch-after-birth specially), or (c) accepting that
winner-take-all-after-a-long-gap is correct fleet semantics and it's the
regression model, not the transfer mechanics, that still needs to change
(rev m's original conclusion, now with the transfer math confirmed
correct on its own terms). Not chased further here — three fixes deep in
one sitting is enough to stop and let Bob pick the direction rather than
keep guessing at numeric knobs.
## Rev o: option (c) chosen — log-linear regression replaced with direct rate recovery, three-arch verified
Bob's call on rev n's three-way fork: option (c) — winner-take-all after
a real dormancy gap is correct fleet semantics; a cooperative
single-threaded fleet where exactly one VM executes at any instant
naturally produces discrete handoffs, not continuous diffusion, so the
regression model was fighting the actual physics rather than modeling
it. Options (a) and (b) would have been numeric knobs tuned to force
step dynamics to look smooth to a model that was never going to fit
them — the same speculative-tuning trap rev l/n already fell into twice.
**New model:** the transfer law (`amount = elapsed_us * slope >> 16`,
clamped to available heat) is already known exactly, so there's no need
to curve-fit a reconstructed trajectory against a synthetic time axis at
all. Each touch now records its own transfer-law inputs —
`{elapsed_us, amount, clamped}` — instead of a derived heat value.
`vm_physics_tick()` inverts the law directly for every **unclamped**
sample (`rate = amount / elapsed_us`, exact, no log transform) and takes
the **median** across the window. Clamped samples (where the actual
transfer was capped by what the rest of the fleet held, including the
`others_total == 0` fully-concentrated case) carry no rate information
and are excluded outright, not zero-filled. Fewer than 8 informative
samples in the 64-deep window: skip the update, keep the current slope
— the same "skip, don't substitute a degenerate value" philosophy as
the existing `is_warm` gate, now also closing the zero-freeze trap at
its source instead of hoping the input data avoids it.
`slope_fit_quality_q48` is now `informative_count / window_depth` — a
real number, replacing the fixed `0.8` placeholder that had been in
place since the original design.
**`VMFleetWindow.heat_history` renamed to `touch_samples`**, now an
array of `VMFleetTouchSample` instead of `uint64_t`. `q48_log_approx` /
`q48_mul` / `q48_add` / `q48_from_u64` are no longer used in this file
(still used elsewhere); `Q48_ONE` remains needed and the include stays.
**Verified, amd64, 60-run DOE campaign** (instrumentation added, checked
removed before commit, same methodology as every fix this session):
```
seed: fleet_transfer_slope_q48=21845 quality=0 (cold)
after warm-up: fleet_transfer_slope_q48=21845 quality=0 (< 8 informative, held)
mid-campaign: fleet_transfer_slope_q48=21840 quality=25600 (~39% informative)
post-campaign (60): fleet_transfer_slope_q48=21829 quality=14336 (~22% informative)
```
No freeze, no runaway, no exact zero. The small drift (21845→21840→21829)
is expected and benign: since `amount` for any given touch was itself
computed from whatever slope was in effect at that moment, an unclamped
touch's recovered rate is close to self-consistent by construction —
this estimator's job is to stay stable and quality-aware under real
fleet conditions, not to discover some independent ground truth the
transfer law doesn't already encode. That's a feature, not a
limitation: the failure modes being fixed here were freezing and
fighting the physics, not "insufficiently novel inference."
**Three-arch acceptance, full protocol** (`make -f Makefile.starkernel
ARCH={amd64,aarch64,riscv64} clean qemu`, one at a time, foreground):
all three booted to `ok>` clean, **976 PASS / 0 FAIL** identically on
every architecture, and identical `dict_hash`
(`0xd5c86b7db55efa58`) across all three — zero-deviation determinism
holds. Logs: `logs/20260706-061201/amd64/`,
`logs/20260706-061320/aarch64/`, `logs/20260706-061455/riscv64/`.
This closes the zero-slope investigation that ran rev k through rev o.
The fleet's transfer mechanics, the sampling that feeds its
self-tuning loop, and the estimator itself are now all consistent with
each other and with the fleet's actual winner-take-all dynamics.
## Rev p: Phase 3 re-run under the fixed model — richer dynamics, same null result
Identical protocol to rev i/j — seed `12345`, `180` runs, one boot per
architecture, `EXEC-FLEET-DOE` — re-run under rev o's fixed estimator to
see whether the earlier null result was an artifact of the frozen slope
or a real property of the fleet.
**Config-triple determinism: still holds, unchanged.** All three
architectures' `FLEETDOE` marker files are byte-identical after sorting
(`diff` on amd64 vs. aarch64 and amd64 vs. riscv64: 0 lines). Expected —
config draws are seeded PRNG arithmetic, untouched by anything in rev
n/o.
**Heat dynamics are now genuinely rich — and now architecture-
dependent, which they weren't before.** Distinct `(hera, hermes,
artemis)` heat-triple states across the 5781-row trace:
| Arch | Rev i (frozen model) | Rev p (fixed model) |
|------|----------------------|----------------------|
| amd64 | 5 | 612 |
| aarch64 | 5 | 2276 |
| riscv64 | 5 | 2273 |
The jump from 5 to hundreds/thousands of distinct states directly
confirms rev o's fix: the fleet is now spending real, observable time in
fractional-transfer states between winner-take-all snaps, not just
teleporting between three fixed corners. The new *cross-architecture
divergence* (612 vs. ~2275) is an expected, previously-invisible
consequence of the same fix: the old step-function model crossed its
full-drain threshold on almost any realistic gap, so the exact
elapsed-time value never mattered and all three architectures produced
byte-identical winner-take-all trajectories under TCG. The new model's
fractional transfers are directly proportional to real elapsed
microseconds between touches, so different emulation speeds now
produce genuinely different heat trajectories — config-sequence
determinism (what gets drawn) is preserved; heat-trajectory numeric
determinism (what happens with it) is not, and arguably should not be,
now that real timing is load-bearing. This doesn't threaten `BIRTH`/
parity determinism: the three-arch acceptance run in rev o already
confirmed identical `dict_hash` across all three architectures — that
guarantee lives in word-level dictionary state, which this doesn't
touch.
**Transition counts, per architecture** (destination distribution):
| Arch | Total transitions | hera | hermes | artemis |
|------|--------------------|------|--------|---------|
| amd64 | 301 | 134 | 91 | 76 |
| aarch64 | 403 | 167 | 162 | 74 |
| riscv64 | 432 | 179 | 173 | 80 |
| *rev j (frozen, single log)* | *318* | *140* | *98* | *80* |
Higher absolute counts than rev j across the board (expected — the
fixed model produces far more state changes, fractional and full,
rather than a handful of clean snaps), but the **relative ordering
(hera > hermes > artemis) holds in all three new logs**, matching rev
j's ordering exactly. A structural property of the fleet (Hera as root
absorbs on retire, per `vm_physics_retire`) surviving a complete
replacement of the transfer/estimator machinery underneath it.
**Permutation tests: same null result, now on a fully-corrected model.**
Re-ran rev j's exact test — 12 factor/VM combinations (`ent`, `cv`,
`tmp`, `stb` × hera/hermes/artemis), Wilson 95% CIs, 50,000-shuffle
permutation test — independently on all three new logs (36 tests
total). **No factor reached significance on any architecture**: p-values
ranged from 0.14 to 1.00, lowest at `artemis`/`ent` (amd64 0.49,
aarch64 0.14, riscv64 0.20) and `artemis`/`cv` (amd64 0.22, aarch64
0.42, riscv64 0.52). This is the same clean null rev j found, but it
now carries more weight: rev j's null could have been explained by the
frozen slope suppressing any real signal (the mechanism literally
couldn't respond to anything). That explanation is no longer available
— the mechanism now visibly responds to real touch timing (612-2276
distinct states), and the null persists anyway. **Confirms rev j's
architectural diagnosis rather than superseding it**: `ssm_apply_mode()`
still only ever writes `vm->ssm_config`; nothing in rev n or rev o gave
L8 config a causal path into `fleet_transfer_slope_q48` or
`vm_physics_touch()`, and this re-run is direct evidence that no
incidental side-channel (e.g. config-dependent timing shifts) creates
one either.
**One directionally-consistent, non-significant pattern worth
flagging:** `artemis`'s destination rate trends *lower* under high
`ent` and `cv` config bits in **all three** independent architecture
replicates (`ent`: amd64 0.042, aarch64 0.058, riscv64 0.053; `cv`:
amd64 0.070, aarch64 0.036, riscv64 0.029). No individual test is
significant, and three consistent signs out of three is not strong
evidence on its own — but it's the kind of pattern that would be worth
a purpose-built follow-up (larger n specifically on this factor/VM
pair) if this ever becomes a priority, rather than the untargeted full
16³ Phase 4 sweep.
**Recommendation: unchanged from rev j, now on firmer ground.**
Escalating to Phase 4 (~123,000 runs) is still not recommended without
first giving L8 config an actual causal pathway into the fleet
mechanism — that was true when the null might have been a broken-model
artifact, and it's still true now that the model is fixed and the null
held anyway. The zero-slope investigation (rev ko) is fully resolved;
whether to build that causal pathway is a new, separate design
decision for Bob, not a continuation of this one.
Data: `experiments/bare_metal/runs/{doe,fleetdoe}-{amd64,aarch64,
riscv64}-20260706-*.csv`. Logs: `logs/20260706-062132/amd64/`,
`logs/20260706-062357/aarch64/`, `logs/20260706-062528/riscv64/`.
## Rev q: L8 given a real causal channel into the fleet — the pathway rev j/p found missing, now built
Bob's question after rev p closed the zero-slope investigation: "shouldn't
L8 be wired in now?" — pointing at `doe_metrics.c`'s `2^7` loop-enable
space (L1-L7) and the Jacquard selector's whole purpose being to
translate observed statistics into live physics adjustment. Confirmed:
`ssm_jacquard.c` already runs a full 128-config adaptive UCB bandit
(`SsmConfigTable`), not the 16-mode legacy selector rev i/j's campaign
actually drove — but three separate gaps kept that from mattering. All
three closed, three-arch verified at each step.
**Gap 1 — L1/L4/L7 silently discarded.** The 128-config bandit always
selected across the full 7-bit space, but `ssm_apply_mode_from_table()`
and the inline apply in `ssm_l8_trial_end()` only ever wrote 4 of those
7 bits (L2/L3/L5/L6) into `ssm_config_t` — L1 (hotwords cache), L4
(pipelining), L7 (adaptive heartrate) were computed by the bandit and
then thrown away before reaching any real gate, governed instead by
whichever compile-time macro the binary was built with. Fixed:
`ssm_config_t` extended to all 7 fields, both write-sites updated, real
runtime checks wired at the two live call sites reachable from the hot
path (L4 in `physics_execution_hooks.c`'s pipelining block; L1 by
syncing the existing `hotwords_cache_set_enabled()` toggle once per L8
tick). L7 left deliberately propagate-only: `HEARTBEAT_THREAD_ENABLED`
governs background-pthread-vs-synchronous-inline dispatch, a
concurrency-model choice the tick machinery runs under either way, not
an on/off feature — live-toggling it would mean spawning/joining a real
OS thread from inside a per-tick physics callback, a different order of
risk than L1/L4's "skip a block of computation," and L7's own DoE
characterization ("ALWAYS ON, beneficial in 71% of top configs") gives
no indication toggling it off would ever be the right move. Commit
`61ac7f5e`.
**Gap 2 — the Fleet DOE drove the wrong mechanism.** `L8-UPDATE`/
`L8-APPLY` (what rev i/j's campaign actually called) drives the legacy
16-mode path — but the VM's own heartbeat always uses the 128-config
adaptive table instead (unconditionally allocated at boot), and
whichever fires more recently wins the shared `ssm_config_t`. Over a
multi-run campaign with plenty of word executions between draws, the
bandit's own periodic tick could silently overwrite a DOE-injected
config before it had any chance to matter — a second, previously-
invisible confound on top of rev j's "no wired path" finding. Fixed:
new `ssm_l8_force_config()` sets the bandit's own `current_config`
directly (an externally-forced trial looks identical to a self-selected
one, so the next trial-end scores it coherently and the bandit's normal
reward loop continues from there), exposed as `L8-TABLE-FORCE
( config_idx -- )`. `capsules/init.4th`'s whole `FDOE-CFG-CMD-LO/HI`
CASE-dispatch table (translating a 0-15 mode index into four synthetic
Q48.16 factor values) is gone — a raw 0-127 index needs no per-value
translation, just a "N L8-TABLE-FORCE" command string built via `PAD` +
`<# #S #>` + `MOVE` for remote `VM-EXEC`. `EXEC-FLEET-DOE` now draws
0-127 per VM. Net six capsule blocks removed, one added. Verified via
direct REPL injection (`VM-EXEC: '17 L8-TABLE-FORCE' -> 'Hermes'`, no
errors) and a 20-run smoke test drawing across the full range. Commit
`a6206acc`.
**Gap 3 — even correctly applied, L8 config had no path to the fleet
mechanism at all.** `fleet_transfer_slope_q48` is computed entirely by
`vm_physics_tick()`'s own rate-recovery estimator (rev o), with zero
code reading anything L8-related — closing gaps 1 and 2 makes L8 config
*apply* correctly, but still can't *reach* the fleet. Fixed: new
`l8_regime_modulation_q16()` reads Hera's own adaptive table's
`current_regime` (the same 3-bit entropy/cv/temporal classification the
bandit already stratifies its own scores by) and scales the
empirically-recovered median rate by `(popcount(regime)+1) * 0.5` —
0.5x-2.0x, richer signal environments track change more aggressively,
quieter ones damp toward the raw recovered rate. No per-VM favoritism:
this scales whatever the estimator already computed uniformly, never
decides which VM receives heat, so "physics is a passive observer,
never a driver" (`VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md`'s
Goal) holds. Floor is 0.5x, never zero, so this can't reintroduce a
freeze on its own. New coupling, named explicitly: `capsule_vm_physics.c`
previously had zero dependency on `vm.h`/`ssm_jacquard.h`; it now reads
Hera's L8 state read-only via `sk_get_mama_vm()`. Unlike
`SsmConfigTable`'s DoE-seeded priors, this modulation formula has no
empirical calibration behind it — a principled default (bounded,
symmetric, reuses an existing signal), not a proven-optimal one.
Verified via a 150-run campaign (instrumentation added, checked,
removed before commit): a window with 44/64 informative samples moved
`fleet_transfer_slope_q48` from the seed (21845) to 2705, with heat
genuinely three-way fractional (11972/31294/22270) rather than
winner-take-all. Commit `5e40808a`.
**Net effect:** L8 config now has an unbroken, verified path from a
VM's own observed workload statistics through to the fleet's transfer
rate — the exact channel rev j's null result identified as missing.
Re-running Phase 3/4 against this would, for the first time, be testing
a mechanism actually capable of showing an effect. Not done here — a
new campaign, not a continuation of this fix.
Three-arch acceptance at every step (976 PASS / 0 FAIL, identical
`dict_hash` within each step) confirms determinism holds throughout.
## Rev r/s: one clock only — collapsing the "trial" batching unit, then discovering (and fixing) why that broke determinism
Bob, after rev q: "yes. let's show it works" — a demonstration campaign
against the new L8-into-fleet wiring. Before running it, instrumentation
revealed the rev q modulation had never actually varied: `current_regime`
requires `SsmConfigTable.trial_tick_count` to reach `SSM_MIN_TRIAL_TICKS`
(2000) before a trial ends and the regime updates, and no VM in a
realistic campaign — Hera ~70 ticks, Hermes ~294, Artemis frozen at
1620 across a 40-run campaign — ever got there. The earlier "150-run
campaign confirms modulation" claim (rev q) was corrected: it confirmed
the multiplication executes without crashing, not that it responds to
anything.
Bob's diagnosis, verbatim: "`current_regime` sounds like some kind of
tight C99 coupling to a given FORTH executing experiment to me" — the
"trial" was an artificial batching unit with no principled tie to
anything, invented to smooth a bandit's reward signal but in practice
just preventing it from ever firing during a VM's real boot lifetime.
Directive, verbatim: "ONE clock only, the heartbeat tick... a tick is a
tick and a tick is our resolution." No batching, no special-casing
between VM ticks and word ticks.
**Rev r.** `ssm_l8_trial_end()` → `ssm_l8_tick_score_and_select()`,
called unconditionally every tick from `ssm_l8_update_table()` instead
of gated behind `trial_tick_count >= SSM_MIN_TRIAL_TICKS`. ANOVA
stability changed from an accumulated fraction
(`anova_exits_this_trial / inference_runs_this_trial`) to a per-tick
signal (this tick's own outcome, neutral when inference didn't run this
tick). `SsmConfigEntry.trial_count` → `tick_count`,
`SsmConfigTable.regime_trials`/`total_regime_trials` →
`regime_ticks`/`total_regime_ticks`, `SSM_MIN_TRIAL_TICKS` removed
entirely. Every "trial" name and comment scrubbed from touched code per
explicit instruction ("anything that says 'trial' must be scrubbed from
the codebase").
Bob, mid-implementation: "here's the thing I'm concerned about. You're
making assumptions based on existing DoE designs. only the DoE concept
exists, the experiment does not yet exist" — a correction against
letting the *old* Fleet DOE protocol's assumptions (a config "holding"
for a run) constrain the clock fix. Get the clock right first; design
whatever experiment fits it after, not before.
**The regression.** Three-arch acceptance under rev r showed `dict_hash`
diverging for the first time all session — amd64 `0x64bb90c39bb94b1f`
vs. aarch64 `0x41c2b512a2dc415b`. Root-caused via temporary
instrumentation (added, used, fully removed before commit) to two
wall-clock-tainted paths, both derived from `vm->decay_slope_q48`
(Loop #6's adaptively-inferred decay slope, itself indirectly
wall-clock-dependent via Loop #3's real-nanosecond heat decay): (1)
`metrics.temporal_decay`, one of L8's three regime-classification
inputs, dormant at the old once-per-2000-ticks cadence, now live and
differently influencing which physics loops are enabled per
architecture every single tick; (2) the reward's window/slope
joint-convergence signal, same root cause. Bob: "must be fixed before
moving forward." A proposed stopgap (drop the temporal dimension from
regime classification) was rejected outright: "accurate code. fix it.
never allow a workaround."
**Rev s: the generalized compudynamics module.** Mid-fix, Bob flagged
the coming shape of the problem: "sooner or later, more sooner, we will
be adding messages and blocks along with VMs and words" — any fix
needed to generalize across all four entity levels, not just patch
words. After a short language-constraint misunderstanding was cleared
up (TRIPOD.md's/ARTEMIS.md's "StarForth dialect ONLY" applies to a thin
administrative/observational word surface; the underlying mechanics are
meant to be C99, same as everything built earlier this session — Bob's
own words: "I think the confusion is you took the FORTH/C interpretation
too literally... My bad"), Bob proposed the concrete shape: "why not
just a generalized compudynamics C module and appropriate tuning knobs
for each of the 4 operational levels?" — confirmed as design-for-all-
four, instantiate-what-exists-now ("two levels now, blocks and messages
later... design for all 4"), then final go-ahead: "yes, build it. that
sounds exactly perfect."
New `include/compudynamics.h` / `src/compudynamics.c`:
- `cd_classify_ids()` — deterministic diversity/volatility/locality
classifier over a recent entity-ID touch sequence (word IDs for word
level today), purely execution-count-derived, zero wall-clock input.
Replaces `temporal_decay` with a `locality_q16` signal fed from
`rolling_window_get_recent_sequence()` (word level's existing
execution-history structure — reused, not duplicated) instead of
`decay_slope_q48`. The reward's joint-convergence signal now compares
tick-to-tick `locality_q16` deltas instead of slope deltas.
- A generic UCB1 config-space bandit (`CDConfigTable`/`CDConfigEntry`,
`cd_config_table_init/seed/tick/force`) lifted out of `ssm_jacquard.c`
verbatim in logic, malloc-sized to `CDTuning.num_configs`/
`num_regimes` rather than hardcoded 128/8.
- `CDTuning` knob struct per level: `cd_tuning_word()` (128 configs,
8 regimes, the exact constants `ssm_jacquard.c` used before this
module existed — a lift, not a retune) is the only level actually
wired; `cd_tuning_vm()` documents classifier-side parameters for a
possible future migration of VM-level regime classification off
Hera's proxy (`capsule_vm_physics.c`'s `l8_regime_modulation_q16()`,
unchanged — it transitively inherits determinism from reading Hera's
now-fixed `current_regime`, needing no direct edit); block and
message levels are not defined at all, deliberately, rather than
filled with invented numbers.
`ssm_jacquard.c` now delegates: `ssm_l8_state_t.table` is a
`CDConfigTable*` instead of a hand-rolled `SsmConfigTable*`;
`ssm_l8_update_table()`'s signature changed from
`(..., uint32_t current_window, uint64_t current_slope)` to
`(..., uint32_t current_window, const uint32_t *recent_word_ids,
uint32_t recent_word_ids_count)`. The legacy 16-mode threshold path
(`ssm_l8_update()`/`ssm_apply_mode()`, dormant except on table-malloc
failure) is untouched — it still reads `metrics.temporal_decay` exactly
as before, deliberately, to keep this fix's blast radius to the path
that was actually live and actually broken.
**A second bug, found chasing the first.** Wiring `cd_tuning_word()`'s
address across translation units (`&CD_TUNING_WORD` as an addressable
`extern const CDTuning`) crashed the kernel on a page fault at a
constant, reproducible wrong address (`CR2=0x0000000800000034`, always
base-plus-field-offset from the same wrong base) the instant
`ssm_l8_init_table()` reached across into `compudynamics.c`. Traced via
direct fault-address arithmetic (matching `CR2` offsets against
`CDTuning`'s field layout) to this freestanding kernel's boot-time
loader not reliably relocating the address of an extern const aggregate
referenced across translation units — confirmed by `readelf -r` on the
*final linked* kernel ELF showing thousands of un-applied relocation
entries (`.rela.text`/`.rela.rodata`/`.rela.data`), meaning this build
is not a normal fully-resolved EXEC image; something at boot is meant
to walk `.rela.dyn` and apply these, and whatever that mechanism is,
this particular relocation type wasn't landing correctly for a plain
data symbol's address. Cross-TU *function calls* — including struct
return-by-value, which uses a caller-local hidden pointer, never a
global's address — worked correctly throughout (confirmed: `malloc()`
and `cd_config_table_init()` both executed correctly across the same
TU boundary). Fix: `CD_TUNING_WORD`/`CD_TUNING_VM` are not addressable
globals at all — `cd_tuning_word()`/`cd_tuning_vm()` return `CDTuning`
by value. This is the aggregate-typed version of a constraint every
scalar tunable elsewhere in this codebase already satisfied by being a
`#define` immediate rather than a referenced global; not a workaround,
an accurate fix for a real, verified constraint of this build target.
**A process failure, corrected mid-fix.** Recovering from a stray log
directory (`rm -rf logs/*/amd64`, glob far too broad — deleted
committed historical audit logs back to June 21) was itself botched:
`git checkout -- .` was run to restore the deleted files, but that
command reverts *every* modified tracked file to HEAD, not just deleted
ones — discarding this session's in-progress `ssm_jacquard.c`/`.h`,
`vm_time.c`, `vm_runtime.c`, and `capsule_vm_physics.c` changes,
including the already-verified relocation fix above. `compudynamics.h`/
`.c` were untracked and survived; the reverted files were reconstructed
exactly from conversation record and reverified with a clean
zero-warning hosted rebuild before re-attempting kernel acceptance.
Both mistakes reported to Bob in full as they were discovered, per
standing instruction.
**Three-arch acceptance.** Hera and Hermes `dict_hash` are now
byte-identical across amd64/aarch64/riscv64
(`0x83c2c109100e2ed6` / `0x4159dcb326d79759`) — the fix holds. Artemis's
`dict_hash` matches on amd64/aarch64 (`0xc8cff7d7710a91fc`) but diverges
on riscv64 (`0x41a705a6b637d492`) — new since this change (rev q's own
riscv64 log shows Artemis at `0xc31f4c8e7ca40740`, matching amd64 at the
same commit, so this is not a pre-existing, already-accepted gap).
Root-cause hypothesis, not yet confirmed by a live instrumented run:
Hera and Hermes do no real I/O during boot — their capsule code is pure
computation, identical word sequence every time regardless of real
elapsed time. Artemis's capsule (`capsules/artemis/init.4th`) does real
`virtio-blk-pci` disk I/O at boot (header read, format-write, two full
alloc→fetch→persist→free round-trips), and
`src/starkernel/virtio/virtio_blk.c`'s completion wait
(`while (s->used->idx == s->last_used_idx) { ...; if (!--spin) return
BLKIO_EIO; }`, lines ~279-284) is a genuine busy-wait whose iteration
count depends on real QEMU device-emulation and TCG-translation timing
per architecture — meaning Artemis's real boot duration is
architecture-variable. Because rev r makes L8 rescore and potentially
reselect on *every* heartbeat tick, and the heartbeat thread fires on a
real wall-clock cadence (`HEARTBEAT_TICK_NS`) rather than an
execution-count cadence, any VM whose real boot duration varies by
architecture gets a correspondingly variable *number* of heartbeat
ticks during boot — which can flip which config is active by the time
boot finishes, independent of whether the classifier itself is
deterministic given fixed input. Under rev q, no VM ever lived long
enough in ticks to reach a trial boundary, so this tension was
structurally inert; rev r's correctness (scoring/reselecting on the
clock that actually exists) is what exposes it. Hera/Hermes apparently
don't hit this in practice — their config trajectory happens to stay
stable regardless of a few ticks' difference — while Artemis's heavier,
I/O-bound boot does. The accurate fix, if this hypothesis holds, is a
heartbeat-architecture decision (tying ticks to execution count rather
than wall-clock time, at least for config-reselection purposes) — out
of scope for this fix and explicitly deferred: Bob's direction was
"commit what's verified, log this as a known gap" rather than open-
ended investigation or a narrow Artemis-side patch.
Hosted build: clean, zero warnings, `-Wall -Werror`, throughout every
step above. Commit `07f72874`.
## Rev t: the Artemis/riscv64 gap was Loop #3, not Artemis — decay converted to tick-based, gap closed
Rev s's closing hypothesis (I/O-driven boot-duration variance flipping
which L8 config is active) was Bob's cue for a sharper correction:
"the Artemis riscv64 gap isn't really a 'riscv64' gap. it's a
compudynamic bug it sounds like. EVERYTHING runs off the adaptive
heartbeat period. there's no exceptions or hidden dependency other than
the adaptive heartbeat on the monotonic wall clock timer." Tracing
precisely rather than re-asserting the rev s framing: `dict_hash`
(`capsule_dict_hash_hook()`) hashes each dictionary entry's name and
`execution_heat` — nothing else. `execution_heat` is decayed by Loop #3
(`physics_metadata_apply_linear_decay()`), and Loop #3's decay amount
has always — since before this session — been computed from real
elapsed nanoseconds (`sf_monotonic_ns()` / `vm_monotonic_ns()` deltas
against `DictPhysics.last_active_ns`/`last_decay_ns`), not from any
tick count. That's the actual hidden dependency: not new, not
riscv64-specific, just newly *live* because rev r made L8 responsive
enough to real conditions to expose it (at rev q, no VM ever lived long
enough in ticks to reach the old trial boundary, so which config was
active — and therefore whether Loop #3 was even gated on — never
changed mid-boot, on any architecture; the wall-clock-tainted decay
path existed the whole time but had no opportunity to diverge).
Bob's diagnostic question, verbatim: "We need some sense of relative
time to compute slope? is that it?" — correctly identifying that decay
inherently needs *some* notion of elapsed time to express a rate, and
asking whether that's the whole story. Answer, confirmed by reading
`vm_tick()` (`vm_runtime.c`): `vm->heartbeat.tick_count` already *is*
a purely execution-driven virtual clock — incremented synchronously
from the interpreter's own dispatch path, gated by word-execution count
(`HEARTBEAT_CHECK_FREQUENCY` in the non-threaded/kernel case), never by
a timer interrupt. Artemis's `virtio_blk.c` completion busy-wait (a
real, timing-variable spin loop; see rev s) is plain C — it never calls
`vm_tick()`, so it advances `tick_count` by exactly zero regardless of
how long it spins in real time. The correct clock for Loop #3 already
existed in the codebase; Loop #3 simply wasn't using it. Loop #6
(decay-slope inference, `inference_engine.c`) was checked and confirmed
to have no direct wall-clock reads at all — the bug was isolated to
Loop #3's decay *application*, not also present in the slope's
*calibration*.
**Fix**, once confirmed ("yes... yup"): added `DictPhysics.last_decay_tick`
(`include/vm.h`) alongside the existing `last_active_ns`/`last_decay_ns`
(kept, but now diagnostics-only — `physics_diagnostic_words.c` prints
them for humans, a legitimate use this fix doesn't touch).
`physics_metadata_apply_linear_decay()`'s signature changed from
`(entry, elapsed_ns, vm)` to `(entry, elapsed_ticks, vm)`; internally,
`decay_amount = (elapsed_ticks * slope_q48) >> 16` replaces
`(elapsed_ns/1000 * slope_q48) >> 16` — no conversion constant between
"per microsecond" and "per tick" introduced, since Loop #6 adaptively
recalibrates `decay_slope_q48` against whatever units it's actually
applied in. The `elapsed_ns < DECAY_MIN_INTERVAL` insignificant-interval
gate became `elapsed_ticks == 0` — semantically identical purpose,
tick-native expression. Decay turned out to be applied at six call
sites across four files, not the one background-batch site originally
scoped: `physics_execution_hooks.c`'s `physics_pre_execute()` (every
word execution) and `physics_on_lookup()` (every word lookup by name),
the kernel's `vm_core.c` inner-interpreter loop with its own inline
duplicate of the same two sites, and the periodic background-sweep
decay in `vm_time.c`/`vm_runtime.c`. All six converted identically:
`elapsed_ticks = vm->heartbeat.tick_count - entry->physics.last_decay_tick`,
then `entry->physics.last_decay_tick = vm->heartbeat.tick_count` after.
A real behavioral consequence, not just a units change: tightly-looping
words (multiple touches within the same heartbeat-tick window) now
correctly see zero decay between touches instead of an accumulating
sub-tick real-time trickle — a word that hasn't gone idle for even one
full tick hasn't earned any decay, which is closer to the intended
"hot word stays hot" model than the old scheme's continuous real-time
erosion.
**Three-arch acceptance, decisive.** All four VM identities now
byte-identical across amd64/aarch64/riscv64: Hera `0x83c2c109100e2ed6`,
Artemis `0x5284ea5cd0f9983c` (the previously-diverging value — riscv64
now matches amd64/aarch64 exactly, closing the rev s gap completely),
Hermes `0x4159dcb326d79759` (both the original birth and the
kill/respawn birth), Mama `0xc88c3c1db6ef601b`. Hosted build clean,
zero warnings, throughout.
## Rev u: Kconfig migration, Phase 2 — real symbol tree + bridge, wired but inert
Separate track from the rev r/s/t decay work above: StarForth's ~25 live
build-time knobs were scattered across two Makefiles with drifted, conflicting
fallback defaults (e.g. `ADAPTIVE_SHRINK_RATE` = 50 in the Makefile, 75 in
`rolling_window_knobs.h`'s own independent fallback — never in sync, never
discoverable). Bob asked for a real Linux-style `menuconfig`/`xconfig`/
`kconfig` system, explicitly choosing to vendor the genuine Linux
`scripts/kconfig` tooling over a lighter Kconfiglib-style alternative, and to
cover both the hosted `Makefile` and `Makefile.starkernel` from the start.
Phase 0 (five dead knobs removed) and Phase 1 (`tools/kconfig/` vendored,
building all three frontends — `conf`/`mconf`/`qconf` — as standalone host
tools) are covered by prior commits. This entry covers Phase 2: the real
Kconfig symbol tree and the Makefile-side bridge that consumes it.
**Symbol tree** (`Kconfig` + `Kconfig.arch`/`Kconfig.variant`/
`Kconfig.physics`/`Kconfig.heartbeat`, four `source`-included submenus):
`ARCH` as a three-way choice (amd64/aarch64/riscv64) with a derived
`ARCH_STRING`; `STARFORTH_VARIANT_HOSTED`/`STARFORTH_VARIANT_KERNEL` as the
top-level choice gating the hosted-only `TARGET=` profile choice and
platform-mode choice inside an `if STARFORTH_VARIANT_HOSTED` block (the
kernel build has no equivalent of either); the physics/SSM knob family
(`STRICT_PTR`, `ENABLE_HOTWORDS_CACHE`, `ENABLE_PIPELINING`,
`TRANSITION_WINDOW_SIZE` depending on `ENABLE_PIPELINING`, the `ADAPTIVE_*`
window-shrink family, `INITIAL_DECAY_SLOPE_Q48`, `DECAY_RATE_PER_US_Q16`,
`HEARTBEAT_INFERENCE_FREQUENCY`); and the heartbeat family
(`HEARTBEAT_THREAD_ENABLED`, `HEARTBEAT_TICK_NS` depending on it,
`HEARTBEAT_CHECK_FREQUENCY`, `HEARTBEAT_WINDOW_TUNING_FREQUENCY`,
`HEARTBEAT_SLOPE_VALIDATION_FREQUENCY`, `EMERGENCY_CONSOLE_ENABLED`). Every
default was hand-diffed against the actual current Makefile/header default
for that symbol, not copied from either Makefile's stale comments — this
caught and preserved the same `ADAPTIVE_SHRINK_RATE`/`ENABLE_HOTWORDS_CACHE`/
`ENABLE_PIPELINING` drift-vs-comment mismatches already known from Phase 0/1
scoping, documented in-line in `Kconfig.physics` rather than silently
"corrected" (Phase 3's job, not Phase 2's).
**One known, deliberately-preserved discrepancy**, documented directly in
`Kconfig.heartbeat`'s help text: `Makefile.starkernel:249` unconditionally
appends `-DHEARTBEAT_THREAD_ENABLED=0` after knob forwarding, regardless of
what's requested — LithosAnanke has no pthreads. `HEARTBEAT_THREAD_ENABLED`
is modeled here as an ordinary hosted-side bool (default y), so a
kernel-variant `.config` will show it as `y` even though the actual kernel
build always forces it off. Resolving this (either a `depends on
!STARFORTH_VARIANT_KERNEL` restriction or a documented kept override) is
Phase 4's named verification target, not Phase 2's.
**Bridge** (`mk/Kconfig.mk`, `include`d by both Makefiles): per-architecture
config under `build/$(ARCH)/.config` rather than one root `.config` — the
project's own QEMU acceptance workflow builds all three kernel architectures
back-to-back in the same working directory, so a shared root config would
desync from whichever arch is actually being built. The two Makefiles
canonicalize `aarch64` differently for their own build-directory naming
(hosted `Makefile`: `aarch64` → `arm64`; `Makefile.starkernel`: `aarch64` →
`aarch64`), so `mk/Kconfig.mk` doesn't re-derive an architecture name at all
— it requires the including Makefile to set `KCONFIG_ARCH_DIR` to whatever
directory name *it* already builds objects under, immediately before the
`include`.
Two things confirmed only by direct observation while building this, not
foreseeable from reading `conf --help`:
1. **`conf`/`confdata.c` reads four separate env vars**, not the one
(`KCONFIG_CONFIG`) the Phase 1 plan had anticipated: `KCONFIG_CONFIG`
(`.config` itself), `KCONFIG_AUTOCONFIG` (`include/config/auto.conf`),
`KCONFIG_AUTOHEADER` (`include/generated/autoconf.h`), and
`KCONFIG_RUSTCCFG` (`include/generated/rustc_cfg`, written unconditionally
even though StarForth has no Rust code). Missing any one of the four lets
that file fall back to a bare `include/config`/`include/generated` path
relative to `$(CURDIR)` — i.e. leaking generated files straight into the
real source tree. Caught twice during validation (once via an interactive
`mconf` test run in Phase 1, once again here via a `%_defconfig` run that
only set three of the four vars) before all four were set together in
`mk/Kconfig.mk`.
2. **Including `mk/Kconfig.mk` before either Makefile's own `all:` target
silently hijacked Make's default goal.** `mk/Kconfig.mk`'s first rule
(`$(KCONFIG_CONF):`, building the `conf` binary) became the *first rule
Make had seen anywhere*, at the point `mk/Kconfig.mk` needed to be
included (immediately after `ARCH` is resolved, well before either
Makefile's own `all:` appears later in the file) — so a bare `make` with
no target built `tools/kconfig/conf` and stopped, never touching `all`.
Caught by literally running `make` after wiring the include and getting
`make: 'tools/kconfig/conf' is up to date.` as the *entire* output. Fixed
with an explicit `.DEFAULT_GOAL := all` in both Makefiles at the
`include mk/Kconfig.mk` site, immune to include order.
**Genuinely inert until opted into.** The `include $(KCONFIG_AUTOCONF_FILE)`
line — the one that would actually pull `CONFIG_*` variables into a Makefile
— is itself guarded behind `ifneq ($(wildcard $(KCONFIG_CONFIG_FILE)),)`. No
`build/$(ARCH)/.config` exists until a developer explicitly runs
`menuconfig`/`xconfig`/`config`/`oldconfig`/a `*_defconfig` target; until
then this guard is false and `mk/Kconfig.mk` contributes zero behavior
change to any existing invocation. This is what makes "wired but inert" true
in practice rather than true in name only — it was verified, not assumed,
by running `make help`/`make clean`/`make` against both Makefiles with no
`.config` present anywhere and confirming byte-identical output and behavior
to before this phase.
Four example defconfigs added (`configs/hosted_standard_defconfig`,
`configs/kernel_{amd64,aarch64,riscv64}_defconfig`), each just the minimal
choice selections (`ARCH_*`, `STARFORTH_VARIANT_*`, `TARGET_STANDARD`,
`PLATFORM_DEFAULT` where applicable) needed to steer `conf --defconfig`;
every other symbol resolves through its Kconfig `default` and was hand-diffed
against today's real Makefile/header defaults for all four combinations
before being trusted.
**No `-D` flag repointed.** Confirmed by grepping the actual link command a
plain hosted `make` emits — still bare `$(VAR)`-driven
(`-DSTRICT_PTR=1 -DENABLE_HOTWORDS_CACHE=0 ...`), no `CONFIG_` prefix
anywhere. Repointing specific knob families to `$(CONFIG_VAR)` is Phase 3+,
one family at a time.
**Three-arch acceptance.** Both Makefiles touched (the `.DEFAULT_GOAL` fix
and the bridge `include` land in `Makefile.starkernel` too), so the full
three-arch QEMU run was required before commit per this project's standing
rule. All four VM identities landed byte-identical across amd64/aarch64/
riscv64, matching rev t's values exactly: Hera `0x83c2c109100e2ed6`,
Artemis `0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births),
Mama `0xc88c3c1db6ef601b`. Hosted build clean, zero warnings.
## Rev v: Kconfig migration, Phase 3 — SSM/physics knob family cutover
First "live" phase: the ~18-symbol physics/SSM family (`STRICT_PTR`,
`ENABLE_HOTWORDS_CACHE`, `ENABLE_PIPELINING`, `ROLLING_WINDOW_SIZE`,
`TRANSITION_WINDOW_SIZE`, the four `ADAPTIVE_*` window-shrink knobs,
`INITIAL_DECAY_SLOPE_Q48`, `DECAY_MIN_INTERVAL`, `DECAY_RATE_PER_US_Q16`,
`HEARTBEAT_INFERENCE_FREQUENCY`, plus five previously-unwired `SSM_*` L8
thresholds) now actually reads from `$(CONFIG_VAR)` when a `.config`
exists, instead of Phase 2's "bridge present but nothing consumes it."
**Bridge macros** (`mk/Kconfig.mk`, `kconfig_bool`/`kconfig_int`): both
compile down to exactly `$(VAR) ?= $(default)` — today's behavior,
unchanged — whenever `KCONFIG_ACTIVE` is unset (no `.config` yet); once
active, `kconfig_bool` maps Kconfig's `y`/absent to `1`/`0` (the C side has
always taken 0/1 integers, never y/n), and `kconfig_int` takes
`$(or $(CONFIG_VAR),$(default))`, covering both "Kconfig inactive" and
"Kconfig active but this symbol's `depends on` was unmet so it never
appears in `auto.conf` at all" (e.g. `TRANSITION_WINDOW_SIZE` when
`ENABLE_PIPELINING=n`) with the same historical fallback rather than an
empty `-D`. Command-line/environment overrides of a knob always win
regardless of which branch fires — ordinary Make semantics, not something
this bridge has to implement itself. Verified directly: `make -n
ADAPTIVE_SHRINK_RATE=99` with an active `.config` requesting 77 still
produces `-DADAPTIVE_SHRINK_RATE=99`.
**Two build systems, two different existing philosophies, same bridge.**
The hosted `Makefile` blanket-forwards every knob unconditionally
(`?=` always emits a `-D`); each bare `?=` line became one
`$(eval $(call kconfig_bool/int,VAR,default))` call in place. The kernel
`Makefile.starkernel` only forwards a knob if a developer explicitly set
it (`origin($1) != undefined`) — Phase 3 preserves that asymmetry exactly:
a new `ifeq ($(KCONFIG_ACTIVE),1)` block pre-assigns each knob from its
`CONFIG_` value *before* the existing `add_vm_flag` foreach loop runs, so
`origin()` becomes `"file"` and the knob gets forwarded as if a developer
had typed it — but only when Kconfig is actually active. With no
`.config`, kernel-build behavior for this family is unchanged: nothing
forwarded unless explicitly requested, same as every kernel build before
this migration.
**A real, live bug found and fixed, not just a documented drift.**
`ADAPTIVE_SHRINK_RATE` (75 vs the Makefile's 50) and — newly discovered
here — `ADAPTIVE_GROWTH_THRESHOLD` (1 vs 5) had conflicting fallback
values in `rolling_window_knobs.h`, but both were confirmed *dormant* for
every consumer (`doe_metrics.c`, `rolling_window_of_truth.c`): both
transitively include `vm.h` before `rolling_window_knobs.h`, so
`starforth_config.h`'s `#ifndef` always wins the race regardless of
whether a `-D` flag was present. `TRANSITION_WINDOW_SIZE` (2 vs 8) in
`physics_pipelining_metrics.h` was not dormant: `physics_pipelining_metrics.c`
includes `physics_pipelining_metrics.h` *before* `vm.h`, and the kernel
Makefile's opt-in-only forwarding means no `-D` flag exists for this
symbol unless a developer explicitly sets one — so a stock kernel build
compiling that one translation unit was silently getting `2`, not the
intended `8`, whenever pipelining happened to be on. Never observed in
practice only because `ENABLE_PIPELINING` defaults off. Fix, applied
uniformly to all three duplicate-fallback headers (`rolling_window_knobs.h`,
`physics_pipelining_metrics.h`, and — extending the same treatment for
consistency, since starforth_config.h's own file-header comment declares
it "the single source of truth for VM build-time defaults" — `ssm_jacquard.h`,
which was not itself drifted but was a second independent copy of the same
five values): delete each header's own `#ifndef X #define X (own value)
#endif` block, add `#include "starforth_config.h"` at the top instead.
Confirmed safe: `starforth_config.h` has zero includes of its own, and
`vm.h` explicitly avoids including either of the other two headers
("avoids circular include" comments already in place), so no cycle risk.
**SSM_\* thresholds promoted, not just documented.** `ssm_jacquard.h`'s
five L8 mode-selector constants (`SSM_ENTROPY_HIGH_THRESHOLD`,
`SSM_CV_HIGH_THRESHOLD`, `SSM_TEMPORAL_DECAY_THRESHOLD`,
`SSM_TEMPORAL_DECAY_LOW_THRESHOLD`, `SSM_HYSTERESIS_TICKS`) had no Makefile
knob and no `-D` forwarding at all before this phase — genuinely unwired,
per the migration plan's own framing, "lowest risk since nothing currently
overrides them." Now wired identically to every other physics knob in both
Makefiles. The four threshold values are C `double`s compared against
`ssm_config_t` fields also typed `double`; Kconfig has no native float
symbol type, so they're modeled as Kconfig `string` symbols
(`default "0.75"` etc.) rather than `int`. Confirmed by direct
experimentation that this needs no quote-stripping on the Make side:
`conf`'s `.config` output keeps a string default quoted
(`CONFIG_FOO="0.75"`), but the `--syncconfig`-generated `auto.conf` that
Make actually `include`s writes string defaults unquoted
(`CONFIG_FOO=0.75`) — deliberately Make-syntax-safe. This means
`kconfig_int` (originally written only for true `int` symbols) already
handles these `string`-typed-but-numeric-literal symbols correctly with no
changes, and StarForth ended up not needing the separate `kconfig_str`
macro/quote-stripping logic originally sketched for this.
**Verification.** Every repointed default was hand-diffed against the
knob's actual pre-Phase-3 value (not the sometimes-stale Makefile comment
beside it — several comments described a different number than the `?=`
line actually set, e.g. `TRANSITION_WINDOW_SIZE`'s comment said "Default: 2"
next to a `?= 8` line; comments corrected in the same edit). End-to-end
round-trip confirmed by hand: generating a `.config` with
`ADAPTIVE_SHRINK_RATE=77`/`SSM_ENTROPY_HIGH_THRESHOLD="0.42"` and observing
`make -n` emit exactly those values in the link command; reverting and
confirming the inactive path emits the identical `-D` flag list, in the
identical order, as the pre-Phase-3 baseline. Both Makefiles touched
(`Makefile.starkernel`'s knob-forwarding block changed), so the full
three-arch QEMU acceptance ran again: all four VM identities landed
byte-identical to rev t/u's baseline — Hera `0x83c2c109100e2ed6`, Artemis
`0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births), Mama
`0xc88c3c1db6ef601b`. Hosted build clean, zero warnings, `-D` flag list
for a plain `make` byte-for-byte identical to before this phase aside from
the five newly-added `SSM_*` flags (new, not a repointing of anything that
existed before).
**One pre-existing, out-of-scope finding, reported rather than fixed per
standing instruction not to fix unrequested bugs:** `DECAY_MIN_INTERVAL`
has a `?=` default in the hosted Makefile but was never actually forwarded
as a `-D` flag in `BASE_CFLAGS` — the Make variable exists, but changing it
via `make DECAY_MIN_INTERVAL=999` has always had zero effect on
compilation, independent of and predating the tick-based decay conversion
(rev t) that made the value non-load-bearing at the C-logic level too.
Left exactly as found; not this phase's job to fix.
## Rev w: Kconfig migration, Phase 4 — heartbeat family cutover, HEARTBEAT_THREAD_ENABLED gate resolved
Repointed `EMERGENCY_CONSOLE_ENABLED`, `HEARTBEAT_THREAD_ENABLED`,
`HEARTBEAT_TICK_NS`, `HEARTBEAT_CHECK_FREQUENCY`,
`HEARTBEAT_WINDOW_TUNING_FREQUENCY`, `HEARTBEAT_SLOPE_VALIDATION_FREQUENCY`
through the same `kconfig_bool`/`kconfig_int` bridge as Phase 3's physics
family, in both Makefiles. Promoted the three `HEARTBEAT_CHECK_FREQUENCY`/
`HEARTBEAT_WINDOW_TUNING_FREQUENCY`/`HEARTBEAT_SLOPE_VALIDATION_FREQUENCY`
knobs into the hosted Makefile for the first time -- they had opt-in
forwarding in `Makefile.starkernel` already, but no `?=` and no `-D`
forwarding at all in the hosted `Makefile` before this phase, so hosted
builds always got `starforth_config.h`'s bare defaults with no override
mechanism. Same "promote a previously-unwired knob" treatment as Phase 3's
`SSM_*` constants.
**The named verification gate.** `Makefile.starkernel`'s
`VM_FEATURE_OVERRIDES += -DHEARTBEAT_THREAD_ENABLED=0` unconditionally
forces the kernel build's heartbeat thread off after knob forwarding,
regardless of what was requested -- LithosAnanke is freestanding, no
pthreads. The plan asked this phase to decide explicitly between modeling
it as a Kconfig `depends on` restriction or keeping the hardcoded
post-override. Chose both, deliberately: added
`depends on STARFORTH_VARIANT_HOSTED` to `HEARTBEAT_THREAD_ENABLED` in
`Kconfig.heartbeat` (so a kernel-variant `.config` never shows or sets it
-- `CONFIG_HEARTBEAT_THREAD_ENABLED` simply doesn't appear in that config's
`auto.conf`, and the `kconfig_bool` bridge correctly resolves the absence
to `0` with no special-casing needed), *and* kept
`Makefile.starkernel`'s own hardcoded override exactly as it was. The
`depends on` gets the Kconfig model right (menuconfig can't offer a choice
the kernel has no way to honor); the override guarantees the fact holds
even if the Kconfig tree, a hand-edited `auto.conf`, or a future refactor
ever disagrees. Belt-and-suspenders on purpose -- a freestanding kernel
image linking against pthreads that don't exist is a failure mode that's
silent until boot and expensive to debug on real hardware, not something
to trust to a single layer.
Verified directly, both branches: with no `.config`, the kernel build's
`-D` list contains only `-DHEARTBEAT_THREAD_ENABLED=0` (from the hardcoded
override) and no other heartbeat flags — identical to before this phase.
With a `kernel_amd64_defconfig`-derived `.config` active,
`-DHEARTBEAT_THREAD_ENABLED=0` still appears (now doubly-sourced: the
`depends on`-driven absence resolving to `0` via `kconfig_bool`, *and* the
hardcoded override — both agree, no redefinition conflict), while
`HEARTBEAT_TICK_NS`/`CHECK_FREQUENCY`/`WINDOW_TUNING_FREQUENCY`/
`SLOPE_VALIDATION_FREQUENCY` all forward correctly with their Kconfig
defaults. The hosted build's active path was checked too, confirming
`HEARTBEAT_THREAD_ENABLED=1` there (unaffected by the kernel-only
`depends on`, since `STARFORTH_VARIANT_HOSTED` config properly makes the
symbol visible again). `HEARTBEAT_TICK_NS` forwarding a fallback value for
a kernel build even though its own `depends on HEARTBEAT_THREAD_ENABLED`
also makes it Kconfig-invisible there is intentional, not a bug -- same
precedent as Phase 3's `TRANSITION_WINDOW_SIZE`-under-disabled-pipelining
case: `kconfig_int`'s `$(or ...)` always falls back to the historical
literal default rather than an empty `-D`, and the C side never reads
`HEARTBEAT_TICK_NS` when `HEARTBEAT_THREAD_ENABLED` is `0` regardless, so
this is a harmless no-op, not a behavior change.
**Repeated the same operational mistake from Phase 1/2, twice more, while
writing this phase.** A direct ad-hoc `tools/kconfig/conf --defconfig=...`
call used only to re-verify the new `depends on` relationship (not routed
through `make`) set `KCONFIG_CONFIG` but not the other three required env
vars, leaking `include/config/`+`include/generated/` into the real source
tree again -- the exact class of mistake already documented as a lesson in
rev u. Caught and cleaned up both times via `git status` before staging
anything. Recording again, more bluntly this time: there is no safe
shorthand for a one-off `conf` invocation outside `make` -- all four
(`KCONFIG_CONFIG`, `KCONFIG_AUTOCONFIG`, `KCONFIG_AUTOHEADER`,
`KCONFIG_RUSTCCFG`) or none of the ad-hoc-testing convenience is worth it.
**Verification.** Both Makefiles touched, so the full three-arch QEMU
acceptance ran again: all four VM identities byte-identical to the
established baseline -- Hera `0x83c2c109100e2ed6`, Artemis
`0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births), Mama
`0xc88c3c1db6ef601b`. Hosted build clean, zero warnings.
## Rev x: Kconfig migration, Phase 5 — platform-mode choice + remaining pipelining constants
Two independent pieces of work, both scoped to this phase by the plan.
**Platform-mode choice wired.** `Kconfig.variant`'s `PLATFORM_DEFAULT`/
`PLATFORM_MINIMAL`/`PLATFORM_L4RE` choice (written inert in Phase 2) now
actually drives the hosted `Makefile`'s `MINIMAL`/`L4RE` variables when a
`.config` selects something other than default. This family can't reuse
`kconfig_bool`/`kconfig_int`: `MINIMAL`/`L4RE` are tested with `ifdef`
downstream, which cares about *definedness*, not value, so unconditionally
defining either to 0 the way `kconfig_bool` does would make `ifdef` true
regardless of which platform was actually selected. Wrote the bridge
directly instead: only assign `MINIMAL := 1` or `L4RE := 1` (never both;
never at all for `PLATFORM_DEFAULT`) when Kconfig is active *and* neither
variable already has an explicit command-line/environment origin -- so
`make L4RE=1` always wins even if `.config` says `PLATFORM_MINIMAL`, a
case ordinary Make command-line precedence doesn't cover on its own since
it protects a variable from being overridden by an assignment to *itself*,
not from a *different* variable also becoming defined and winning an
`ifdef`/`else ifdef` chain.
**A real ordering bug, caught by testing, not by inspection.** First
placement of this bridge was directly above the `ifdef MINIMAL` platform
block near the bottom of the file -- seemed natural, right next to the
code it feeds. Testing a `PLATFORM_L4RE` config immediately showed
`HEARTBEAT_THREAD_ENABLED=1` and `-pthread` still present in the link
command, when both should be suppressed once `L4RE` is set. Root cause:
Make evaluates a file top-to-bottom, and `L4RE`'s existing
`ifeq ($(L4RE),1) HEARTBEAT_THREAD_ENABLED := 0 endif` heartbeat-override
check sits much earlier in the file (right after the heartbeat knobs) --
so at the point that check ran, `L4RE` hadn't been set yet by a bridge
placed near the bottom. Fixed by moving the whole platform-mode bridge to
immediately after `include mk/Kconfig.mk`, before anything else in the
file branches on `MINIMAL`/`L4RE`. Verified all three cases directly by
generating a `.config` for each and inspecting `make -n`'s actual command
line: `PLATFORM_MINIMAL` produces `-DSTARFORTH_MINIMAL=1 -nostdlib
-ffreestanding`; `PLATFORM_L4RE` produces `-D__l4__=1`,
`-DHEARTBEAT_THREAD_ENABLED=0`, and no `-pthread`; `PLATFORM_DEFAULT`
produces neither; and `make L4RE=1` against a `PLATFORM_MINIMAL` config
produces `-D__l4__=1` with no `STARFORTH_MINIMAL`/`ffreestanding` at
all -- command-line intent wins cleanly.
**Remaining `physics_pipelining_metrics.h` constants promoted.**
`SPECULATION_THRESHOLD_Q48`, `SPECULATION_DEPTH`,
`MIN_SAMPLES_FOR_SPECULATION`, `MISPREDICTION_COST_Q48`, and
`MINIMUM_PREFETCH_ROI` were plain, unconditional `#define`s with no
override mechanism at all before this phase (unlike `TRANSITION_WINDOW_SIZE`
and `ENABLE_PIPELINING` in the same file, already `#ifndef`-guarded and
cut over in Phase 3). Converted all five to the established pattern:
`#ifndef`-guarded, default sourced from `starforth_config.h`, wired through
`kconfig_int` in both Makefiles as Kconfig `hex` (the three Q48.16-encoded
values) or `int` (the two plain counts) symbols, gated `depends on
ENABLE_PIPELINING` alongside `TRANSITION_WINDOW_SIZE`. Verified the header
change alone (no Kconfig involved) with a clean hosted build at both
`ENABLE_PIPELINING=0` (today's default) and `=1`, before touching either
Makefile -- zero warnings both ways, confirming the `-D`-flag-always-wins
`#ifndef` pattern holds regardless of pipelining state.
**A second real bug found, not fixed.** Computing the exact Q48.16
encodings to transcribe into Kconfig defaults (rather than trusting the
existing C comments) surfaced a second drift bug, independent of anything
found in Phase 3: `MINIMUM_PREFETCH_ROI`'s own comment says
`1.10 = 1.10 * (1 << 16) = 0x11999AL`, but `1.10 * 65536 = 72089.6 ≈
0x1199A` (5 hex digits) -- the shipped constant `0x11999AL` (6 hex digits)
is `1153434`, which is `≈17.6` in Q48.16, not `1.10`. An order-of-magnitude
error that's been silently shipping in every build with pipelining enabled
since this constant was introduced, since ROI comparisons like
`(prefetch_latency_saved / total_attempts) > MINIMUM_PREFETCH_ROI` would
essentially never trigger speculation at 17.6 as the bar instead of 1.10.
Per standing instruction not to fix bugs found while doing unrelated work,
preserved the exact shipped value (`0x11999A`) as the Kconfig default and
`starforth_config.h` fallback, with an explicit comment at both sites
documenting the discrepancy and pointing back here. Flagging directly:
**Bob may want to fix `MINIMUM_PREFETCH_ROI` to `0x1199A` in a future
session** -- not done as part of this migration.
**Verification.** Both Makefiles touched, so the full three-arch QEMU
acceptance ran again: all four VM identities byte-identical to the
established baseline -- Hera `0x83c2c109100e2ed6`, Artemis
`0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births), Mama
`0xc88c3c1db6ef601b`. Hosted build clean, zero warnings.
## Rev y: MINIMUM_PREFETCH_ROI fixed on Bob's explicit instruction
Rev x reported, but deliberately did not fix, that `MINIMUM_PREFETCH_ROI`
had been shipping as `0x11999A` (~17.6 in Q48.16) against a documented
intent of `1.10` (correct encoding `0x1199A`) -- an order-of-magnitude
error that's been silently suppressing prefetch speculation (the ROI bar
was ~16x higher than intended) in every build with `ENABLE_PIPELINING=1`
since the constant was introduced. Bob's instruction: fix it, then
continue to Phase 6.
Corrected the value at all four sites that had inherited the wrong
default while it was merely being made overridable (not yet fixed) in
Phase 5: `include/starforth_config.h`'s
`STARFORTH_CONFIG_MINIMUM_PREFETCH_ROI_DEFAULT`, `Kconfig.physics`'s
`MINIMUM_PREFETCH_ROI` symbol default, and the `kconfig_int` fallback
literal in both `Makefile` and `Makefile.starkernel`. Updated each site's
comment from "preserved as shipped, not fixed" to a plain statement of the
correct derivation, removing the now-stale "flagged, not fixed" language.
`physics_pipelining_metrics.h`'s own doc comment already stated the
correct `0x1199A` derivation (only its two upstream defaults were wrong),
so needed no correction, just removal of the discrepancy note pointing at
`starforth_config.h`.
Verified: a `.config` with `ENABLE_PIPELINING=y` now reports
`CONFIG_MINIMUM_PREFETCH_ROI=0x1199A`; a hosted build with
`ENABLE_PIPELINING=1` emits `-DMINIMUM_PREFETCH_ROI=0x1199A` in the actual
link command, zero warnings. `grep -rn 11999A` across the tree returns
only the historical-note comments explaining the fix, no live default.
Both Makefiles touched (again), so the full three-arch QEMU acceptance ran
once more: byte-identical to the established baseline -- Hera
`0x83c2c109100e2ed6`, Artemis `0x5284ea5cd0f9983c`, Hermes
`0x4159dcb326d79759` (both births), Mama `0xc88c3c1db6ef601b`. Unsurprising
that dict_hash is unaffected either way -- `MINIMUM_PREFETCH_ROI` governs
pipelining speculation decisions, not dictionary word heat, and the
kernel's default build has `ENABLE_PIPELINING=0` regardless -- but the
acceptance bar is unconditional per project rule, so it ran anyway.
## Rev z: Kconfig migration, Phase 6 — kernel-only flags, plan complete
Final phase. `PARITY_MODE`, `STARFORTH_ENABLE_VM`, `SK_PARITY_DEBUG`,
`HEARTBEAT_DOE_LOG` -- all kernel-only, no hosted equivalent -- repointed
through the Kconfig bridge in a new `Kconfig.kernel` (sourced from the
root `Kconfig`, wrapped in `if STARFORTH_VARIANT_KERNEL ... endif` so
these four symbols simply don't exist for a hosted-variant `.config`).
`STARFORTH_ENABLE_VM` and `PARITY_MODE` had bare `?=` defaults already
(blanket-forwarded); repointed via `kconfig_bool` in place, same as every
other blanket-forwarded knob this migration. `HEARTBEAT_DOE_LOG` likewise.
`SK_PARITY_DEBUG` had no bare default at all -- opt-in-only via
`VM_FEATURE_FLAG_VARS`, and (per this project's convention for that
family) pre-assigned from its `CONFIG_` value inside the existing
`ifeq ($(KCONFIG_ACTIVE),1)` block, same treatment as the `SSM_*`/
pipelining constants in Phases 3/5.
**Doesn't touch the hosted `Makefile` at all**, exactly as the plan
specified -- confirmed directly, not just by omission: `git status` after
this phase's edits shows only `Kconfig`, `Kconfig.kernel` (new), and
`Makefile.starkernel` changed, and a full hosted `make clean && make` ran
clean with zero warnings and a byte-identical link command to before this
phase. This doubles as the regression check the plan asked for: five
phases of Kconfig work landing entirely inside `Makefile.starkernel`,
`mk/Kconfig.mk`, and the `Kconfig*` tree, with the hosted build path
provably undisturbed.
Verified both branches on the kernel side directly, not just by
inspection: with no `.config`, the kernel `-D` list is
`PARITY_MODE=0 STARFORTH_ENABLE_VM=1 HEARTBEAT_DOE_LOG=1` and no
`SK_PARITY_DEBUG` flag at all -- byte-identical to pre-Phase-6 behavior.
With a `kernel_amd64_defconfig`-derived `.config`, the same defaults
appear plus `-DSK_PARITY_DEBUG=0` (now forwarded, correctly at its
default). A hand-built `.config` flipping all four
(`PARITY_MODE=y SK_PARITY_DEBUG=y HEARTBEAT_DOE_LOG=n`) produced exactly
`-DPARITY_MODE=1 -DSK_PARITY_DEBUG=1 -DHEARTBEAT_DOE_LOG=0
-DSTARFORTH_ENABLE_VM=1` in the actual `make -n` command line.
Both Makefiles touched by the migration as a whole across all six phases,
kernel-only this phase specifically, so the full three-arch QEMU
acceptance ran one final time: all four VM identities byte-identical to
the baseline established at rev t and held through every phase since --
Hera `0x83c2c109100e2ed6`, Artemis `0x5284ea5cd0f9983c`, Hermes
`0x4159dcb326d79759` (both births), Mama `0xc88c3c1db6ef601b`. Hosted
build clean, zero warnings.
**Kconfig migration plan complete.** Six phases, six commits (plus one
out-of-band fix commit at Bob's request between Phase 5 and Phase 6),
zero behavioral regressions at any step, two real pre-existing bugs found
and reported (`TRANSITION_WINDOW_SIZE`'s live-not-latent kernel-build
drift, Phase 3; `MINIMUM_PREFETCH_ROI`'s order-of-magnitude Q48.16
encoding error, Phase 5), one of the two fixed on explicit instruction
(rev y). Every physics/SSM/heartbeat/pipelining/kernel-only knob that had
a Makefile presence before this migration still has one, now sourced from
a single discoverable Kconfig tree (`Kconfig` + `Kconfig.arch`/
`Kconfig.variant`/`Kconfig.physics`/`Kconfig.heartbeat`/`Kconfig.kernel`)
when a developer opts in via `menuconfig`/`xconfig`/`config`/`oldconfig`/
a `*_defconfig` target, and unchanged when they don't. `CDTuning` values
(`compudynamics.h`) and runtime-only DoE/profiling parameters remain
explicit non-goals, as scoped from the start.