2066 lines
121 KiB
Markdown
2066 lines
121 KiB
Markdown
# VM Fleet Attractor Experiment — Design Doc
|
||
**Date:** 2026-07-05 (rev g)
|
||
**Branch:** `lithosananke`
|
||
**Status:** Phase 1 and Phase 2 implemented and verified three-arch. Rev f
|
||
found the "100% heat at Hera" result from revs b/d/e was never an attractor
|
||
finding — a seeding bug left the redistribution mechanism completely inert.
|
||
Rev g fixed that seed and a second bug it exposed (wrong time-source in
|
||
`vm_physics_touch`'s call sites), both verified amd64-clean, but found a
|
||
third, deeper, pre-existing bug underneath in `src/starkernel/hal/
|
||
host_services.c`'s `kernel_monotonic_ns()` — it doesn't handle the
|
||
RELATIVE-timer-mode fallback the rest of the kernel already uses when TSC
|
||
calibration fails (as it does under this QEMU/TCG environment). **Not yet
|
||
fixed** — likely affects word-level background decay too, not just VM-fleet
|
||
physics. Phase 3 stays blocked until heat can actually be observed moving.
|
||
**Author:** Captain Bob / Claude Code
|
||
|
||
---
|
||
|
||
## Research question
|
||
|
||
Does the VM fleet — Hera, Hermes, Artemis, and any future children — find an
|
||
attractor the same way individual FORTH words do?
|
||
|
||
This is not a rhetorical question. It's directly motivated by the
|
||
**L8 Attractor Map campaign** (`docs/working/experiments/campaigns/l8_attractor_map/SESSION_REPORT_2025-12-10.md`,
|
||
180 runs, 6 workloads × 30 replicates): word-level execution heat, under the
|
||
existing physics engine (Loops #1–#7), was empirically observed to converge
|
||
deterministically to a steady-state "coldest" configuration — low
|
||
coefficient-of-variation, workload-independent convergence time
|
||
(23.3 ± 2.61 ticks), and a candidate universal oscillation frequency
|
||
(ω₀ ≈ 13.5 Hz, later suspected to be CPU-clock-linked rather than a true
|
||
constant). That finding, plus the independent 38,400-run 2^7 factorial DoE,
|
||
is the empirical basis for the whole "physics-grounded adaptive runtime"
|
||
claim this project is built on.
|
||
|
||
`capsule_vm_physics.c` (added this session, see
|
||
`VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md`) is a deliberate structural
|
||
mirror of that same mechanism — per-entity execution heat, a rolling
|
||
touch-history window, regression-inferred slope — applied to VMs instead of
|
||
words. The mirroring was designed in, but never tested for the property that
|
||
motivated the original mechanism: does it converge? Does the fleet's heat
|
||
distribution settle into a repeatable steady state, the way word heat did?
|
||
Nobody knows yet. This doc is step one toward finding out.
|
||
|
||
## Structural note: the shared-slope question is not a blocker
|
||
|
||
An earlier framing of this question (see prior conversation) suggested that
|
||
`fleet_transfer_slope_q48` being a single fleet-wide value, rather than
|
||
per-VM, might mean the mechanism *can't* produce genuine per-VM attractors.
|
||
That framing was checked against the word-level mechanism and doesn't hold:
|
||
`decay_slope_q48` (`include/vm.h:504`) is likewise a single global value
|
||
applied uniformly to every word's individually-varying execution heat — one
|
||
shared decay rate, many independently-converging heat values. That's exactly
|
||
the structure the L8 campaign found attractor behavior in. So the VM-fleet
|
||
mechanism has the same shape, and the question is genuinely open rather than
|
||
foreclosed by the shared slope — worth testing, not worth assuming either way.
|
||
|
||
## Design goal: transparency first
|
||
|
||
Per explicit direction: **transparency is a design goal of this system, not
|
||
an incidental nice-to-have.** Before any experiment can answer "does the
|
||
fleet find an attractor," the mechanism has to be fully observable —
|
||
individual VM state, not just an aggregate pass/fail. Today it is not:
|
||
|
||
`vm_physics_status()` (`capsule_vm_physics.c`) prints exactly four things:
|
||
fleet heat sum, conservation verdict (CONSERVED/DRIFTED), the one shared
|
||
`fleet_transfer_slope_q48`, its fit quality, and whether the touch-history
|
||
window is warm. **No per-VM heat is ever surfaced.** You cannot currently
|
||
ask "what is Hermes's heat right now" from FORTH or from a boot log — only
|
||
"what is the fleet's total" (which conservation forces to Q.1 always, so it
|
||
carries no information about *distribution*).
|
||
|
||
This is the opposite of transparent. A conservation check that always reads
|
||
"1.0" tells you nothing about whether the fleet is balanced 33/33/33 or
|
||
99/0.5/0.5 among Hera/Hermes/Artemis. If the goal is to observe attractor
|
||
behavior — or, just as importantly, to *rule it out* — the system has to
|
||
expose per-VM heat as a first-class, always-available fact, not something
|
||
inferred indirectly or reconstructed after the fact from birth/kill/touch
|
||
parity logs.
|
||
|
||
**Concretely, "transparent" means, at minimum:**
|
||
1. Any live VM's current `execution_heat_q48` is readable by ID or by name,
|
||
at any time, without needing to halt or inspect via debugger.
|
||
2. The fleet-wide state (`fleet_transfer_slope_q48`, fit quality, window
|
||
warmth) remains visible as it is today — this doc adds to that, doesn't
|
||
replace it.
|
||
3. A time-series of the above, not just an instantaneous snapshot, since
|
||
"attractor" is a claim about behavior *over ticks*, not a single reading.
|
||
4. The instrumentation itself must be a passive observer, per the existing
|
||
VM-fleet-physics design principle — recording and reporting only,
|
||
changing nothing about when or how VMs actually execute.
|
||
|
||
## Current instrumentation gap (concrete)
|
||
|
||
| What's needed | What exists today |
|
||
|---|---|
|
||
| Per-VM heat readable by ID/name | Not exposed. `VMPhysics.execution_heat_q48` lives inside the kmalloc-backed registry linked list, private to `capsule_vm_physics.c`. |
|
||
| Time-series export (per-tick, per-VM) | Nothing. The hosted VM's `--doe` CSV pipeline (`doe_metrics.c`, `heartbeat_export.c`) has no kernel-build equivalent — bare-metal experiments stream CSV-shaped rows to serial from FORTH capsules instead (see `experiments/bare_metal/`), and no such capsule exists yet for VM-fleet physics. |
|
||
| A controllable experimental factor | The mechanism is a passive observer by design — it never decides who runs. The *experiment* needs something else (a driving capsule) to vary dispatch patterns across Hera/Hermes/Artemis in a controlled way; no such capsule exists (the old `doe-campaign.4th`, which attempted something adjacent, is broken and being superseded by this doc — see below). |
|
||
| Analysis pipeline | The L8 campaign's `l8_analysis_safe.R`-style pipeline (CV, convergence time, phase-space portraits) is workload/word-shaped, not VM-fleet-shaped. Would need a new script, though the statistical *methodology* — CV over time, convergence-tick counting, ANOVA across factor levels — transfers directly. |
|
||
|
||
## Relationship to `doe-campaign.4th`
|
||
|
||
This design doc supersedes migrating `doe-campaign.4th` as originally scoped.
|
||
That capsule's model — manually setting VM heat via `VM-HEAT!` and manually
|
||
pumping ticks via `K-BUMP` — has no equivalent under the current mechanism,
|
||
which deliberately has no manual heat-injection point (heat only moves in
|
||
response to real `VM-EXEC`/`VM-CALL`/`VM-STEP` dispatch, by design, to keep
|
||
the physics a passive observer rather than a puppet). Reviving the old
|
||
capsule's *mechanics* would mean re-adding exactly the kind of synthetic
|
||
heat-injection surface the current design deliberately removed. The old
|
||
capsule's *goal* — drive the fleet through controlled scenarios and observe
|
||
the result — is exactly this experiment's goal, so its ideas carry forward
|
||
here; its code does not. `doe-campaign.4th` itself should be deleted once
|
||
this experiment has a working replacement, not before.
|
||
|
||
## Rev b: minimum step implemented, and a first real observation
|
||
|
||
Per open question #4, the smallest useful step: `vm_physics_heat_of(uint32_t
|
||
vm_id)` (`capsule_vm_physics.c`/`.h`) and an extension to the existing
|
||
`vm_physics_status()` that walks the live-VM list and prints
|
||
`vm_id=N name=X heat_q48=Q` for each, using `capsule_vm_registry_get()` for
|
||
the name. No new FORTH word needed — `VM-PHYSICS-STATUS` already existed and
|
||
now carries the per-VM breakdown automatically.
|
||
|
||
Verified by boot-injecting `VM-PHYSICS-STATUS` at the `ok>` prompt via the
|
||
serial socket (`socat - UNIX-CONNECT:<sock>`, same mechanism the Makefile's
|
||
`DOE_INJECT` uses for `EXEC-DOE`) right after a normal TRIPOD-TEST run. Output:
|
||
|
||
```
|
||
[Hera] VM-PHYSICS: fleet_heat_sum=65536
|
||
[Hera] VM-PHYSICS: conserved=CONSERVED
|
||
[Hera] VM-PHYSICS: fleet_transfer_slope_q48=0
|
||
[Hera] VM-PHYSICS: slope_fit_quality_q48=0
|
||
[Hera] VM-PHYSICS: fleet_window_warm=YES
|
||
[Hera] VM-PHYSICS: vm_id=3 name=Hermes heat_q48=0
|
||
[Hera] VM-PHYSICS: vm_id=1 name=Artemis heat_q48=0
|
||
[Hera] VM-PHYSICS: vm_id=0 name=Hera heat_q48=65536
|
||
```
|
||
|
||
**This is already a real data point, not just a smoke test.** At the end of
|
||
a normal boot (after `PASS: fleet K`, `PASS: Hermes liveness`, `PASS: Artemis
|
||
ready`, the Hermes kill/rebirth soak, and `PASS: E2E msg flow`), 100% of
|
||
fleet heat sits at Hera — the root — and both children read exactly zero.
|
||
That's one snapshot, not a trajectory, so it doesn't answer the research
|
||
question by itself, but it's suggestive: if this holds up as the fleet's
|
||
actual resting state rather than an artifact of a short boot sequence, the
|
||
"attractor" this mechanism finds may not look like the word-level one
|
||
(distributed heat settling into a stable per-word split) — it may look like
|
||
total heat drift back to the structural root every time activity quiets
|
||
down, which would itself be worth understanding (`vm_physics_retire`'s
|
||
parent-pointer walk to Hera on kill is one obvious contributor; whether
|
||
`vm_physics_touch`'s pull-toward-the-touched-VM mechanic ever produces a
|
||
*sustained* non-root distribution during activity, only draining afterward,
|
||
is exactly what a time-series would show and a single snapshot can't).
|
||
|
||
`fleet_window_warm=YES` but `fleet_transfer_slope_q48=0` in the same
|
||
snapshot is also worth noting for rev c: the touch-history window has
|
||
enough samples to be "warm," yet the inferred slope is still zero. Worth
|
||
checking whether that's expected (e.g. inference hasn't been triggered by a
|
||
`vm_physics_tick` call recently enough) or a sign the inference path needs
|
||
its own scrutiny before trusting slope-based conclusions later.
|
||
|
||
## Experiment roadmap (rev c)
|
||
|
||
Four phases, each a prerequisite for the next. Escalate only as far as
|
||
needed — do not build phase N+1 until phase N's result says it's necessary.
|
||
|
||
### Phase 1 — Readiness handshake (existing primitives only)
|
||
|
||
**Goal:** the simplest possible rendezvous, as a clean baseline before any
|
||
workload variation is introduced. No new broadcast machinery — this phase
|
||
deliberately uses only what already exists (`MSG-SEND`, `MSG-DELIVER-ALL`,
|
||
`HERMES-TICK`, `MSG-ACK-LAST`), per the finding above that real multi-member
|
||
delivery doesn't exist yet and shouldn't be built just for this.
|
||
|
||
**Mechanism:** after each VM's own `CD-INIT` completes, it sends a `READY`
|
||
message (new event-code constant, alongside the existing `SPAWN-EVENT`/
|
||
`PAUSE-EVENT`/`RESUME-EVENT`/`KILL-EVENT` in `hermes/init.4th` block 4100).
|
||
The receiving side (whichever VM hosts the message arena — needs
|
||
confirming against the actual capsule wiring when implementation starts,
|
||
not assumed here) drains the queue via repeated `HERMES-TICK`/
|
||
`MSG-DELIVER-ALL` passes and counts distinct `READY` senders. Once all
|
||
three are accounted for, it sends each an ACK back via the existing
|
||
`MSG-ACK-LAST` pattern (`common:msg.4th`'s `HERMES-ACK`).
|
||
|
||
**Observation:** snapshot `VM-PHYSICS-STATUS` immediately after all three
|
||
ACKs land. This is the cleanest possible baseline reading — no workload
|
||
variation, no DoE — and directly extends the rev b observation (which was
|
||
taken after a full TRIPOD-TEST run, not a clean handshake). Compare the two:
|
||
does heat still end up 100% at Hera after *just* a handshake, or was rev
|
||
b's observation an artifact of everything else TRIPOD-TEST does (the K
|
||
soak kill/rebirth in particular)?
|
||
|
||
**Explicitly out of scope for phase 1:** any workload, any DoE, any
|
||
multi-member broadcast. Point-to-point messages and a counter only.
|
||
|
||
### Phase 2 — Real broadcast delivery
|
||
|
||
**Goal:** implement actual multi-member channel delivery, consuming the
|
||
existing but currently-unused `CH-ADD-MBR`/`CH-MBRS`/`MBR-VM@` scaffolding
|
||
in `hermes/init.4th`. A message sent to a channel should reach every
|
||
member, not just one `to` VM.
|
||
|
||
**Why after phase 1, not before:** phase 1 proves the simpler point-to-point
|
||
handshake works and gives a clean baseline reading before adding new
|
||
delivery machinery to the mix. If phase 1's baseline is already confusing
|
||
or unexpected, that's a signal to understand *before* introducing a second
|
||
new mechanism on top of it.
|
||
|
||
**Not yet designed:** fan-out delivery semantics (does a channel message
|
||
get one arena slot copied N times, or N independent slots?), how
|
||
`MSG-DELIVER-ALL`'s scan loop changes to walk `CH-MBRS` per channel message
|
||
instead of a single `to` field, and whether `COMMON-CH`'s existing
|
||
heat-floor-only role expands or stays separate from its new delivery role.
|
||
This needs its own design pass when phase 1 is done and verified — not
|
||
specified further here.
|
||
|
||
### Phase 3 — DoE small: one config per VM per run
|
||
|
||
**Goal:** the actual attractor question, at the smallest scale that can
|
||
plausibly answer it.
|
||
|
||
**Factor:** `doe.4th`'s existing 2⁴ factorial (`CFG-ENT`/`CFG-CV`/`CFG-TMP`/
|
||
`CFG-STB`, `CURR-CFG` 0–15 via `APPLY-CFG`) — the "16 workloads." Confirmed
|
||
choice over the 9 separate `init-*.4th` capsule files, which are alternate
|
||
*Mama personalities* (mutually exclusive, one per boot) and structurally
|
||
wrong for "assign each of 3 co-existing VMs its own workload."
|
||
|
||
**Design:** each run assigns one of the 16 configs to each of Hera/Hermes/
|
||
Artemis independently (with or without replacement across the 3 — TBD when
|
||
implementation starts), shuffled/replicated across many runs — the same
|
||
scale and randomization-for-temporal-bias-elimination approach as the
|
||
180-run L8 campaign. **Not** the full 16³ cross product (see Phase 4).
|
||
Response variable: per-VM heat trajectory over the run (via the phase-1/
|
||
rev-b transparency primitives), watched for CV convergence, stabilization
|
||
time, anything resembling the word-level campaign's findings — exact
|
||
statistical treatment TBD, likely adapting `l8_analysis_safe.R`'s approach
|
||
(CV by group, ANOVA, phase-space portraits) rather than inventing new
|
||
methodology.
|
||
|
||
**Escalation criterion:** if Phase 3's results are inconclusive — no clear
|
||
convergence signal, too much noise to distinguish configurations, etc. —
|
||
proceed to Phase 4. If Phase 3 already shows something as clean as the
|
||
word-level campaign's findings, Phase 4 may not be needed at all.
|
||
|
||
### Phase 4 — DoE large: full 16×16×16 cross product (only if Phase 3 is inconclusive)
|
||
|
||
Every combination of (Hera's config, Hermes's config, Artemis's config),
|
||
4,096 distinct combinations. At 30 replicates each (the L8 campaign's
|
||
convention) that's ~123,000 runs — a real undertaking, not a first pass.
|
||
Explicitly gated on Phase 3 not being sufficient; not scheduled otherwise.
|
||
|
||
---
|
||
|
||
## Open questions (carried forward, narrowed by the roadmap above)
|
||
|
||
1. ~~**Driving mechanism.**~~ Resolved by the roadmap: Phase 1 uses existing
|
||
point-to-point messaging; Phase 2 adds real broadcast; Phases 3/4 drive
|
||
workload via `doe.4th`'s existing `APPLY-CFG`. Remaining detail: which
|
||
VM hosts the message arena / where `READY` gets received in Phase 1 —
|
||
deferred to implementation time, not blocking this doc.
|
||
2. **What counts as "found an attractor" for a VM fleet?** Still open.
|
||
Phase 1's clean-baseline reading and Phase 3's trajectory data are both
|
||
needed before this can be answered rather than guessed at.
|
||
3. **Where does this live?** Still open — `docs/working/experiments/campaigns/`
|
||
following the `l8_attractor_map` precedent is the likely answer, but no
|
||
campaign directory has been created yet since no phase has code yet.
|
||
4. ~~**Minimum instrumentation vs. full pipeline.**~~ Resolved in rev b:
|
||
smallest step first. The four-phase roadmap above is the same principle
|
||
applied one level up — each phase is itself the smallest step toward
|
||
the next.
|
||
|
||
---
|
||
|
||
## Rev d: Phase 1 implemented and verified (three-arch)
|
||
|
||
**Implementation.** Hermes gained `READY-EVENT`/`READY-COUNT`/`NOTE-READY`/
|
||
`READY-ALL?`/`ENQUEUE-READY`/`HERMES-ANNOUNCE-READY`/`READY-ACK` (block 4118,
|
||
previously empty). Artemis gained `ARTEMIS-ANNOUNCE-READY`/`ARTEMIS-READY-ACK`
|
||
(new block 4129, previously unassigned) — since Artemis has no message arena
|
||
of its own, its announcement is a real `MSG-SEND` enqueued via a single-level
|
||
`VM-EXEC` call into Hermes (`S" 2 ENQUEUE-READY" S" Hermes" VM-EXEC`), not a
|
||
nested command string. `init.4th` gained `READINESS-HANDSHAKE` (new block
|
||
2053, defined before its call site per file-order execution — same pattern
|
||
`BOOT-BANNER`'s block 2057 already uses) and now calls it right after
|
||
`BOOT-BANNER`, before `TRIPOD-TEST` runs.
|
||
|
||
**A sequencing constraint the doc's rev-c text didn't anticipate:** Artemis
|
||
is birthed before Hermes in the existing boot order, so it can't announce
|
||
its own readiness at the tail of its own `CD-INIT` — Hermes (the message
|
||
host) doesn't exist yet at that point. Resolved by having Hera trigger all
|
||
three announcements explicitly, once both children are alive, rather than
|
||
each VM self-announcing inline in its own `CD-INIT`.
|
||
|
||
**Verified three-arch, first try:** amd64/aarch64/riscv64 all show
|
||
`Hermes: ready-ack`, `Artemis: ready-ack`, `PASS: readiness handshake`, then
|
||
every existing TRIPOD-TEST gate (`fleet K`, `Hermes liveness`, `Artemis
|
||
ready`, `reap`, `K soak`, `E2E msg flow`) — zero `FAIL` lines, identical
|
||
`dict_hash` (`0xcc590996ecda7654`) across all three.
|
||
|
||
**The observation holds up, and gets cleaner.** The `VM-PHYSICS-STATUS`
|
||
snapshot taken immediately after the handshake — before `TRIPOD-TEST`, before
|
||
any K-soak kill/rebirth — shows the *same* pattern as rev b's post-boot
|
||
reading: 100% of fleet heat at Hera, both children at exactly zero.
|
||
`fleet_window_warm=NO` this time (vs. `YES` in rev b), consistent with this
|
||
being a much shorter, lower-activity sequence. This rules out the rev-b
|
||
hypothesis that the all-heat-at-root reading was an artifact of TRIPOD-TEST's
|
||
kill/rebirth soak — it isn't. Even the simplest possible sequence (birth two
|
||
children, exchange three point-to-point messages, done) settles heat
|
||
entirely at the structural root. Two independent readings now agree; this
|
||
looks like the mechanism's actual resting behavior, not noise.
|
||
|
||
Next: Phase 2 (real broadcast delivery), still not started.
|
||
|
||
## Rev e: Phase 2 implemented and verified (three-arch)
|
||
|
||
**Resolves the two "not yet designed" questions from rev c's Phase 2
|
||
section.** Fan-out semantics: **N independent message slots**, not one
|
||
slot copied/shared — `MSG-BROADCAST` is sugar that walks `CH-MBRS` at send
|
||
time and calls the existing single-recipient `MSG-SEND` once per member.
|
||
This means `MSG-DELIVER`/`MSG-DELIVER-ALL` needed **zero changes** — every
|
||
fanned-out message still has a normal single `to` field, so the entire
|
||
existing delivery/ack/type-state machinery works unmodified. `COMMON-CH`'s
|
||
existing heat-floor role stays untouched; `CH-ADD-MBR` only touches
|
||
`CH-MBRS`, never `CH-HEAT`.
|
||
|
||
**Implementation.** Hermes gains `MSG-BROADCAST ( type from paddr plen ch -- )`
|
||
(new block 4151 — walks the channel's member list via `CH-MBRS@`/`MBR-NEXT@`,
|
||
calling `MSG-SEND` once per `MBR-VM@`) and `REGISTER-COMMON-MEMBERS`/
|
||
`BCAST-RECV`/`SEND-BROADCAST-TEST` (new block 4152). Artemis and Hera each
|
||
get their own `BCAST-GOT`/`BCAST-RECV` pair (own block 4129 addition;
|
||
`init.4th` new block 2054) — same word names, independent per-VM state,
|
||
mirroring how `NOTE-READY`/`READY-COUNT` worked in Phase 1. `init.4th` gains
|
||
`BROADCAST-TEST`, called right after `READINESS-HANDSHAKE`: registers all
|
||
three VMs as `COMMON-CH` members, triggers one broadcast from Hermes, drains
|
||
via `HERMES-TICK`, then checks each VM's own `BCAST-GOT` counter via
|
||
`VM-CALL` to confirm genuine 3-way delivery (not just "a message was sent").
|
||
|
||
**Verified three-arch, first try:** amd64/aarch64/riscv64 all show
|
||
`PASS: broadcast reached all 3` immediately after the Phase 1 handshake,
|
||
then every existing TRIPOD-TEST gate — zero `FAIL`, identical `dict_hash`
|
||
(`0x23c8d067ec5709b6`) across all three, no crashes or hangs on any
|
||
architecture.
|
||
|
||
**Third independent reading, same result.** The `VM-PHYSICS-STATUS`
|
||
snapshot after the broadcast — now after *two* real message exchanges
|
||
(3 point-to-point readiness messages, then 3 more from the broadcast
|
||
fan-out; 6 total) — shows the identical pattern as revs b and d: 100% of
|
||
fleet heat at Hera, both children at exactly zero. Three independent
|
||
readings, three different activity levels, same resting state every time.
|
||
This is no longer "suggestive" — it's a reproducible finding: whatever
|
||
activity the fleet does, heat returns entirely to the structural root
|
||
between observations. Whether that's `vm_physics_touch`'s pull-toward-the-
|
||
touched-VM never producing a *sustained* distribution, or something else,
|
||
is squarely a question for Phase 3's trajectory data (a single post-hoc
|
||
snapshot, however many times repeated, still can't show what happens
|
||
*during* activity — only what's left after it settles).
|
||
|
||
Next: Phase 3 (DoE small — one of `doe.4th`'s 16 configs per VM per run),
|
||
not yet started.
|
||
|
||
## Rev f: the time-series gap was already closed, and it found a real bug
|
||
|
||
Rev a's instrumentation-gap table claimed no kernel-build time-series
|
||
export existed. **That was wrong.** `src/starkernel/doe_log.c` already
|
||
streams a genuine per-tick (100Hz), 15-column CSV row to serial on every
|
||
single boot — this is what has been producing the "N data rows" line at
|
||
the end of every acceptance run all session (`experiments/bare_metal/runs/
|
||
doe-*.csv`). It's word-level dictionary metrics only (hot word count, avg
|
||
word heat, window width, etc.) — no per-VM fleet heat. Extended it to 18
|
||
columns, adding `hera_heat_q48`/`hermes_heat_q48`/`artemis_heat_q48`,
|
||
looked up by name each tick via the existing `vm_physics_heat_of()`
|
||
accessor (robust to vm_id changes across kill/rebirth). Fixed to the three
|
||
known Tripod VMs — a logging-schema choice, not a change to the physics
|
||
mechanism itself, which stays VM-count-agnostic.
|
||
|
||
**This immediately surfaced why all three prior readings agreed: the
|
||
mechanism has never once activated.** All 1797 ticks across a full ~18s
|
||
boot (handshake, broadcast, TRIPOD-TEST, K-soak, E2E messaging — every
|
||
phase run so far) show the *exact same three values on every single tick*:
|
||
`hera=65536, hermes=0, artemis=0`. Not "returns to this" — never deviates.
|
||
|
||
**Root cause, traced to `capsule_vm_physics.c` (this session's own code):**
|
||
`fleet_transfer_slope_q48` is seeded at `0` (`capsule_vm_physics.c:82`).
|
||
`vm_physics_touch`'s transfer amount is `elapsed_ns * slope >> 16` — at
|
||
slope=0, every touch moves exactly zero heat, permanently. Meanwhile
|
||
`vm_physics_tick`'s regression is supposed to *infer* a better slope from
|
||
the touched VMs' heat trajectory, but those VMs' heat has been frozen at
|
||
zero by the very fact that slope=0 — and zero-heat samples are skipped
|
||
outright in the log-linear regression (`if (trajectory[i]==0) continue`).
|
||
Closed loop: zero slope → touched VMs never receive heat → regression sees
|
||
no signal to fit → slope stays zero. It cannot bootstrap itself out of the
|
||
starting condition. The word-level mechanism this was mirrored from avoids
|
||
exactly this by seeding `decay_slope_q48` at a nonzero default (2:1 =
|
||
131072, `include/vm.h:504`) — the one place this session's mirroring
|
||
wasn't faithful to the pattern it was copying.
|
||
|
||
**Every finding from rev b through rev e needs re-reading in this light.**
|
||
"100% of heat sits at Hera" was never evidence about where the fleet's
|
||
attractor is — it was evidence that the redistribution mechanism has been
|
||
inert since it was written. The `doe_log.c` extension itself is correct
|
||
and worth keeping regardless (it's what exposed this); the flat CSV data
|
||
it's been recording is the artifact to fix, not a finding to interpret.
|
||
Reported, not yet fixed — awaiting a decision on the nonzero seed value
|
||
before touching `capsule_vm_physics.c` again.
|
||
|
||
---
|
||
|
||
*Rev a posed the question and found the instrumentation gap — incorrectly,
|
||
as rev f later discovered. Rev b closed the per-VM-heat half of that gap
|
||
and got one real observation. Rev c documented the full four-phase
|
||
roadmap. Rev d implemented and verified Phase 1, getting a second reading.
|
||
Rev e implemented and verified Phase 2, getting a third. Rev f extended
|
||
the time-series export that already existed and discovered all three
|
||
readings agreed because the redistribution mechanism has never activated —
|
||
a seeding bug in this session's own `capsule_vm_physics.c`, not a finding
|
||
about attractors. Rev g fixed the seed and a second bug it exposed, then
|
||
found a third, deeper, pre-existing bug blocking both — not yet resolved.*
|
||
|
||
## Rev g: two real fixes applied, a third bug found underneath, not yet fixed
|
||
|
||
**Fix 1 (applied):** `fleet_transfer_slope_q48` reseeded from `0` to
|
||
`65536/3`, matching the word-level `decay_slope_q48`'s actual bootstrap
|
||
value (`src/vm_bootstrap.c`: `(1ULL << 16) / 3` — the `include/vm.h:504`
|
||
comment claiming "2:1 = 131072" is stale relative to the real code).
|
||
Confirmed via direct debug instrumentation that the seed takes effect
|
||
(`VM-PHYSICS-STATUS` reads `fleet_transfer_slope_q48=21845` throughout).
|
||
**On its own, insufficient** — heat still never moved.
|
||
|
||
**Fix 2 (applied):** traced the continued flatness to `vm_physics_touch`'s
|
||
three call sites (`mama_word_vm_exec`/`vm_call`/`vm_step` in
|
||
`mama_forth_words.c`) using `timer_now_ns()` — a low-level raw timer
|
||
function distinct from `vm_monotonic_ns(vm)`, the time source the rest of
|
||
the kernel's physics/heartbeat code actually uses successfully. Switched
|
||
all three to `vm_monotonic_ns(vm)`. Needed `#include "starkernel/vm/
|
||
vm_internal.h"` (qualified path required: an unrelated, unguarded-by-name
|
||
`src/vm_internal.h` for the *hosted* build exists too, and `-Isrc` searches
|
||
before `-I<KERNEL_SRC>/vm`, so a bare `#include "vm_internal.h"` silently
|
||
resolved to the wrong file and produced an implicit-declaration error with
|
||
no hint why).
|
||
|
||
**Bug 3 (found, not yet fixed — the actual remaining blocker):**
|
||
`vm_monotonic_ns(vm)` *also* returns 0, unconditionally, confirmed via
|
||
direct instrumentation showing `hz=0` on all 933,702 calls across a full
|
||
boot. Root cause is in `src/starkernel/hal/host_services.c`'s
|
||
`kernel_monotonic_ns()`: it checks `if (timer_tsc_hz() == 0) return 0`
|
||
with no further fallback. Under this QEMU/TCG environment, TSC frequency
|
||
calibration genuinely fails (confirmed in the boot log: `Timer: WARNING:
|
||
could not derive TSC frequency; PM Timer will be used for RELATIVE ns` /
|
||
`Timer: trust=1 (0=NONE,1=REL,2=ABS), TSC=0 Hz`) — and the timer subsystem
|
||
already has a correct, working fallback for exactly this case
|
||
(`timer_now_ns()`'s `calib_record.vm_mode && trust < TIMER_TRUST_ABSOLUTE`
|
||
branch, calling `timer_now_ns_vm_relative()`). `kernel_monotonic_ns()`
|
||
just never learned about it.
|
||
|
||
**This is pre-existing infrastructure, not introduced this session, and
|
||
it's bigger than VM-fleet physics.** `vm_tick_apply_background_decay(vm,
|
||
vm_monotonic_ns(vm))` — word-level heat decay, called every heartbeat
|
||
tick — goes through the identical broken function. If this environment's
|
||
TSC calibration failure is representative of real hardware this project
|
||
targets (not just this QEMU/TCG dev environment), word-level background
|
||
decay may have been silently inert under the same conditions as the
|
||
VM-fleet mechanism.
|
||
|
||
**Not yet resolved:** even `timer_now_ns()` itself returned 0 in the same
|
||
debug session, before the fix to use `vm_monotonic_ns` was applied — so
|
||
either `timer_now_ns_vm_relative()` has its own issue, or its trigger
|
||
condition wasn't being met the way expected. That second layer wasn't
|
||
traced further; scope check requested before continuing, since a
|
||
`host_services.c` fix has a much wider blast radius than the two capsule
|
||
files this whole investigation started in.
|
||
|
||
**Phase 3 stays blocked** — pending a decision on how far to take the
|
||
timer investigation, plus fixes 1-2 above still can't be verified to
|
||
actually move heat until fix 3 (or an equivalent) lands.
|
||
|
||
## Rev h: bugs 3, 4, and a fifth found underneath — all fixed, heat finally moves, three-arch verified
|
||
|
||
**Fix 3 (applied):** `kernel_monotonic_ns()` (`src/starkernel/hal/
|
||
host_services.c`) reimplemented TSC→ns conversion directly and gave up
|
||
(returned 0) whenever `timer_tsc_hz()` was 0 — which it reliably is under
|
||
QEMU/TCG. Replaced the whole body with a direct delegation to
|
||
`timer_now_ns()`, which already has the correct `TIMER_TRUST_RELATIVE`
|
||
fallback via the ACPI PM Timer (`timer_now_ns_vm_relative()`) that this
|
||
function never used.
|
||
|
||
**Bug 4 (found and fixed):** `vm_tick_inference_engine()` (word-level) and
|
||
`vm_physics_tick()` (VM-fleet) shared the same `last_inference_tick`
|
||
counter in `HeartbeatState`. The word-level call runs first in `vm_tick()`
|
||
and resets that counter every `HEARTBEAT_INFERENCE_FREQUENCY` ticks,
|
||
so the VM-fleet gate immediately after it never saw its own threshold
|
||
satisfied — starved permanently. Added a dedicated
|
||
`last_fleet_inference_tick` field to `HeartbeatState` (`include/vm.h`),
|
||
initialized in both `src/starkernel/vm/vm_bootstrap.c` and the hosted
|
||
`src/vm_bootstrap.c` (shared header). Checked `test_contracts.c` first —
|
||
its axiom snapshot only captures `tick_target_ns`, not the whole struct,
|
||
so the new field carries no contract risk.
|
||
|
||
**Bug 5 (found and fixed — the actual remaining blocker):** even with
|
||
fixes 3 and 4 in place, heat *still* never moved. Traced end to end with
|
||
temporary instrumentation (added, used, fully removed — confirmed via
|
||
`git diff --stat` showing zero net change beyond the three real fixes):
|
||
|
||
1. `vm_physics_touch()` received `now_ns=0` on every call.
|
||
2. `kernel_monotonic_ns()` → `timer_now_ns()` returned `0` on all ~930,000
|
||
calls across a full boot, not just early ones.
|
||
3. Inside `timer_now_ns()`: `trust=1` (RELATIVE) and `vm_mode=1` were
|
||
correctly set the whole time, so it *did* take the
|
||
`timer_now_ns_vm_relative()` branch.
|
||
4. Inside `timer_now_ns_vm_relative()`: `pmtimer_read()` returned
|
||
`0xFFFFFF` (masked all-ones) on *every* call, including the very first
|
||
one used to set `vm_pm_start` — so the tick delta was permanently 0.
|
||
|
||
An all-ones port read is the signature of an unmapped I/O port.
|
||
`PMTIMER_IO_PORT` was hardcoded to `0x408` — the legacy PIIX4/i440fx ACPI
|
||
PM Timer address — but `Makefile.starkernel` boots amd64 with `-machine
|
||
q35,accel=tcg`, and Q35's ICH9 LPC bridge puts the ACPI PM Timer at
|
||
`0x608`. The kernel had been reading a dead port the entire time; this
|
||
had nothing to do with fixes 3/4 and would have defeated any timer fix
|
||
built on top of it.
|
||
|
||
**The real fix, not a machine-specific patch:** added
|
||
`fadt_find_pm_tmr_port()` to `src/starkernel/arch/amd64/timer.c`, which
|
||
walks RSDP → XSDT → FADT ("FACP") — the same pattern already used by
|
||
`pci.c`'s MCFG lookup, independently duplicated here since the kernel has
|
||
no shared ACPI header and touching already-working PCI enumeration code
|
||
wasn't worth the risk. Prefers the ACPI 2.0+ `X_PM_TMR_BLK` Generic
|
||
Address Structure (offset 208) when it names a valid SystemIO address,
|
||
falls back to the legacy 32-bit `PM_TMR_BLK` field (offset 76), and only
|
||
falls back further to the old hardcoded `0x408` if FADT parsing fails
|
||
entirely. `pmtimer_port` is now a runtime variable, discovered once in
|
||
`timer_init()` from `boot_info->acpi_table`. This is machine-type-agnostic
|
||
by construction — it would have found `0x608` on Q35 or `0x408` on
|
||
i440fx without needing to know which one it's running on.
|
||
|
||
**A sixth bug found underneath *that*:** the FADT lookup initially failed
|
||
outright (`Timer: WARNING: FADT PM_TMR_BLK lookup failed`). Root cause:
|
||
`src/starkernel/boot/uefi_loader.c`'s ACPI table discovery loop matched
|
||
`EFI_ACPI_20_TABLE_GUID` *or* `EFI_ACPI_TABLE_GUID` and broke on
|
||
whichever came first in the firmware's configuration table — silently
|
||
handing back the legacy ACPI 1.0 RSDP (revision 0, no XSDT) even though
|
||
OVMF also publishes an ACPI 2.0 one. Any XSDT-based table lookup
|
||
(MCFG *or* FADT) fails against that pointer. Fixed with two explicit
|
||
passes: ACPI 2.0 first, ACPI 1.0 only as a fallback if 2.0 isn't found.
|
||
This is shared boot code (all three architectures), so it was subject to
|
||
full three-arch acceptance, not just amd64.
|
||
|
||
**Verification — all three architectures, clean acceptance runs:**
|
||
|
||
- amd64: `Timer: PM_TMR_BLK discovered from FADT at port 1544` (0x608,
|
||
exactly the expected Q35 address). TSC calibration also started
|
||
succeeding as a side effect (`TSC=2305485654 Hz`, previously always 0)
|
||
— `calibrate_tsc_with_pmtimer()` depends on the same working PM Timer
|
||
read.
|
||
- All three (amd64/aarch64/riscv64): identical `dict_hash=
|
||
0x23c8d067ec5709b6`, `PASS: E2E msg flow`, 0 `FAIL` lines, no panics/
|
||
faults, virtio-blk/Artemis attach unaffected by the `uefi_loader.c`
|
||
change.
|
||
- **Heat now actually moves.** DOE CSV `(hera_heat_q48, hermes_heat_q48,
|
||
artemis_heat_q48)` triples show three distinct states across all three
|
||
architectures — `(65536,0,0)`, `(0,65536,0)`, `(0,0,65536)` — tracking
|
||
which VM was touched most recently, with the conservation invariant
|
||
intact (sum always exactly `65536` = `Q48_ONE`). This was a flat,
|
||
unmoving trace on every single run before this fix chain.
|
||
|
||
**Phase 3 is now unblocked.** The observation mechanism this whole
|
||
investigation was blocked on — being able to see real heat trajectories
|
||
move between VMs — is confirmed working on all three architectures.
|
||
|
||
## Rev i: Phase 3 implemented and run — three-arch, 180 runs each
|
||
|
||
**Design, as implemented:** `EXEC-FLEET-DOE ( seed n-runs -- )`, added to
|
||
Hera's `init.4th`. Each run draws `(hera_cfg, hermes_cfg, artemis_cfg)`
|
||
independently and uniformly from 0–15 (with replacement — consistent
|
||
with Phase 4's full 16³ model, not a permutation of distinct values),
|
||
applies each, drives one cheap touch per remote VM to exercise
|
||
`vm_physics_touch`, and emits a `FLEETDOE,run_id,hera_cfg,hermes_cfg,
|
||
artemis_cfg` marker row. The already-working `doe_log.c` per-tick heat
|
||
CSV is the response variable; no new instrumentation was needed for
|
||
that half. Not run at boot — invoked post-boot via the same serial-socat
|
||
injection pattern the existing `DOE_INJECT`/`EXEC-DOE` mechanism uses,
|
||
with `12345 180 EXEC-FLEET-DOE` as the command.
|
||
|
||
**Three more capsule/runtime bugs found and worked around, not fixed at
|
||
the root (each would require touching C subsystems well beyond this
|
||
investigation's scope, per the standing rule to report and confirm
|
||
before that kind of change):**
|
||
|
||
1. **Cross-capsule `EXEC`/`USE` doesn't reliably resolve words when
|
||
triggered by a remote `VM-EXEC` call after boot.** Confirmed
|
||
pre-existing, not introduced here: `common:msg.4th`'s `USE` already
|
||
silently fails inside Hermes's own `CD-INIT` in every prior accepted
|
||
run (`[CAPSULE][DEFER] ... UNKNOWN WORD: 'USE'`) — it just never
|
||
mattered because nothing calls `HERMES-ACK`/`HERMES-NACK`. Worked
|
||
around: Hermes and Artemis's config application bypasses `APPLY-CFG`
|
||
entirely. Hera computes the four Q48.16 factor values locally and
|
||
sends them as literals straight to `L8-UPDATE`/`L8-APPLY` (both C
|
||
primitives, present in every VM) via a `FDOE-CFG-CMD-LO`/`-HI`
|
||
lookup-table dispatcher (split across two words/blocks to fit
|
||
`mkcapsule`'s 16-content-line-per-block limit at this string length).
|
||
|
||
2. **`VARIABLE`/`CREATE` storage allocated via nested `EXEC` of a
|
||
separately-named capsule fails the VM's own `vm_addr_ok()` bounds
|
||
check on later `!`/`@`.** Found by elimination while debugging bug 1's
|
||
workaround: a shared `common:doe-cfg.4th` capsule (mirroring
|
||
`common:msg.4th`'s pattern) was tried for Hera's own `APPLY-CFG`.
|
||
`0 APPLY-CFG` typed directly at Hera's live REPL failed with a bare
|
||
`ERROR` (no message — `memory_word_store`/`memory_word_fetch` set
|
||
`vm->error=1` silently on a failed bounds check). Isolated via direct
|
||
REPL testing: `0 CFG-ENT .` and `0 0 0 0 L8-UPDATE L8-APPLY` both
|
||
worked fine standalone; `VARIABLE FOO 5 FOO ! FOO @ .` (a brand-new
|
||
variable declared *interactively*, no capsule involved) also worked
|
||
fine. Only `CURR-CFG`, defined via the nested `EXEC`/`USE` of
|
||
`common:doe-cfg.4th`, failed on store/fetch — a distinct bug from #1,
|
||
specific to `VARIABLE`/`CREATE` addressing through that load path.
|
||
Fixed by deleting the shared capsule and inlining `APPLY-CFG` (plus
|
||
its `CURR-CFG` variable) directly into Hera's own `init.4th`, where
|
||
the identical pattern works.
|
||
|
||
3. **`ART-TICK` scans all 22,998 Artemis data blocks per call** — far
|
||
too expensive as a per-run "touch" once multiplied by 180 runs under
|
||
QEMU/TCG with full execution logging; the first real attempt stalled
|
||
rather than erroring. Replaced with a new one-line no-op `ART-PING`
|
||
word added directly to Artemis's own capsule.
|
||
|
||
**A fourth finding, not a capsule bug:** the kernel's `--log-level` boot
|
||
flag is parsed (`cmdline.c`) but never wired to `log_set_level()` —
|
||
`sk_vm_bootstrap.c` hardcodes `log_set_level(LOG_TEST)` unconditionally
|
||
for Hera, so the flag has no effect on the kernel target (it does work
|
||
on the hosted build). This made the ECW/INFO per-word execution trace —
|
||
appropriate for boot-time POST visibility — impossible to suppress for
|
||
a long post-boot campaign, and a first 20-run smoke test only completed
|
||
6 runs in 600 seconds as a result. Rather than touch `sk_vm_bootstrap.c`,
|
||
`EXEC-FLEET-DOE` now brackets its run loop with the already-existing
|
||
`LOG-LEVEL!` FORTH word (`1 LOG-LEVEL!` / `3 LOG-LEVEL!`, i.e. WARN
|
||
during the loop, back to TEST after) — a capsule-only fix with zero data
|
||
impact, since `FLEETDOE` rows and the `doe_log.c` CSV are raw
|
||
`console_puts` output, not gated by log level. Cut the same 20-run smoke
|
||
test to well under two minutes.
|
||
|
||
**Verification — three-arch, 180 runs each, single boot per
|
||
architecture, seed 12345:**
|
||
|
||
| Arch | FDOE errors | `doe_log.c` rows | FLEETDOE rows extracted |
|
||
|------|-------------|------------------|--------------------------|
|
||
| amd64 | 0 | 5781 | 124 / 180 |
|
||
| aarch64 | 0 | 5781 | 124 / 180 |
|
||
| riscv64 | 0 | 5781 | 124 / 180 |
|
||
|
||
Identical row counts across all three architectures, consistent with
|
||
deterministic same-seed behavior. Config draws land roughly uniformly
|
||
across 0–15 (5–12 draws each per factor, out of 180 — expected mean
|
||
11.25, normal sampling variance).
|
||
|
||
**The missing 56/180 FLEETDOE rows per run, every time:** the
|
||
interrupt-driven 100Hz heartbeat tick's `doe_log.c` CSV emission can
|
||
preempt Hera's foreground `FLEET-DOE-ROW` print mid-string on the shared
|
||
serial line (no locking between the two console-write paths), splicing
|
||
an embedded `[HADES][DOE ] ...` fragment into the run marker. The
|
||
extraction script (`extract_fleetdoe.py`) detects any `FLEETDOE,` line
|
||
that doesn't split into exactly 5 clean fields and drops it rather than
|
||
guessing — a known, logged, non-silent loss, not a fix to the underlying
|
||
console race (which would mean adding locking to an interrupt-context
|
||
console path, out of scope here).
|
||
|
||
**Correction to an earlier claim in this rev:** this section originally
|
||
reported "5767 of 5781 distinct heat-triple states" as evidence of rich,
|
||
continuous trajectory data. That number was wrong — a measurement
|
||
error, not a finding. Rechecked directly against
|
||
`experiments/bare_metal/runs/doe-amd64-20260705-223649.csv`: there are
|
||
exactly **5** distinct `(hera_heat_q48, hermes_heat_q48,
|
||
artemis_heat_q48)` states in the entire 5781-row trace — `(65536,0,0)`,
|
||
`(0,65536,0)`, `(0,0,65536)`, and two single-row transient states
|
||
(`(59317,6219,0)` and `(63042,2494,0)`) caught mid-transfer. The
|
||
mechanism is overwhelmingly winner-take-all, exactly as the transfer
|
||
formula (`amount = elapsed_ns * slope >> 16`, clamped to
|
||
`others_total`) predicts for any touch separated by a realistic time
|
||
gap. The likely source of the earlier wrong number: counting distinct
|
||
full CSV rows (near-unique by construction, since `tick_number`/
|
||
`elapsed_ns` differ every row) rather than distinct heat-triples
|
||
specifically — a one-off shell-pipeline mistake, not a code or data
|
||
bug. Recorded here rather than silently fixed, per this document's own
|
||
transparency-first design goal.
|
||
|
||
## Rev j: statistical analysis of the 180-run campaign — a clean null result
|
||
|
||
**Response-variable reconstruction.** `FLEET-DOE-ROW` and
|
||
`doe_log.c`'s per-tick heartbeat CSV emission share one
|
||
un-synchronized serial line — the interrupt-driven 100Hz tick can
|
||
preempt Hera's foreground print mid-string, which is also what corrupts
|
||
56/180 `FLEETDOE` rows per run (Rev i). Since no explicit join key was
|
||
recorded at collection time, a join was reconstructed after the fact by
|
||
walking each redacted log in original line order and tagging every
|
||
per-tick heat sample with whichever `FLEETDOE` config triple was most
|
||
recently seen before it. This is an approximation (a sample shortly
|
||
after a config change may still reflect the previous config), not a
|
||
precise per-run isolation, and it degrades further whenever several
|
||
consecutive `FLEETDOE` markers are corrupted (the samples in between
|
||
get attributed to the last *good* marker, silently spanning more than
|
||
one real run) — good enough for point-in-time transition analysis,
|
||
not for clean per-run window statistics. Both scripts
|
||
(`join_fleetdoe.py`, `analyze_transitions.py`) live in session
|
||
scratch, not committed to the repo; numbers below are reproducible from
|
||
the committed logs.
|
||
|
||
**Cross-architecture determinism: confirmed, cleanly.** Comparing the
|
||
124 runs with usable config markers across all three architectures:
|
||
**0 config-triple mismatches.** amd64/aarch64/riscv64 drew the identical
|
||
sequence of `(hera_cfg, hermes_cfg, artemis_cfg)` triples from the same
|
||
seed. This is the one unambiguous, load-bearing result of the campaign.
|
||
|
||
**Heat dynamics during the actual 180-run campaign: sparse, winner-take-
|
||
-all, no fractional states.** Counting state changes directly from the
|
||
log (not through the run-window join, so unaffected by marker
|
||
corruption): **318 total transitions** across the whole campaign, landing
|
||
on hera/hermes/artemis 140/98/80 times respectively — a real, uneven
|
||
distribution, but see below on causal weight.
|
||
|
||
**Does any config bit predict which VM a transition lands on?** For each
|
||
of the 12 binary factors (entropy/cv/temporal/stability × hera/hermes/
|
||
artemis), a permutation test (10,000 shuffles) compared the rate at
|
||
which transitions landed on VM *X* when *X*'s own factor bit was high
|
||
vs. low, using the 306/318 transitions with a known config at that log
|
||
position. **No factor reached significance** (all p > 0.3; most
|
||
differences were 0.00–0.06 in either direction, well inside permutation
|
||
noise for n≈150 per group). This is a clean null, not an ambiguous one.
|
||
|
||
**Why a null result is the architecturally expected outcome, not a
|
||
surprise:** tracing `APPLY-CFG`/`L8-UPDATE` → `ssm_l8_update()` →
|
||
`ssm_apply_mode()` shows the DoE config only ever writes into
|
||
`vm->ssm_config` (word-level physics tuning — window width, decay-slope
|
||
inference). `fleet_transfer_slope_q48`
|
||
(`src/starkernel/capsule/capsule_vm_physics.c:94`) — the only knob that
|
||
controls how much heat moves per touch — is updated exclusively by
|
||
`vm_physics_tick()`'s own touch-history regression
|
||
(`capsule_vm_physics.c:347`), with no code path connecting it to L8
|
||
state at all. **There is no causal channel, direct or indirect through
|
||
the shared-slope mechanism, from the DoE's manipulated factor to the
|
||
fleet-heat transfer rate.** Any effect the campaign could possibly have
|
||
detected would have to be mediated through incidental timing side
|
||
effects (does a different L8 window-width config change how long
|
||
`HERMES-TICK`/`ART-PING` takes, shifting `elapsed_ns` between touches
|
||
enough to matter) — a confound, not the intended manipulation, and one
|
||
weak enough that 318 transitions found no trace of it.
|
||
|
||
**What this means for the escalation criterion.** The design doc's
|
||
Phase 3 → Phase 4 escalation criterion was "inconclusive → escalate to
|
||
the full 16³ campaign." This result is not inconclusive — it's a clean
|
||
null with an identified structural cause. Running Phase 4 (≈123,000
|
||
runs) against the same architecture would almost certainly reproduce
|
||
the same null at a hundredfold cost, because the missing causal channel
|
||
doesn't appear at a larger sample size. **Escalating to Phase 4 is not
|
||
recommended without first giving L8 config an actual causal pathway
|
||
into `fleet_transfer_slope_q48`** (e.g., letting `ssm_apply_mode()` or
|
||
an equivalent feed a per-VM or fleet-wide slope adjustment) — a real
|
||
design decision, not a bug fix, and squarely a call for Bob rather than
|
||
something to implement speculatively here.
|
||
|
||
## Rev k: the real defect — the fleet's self-tuning loop is starved, not just disconnected from config
|
||
|
||
Bob's instinct after seeing the rev j numbers ("looks like a coin toss
|
||
or not enough data... probably a miswired feedback loop") was right,
|
||
and more precise than rev j's framing. Rev j established that config
|
||
has no path *into* the fleet mechanism. This rev establishes something
|
||
underneath that: **the fleet mechanism's own self-tuning loop —
|
||
`vm_physics_tick()`'s regression that's supposed to periodically re-fit
|
||
`fleet_transfer_slope_q48` from the fleet's own touch history — never
|
||
successfully runs, for any campaign of realistic size, regardless of
|
||
config.** Confirmed by direct instrumentation (added, verified, fully
|
||
removed — `git diff --stat` on both touched files showed zero residual
|
||
change before the finding was committed):
|
||
|
||
**Instrumenting `vm_physics_tick()` directly: zero calls, ever.**
|
||
Across an entire 20-run smoke test — full boot, POST, readiness
|
||
handshake, broadcast test, `TRIPOD-TEST`, and the campaign itself — a
|
||
call counter placed at the top of `vm_physics_tick()` never
|
||
incremented once. Not "called but bailing on `traj_len < 2`" — never
|
||
entered at all.
|
||
|
||
**Tracing the call site (`vm_tick()`, `vm_runtime.c:137`) to why:** the
|
||
fleet-inference gate is `(VM*)sk_get_mama_vm() == vm && (tick_count -
|
||
last_fleet_inference_tick) >= HEARTBEAT_INFERENCE_FREQUENCY`.
|
||
`HEARTBEAT_INFERENCE_FREQUENCY` is `1000` (`starforth_config.h:72`) —
|
||
1000 invocations of `vm_tick()`, which itself only runs once every
|
||
`HEARTBEAT_CHECK_FREQUENCY` (`256`) word executions **on that specific
|
||
VM instance**. Instrumenting `vm_tick()` itself, filtered to calls
|
||
where `vm` is Hera: across the full 20-run campaign, **Hera's own
|
||
`vm_tick()` fired a total of 61 times** (`tick_count` and the
|
||
gate's `diff` both at 61 by the end, `last_fleet_inference_tick` still
|
||
0 throughout — never once satisfied). A parallel unfiltered sample
|
||
showed Hermes and Artemis's own tick counts climbing far faster (233,
|
||
863, 1363 within the same window) — Hera's *own* word-execution rate is
|
||
the bottleneck, not the fleet's overall activity level.
|
||
|
||
**Why Hera's own rate is so low:** the gate was inherited unchanged
|
||
from the word-level inference engine, where it's well-calibrated —
|
||
Hera's dictionary handles millions of word executions during a real
|
||
boot's POST phase, so 256,000 words between inference attempts is
|
||
nothing. But Hera, in her role as fleet orchestrator, is mostly
|
||
*dispatching* — `VM-EXEC`/`VM-CALL` calls that block while the
|
||
*target* VM (Hermes/Artemis) does the actual work in *its own*
|
||
`heartbeat.check_counter`, which never counts toward Hera's. A DoE run
|
||
loop built from `RANDOM`, stack shuffling, a `CASE` dispatch, and a
|
||
handful of `VM-EXEC` calls executes on the order of a few dozen of
|
||
Hera's *own* words per run — nowhere near the volume the constant
|
||
assumes. At the observed rate (~3 of Hera's own ticks per run), the
|
||
180-run campaign accumulates roughly 540 — confirmed directly:
|
||
`VM-PHYSICS-STATUS` read before and after the full 180-run campaign (a
|
||
separate diagnostic run, not the debug-instrumented one) showed
|
||
`fleet_transfer_slope_q48=21845` and `slope_fit_quality_q48=0`
|
||
**identically, before and after** — the seed value, untouched, despite
|
||
`fleet_window_warm=YES` (satisfied almost immediately from a handful of
|
||
early boot-time touches) the entire time. Both gates exist; only one of
|
||
them is reachable at DoE-campaign scale.
|
||
|
||
**This means Phase 3's premise was compromised in a way rev j didn't
|
||
yet capture:** it's not just that config can't reach the fleet
|
||
mechanism — the fleet mechanism's own adaptive half never engages at
|
||
all during a targeted campaign like this one. `fleet_transfer_slope_q48`
|
||
behaved as a hardcoded constant (`65536/3`, the bootstrap seed) for the
|
||
entire 180-run run, on all three architectures. Every transition
|
||
observed in Phase 3's data was governed by that one frozen number, not
|
||
by anything resembling the "regression-inferred, self-tuning" behavior
|
||
the design doc specifies. This is a stronger explanation for the null
|
||
result than "no causal channel from config" alone — even a campaign
|
||
that *did* somehow influence timing enough to matter would be pushing
|
||
against a slope that never moves.
|
||
|
||
**Not fixed here.** This is a real defect in shared kernel code
|
||
(`vm_runtime.c`'s heartbeat gate), not a capsule-side workaround
|
||
target, and warrants its own decision on the right fix — a fleet-scoped
|
||
inference frequency independent of `HEARTBEAT_INFERENCE_FREQUENCY`
|
||
(most direct), gating on fleet touch count rather than Hera's own tick
|
||
count (matches the window's own warm-up signal, which already fires
|
||
promptly), or something else — flagged for Bob rather than implemented
|
||
speculatively.
|
||
|
||
## Rev l: rev k's gate fixed with a shared counter — and a second defect found underneath
|
||
|
||
Bob's diagnosis of rev k's null result ("miswired feedback loop") led
|
||
straight to the fix, and to a spec passage that had already ruled out
|
||
the broken version: `VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md`
|
||
(line 82) specifies that `fleet_transfer_slope_q48` must be inferred
|
||
"from aggregate fleet statistics... not one per VM" — the same
|
||
principle applies to the *readiness signal* that gates computing it,
|
||
not just the regression's own input trajectory. Gating on Hera's own
|
||
`heartbeat.tick_count` was exactly the "one per VM" framing the spec
|
||
already rules out for this parameter — not an missing case, a
|
||
mis-implementation of a spec that said the opposite.
|
||
|
||
**Fix, implemented and committed:** `vm_physics_heartbeat_tick()`
|
||
(`capsule_vm_physics.c`), a new fleet-wide tick counter distinct from
|
||
any VM's own `HeartbeatState`, called from every VM's own `vm_tick()`
|
||
(not gated to Hera specifically). Removes the now-dead
|
||
`last_fleet_inference_tick` per-VM field this replaces. Builds clean
|
||
(zero warnings on touched files) on the hosted build and all three
|
||
kernel architectures.
|
||
|
||
**Confirmed working, for what it fixes:** `vm_physics_tick()` now
|
||
fires and completes — `slope_fit_quality_q48` moved from `0` to the
|
||
fit-succeeded placeholder `52429`, confirmed via `VM-PHYSICS-STATUS`
|
||
before/after a 180-run campaign. The starvation defect is real and
|
||
gone.
|
||
|
||
**A second, distinct defect found immediately while verifying the
|
||
first:** the regression computed `fleet_transfer_slope_q48 = 0`.
|
||
`vm_physics_touch()` requires `slope > 0` to transfer any heat at all
|
||
— so the moment the now-functional loop fires and lands on exactly
|
||
zero, it freezes heat movement permanently, reproducing Bug 1's
|
||
original zero-slope deadlock (`rev g`) through a different door.
|
||
Confirmed directly: **zero heat-state transitions across an entire
|
||
180-run verification campaign** — one state, `(0,0,65536)`, from
|
||
injection to completion, with the slope already at 0 before the
|
||
campaign even started (from earlier boot-time touches).
|
||
|
||
Likely cause, not yet distinguished: the regression assumes smooth
|
||
exponential decay (`heat(t) = h0·e^(-slope·t)`), but the fleet's actual
|
||
dynamics — confirmed back in the correction to rev i — are
|
||
winner-take-all step functions: one VM at full heat, the others at
|
||
zero, snapping abruptly rather than decaying smoothly. Fitting a
|
||
smooth-decay model to a step-function trajectory could plausibly
|
||
produce an exact-zero fitted slope as a genuine (if unhelpful) result,
|
||
not a bug in the fitting code. Q48.16 integer-division truncation of a
|
||
small nonzero result to 0 hasn't been ruled out either — the two would
|
||
need different fixes (a model that matches the actual dynamics, vs. a
|
||
precision fix in the existing division). **Not chased further here** —
|
||
flagged as the next open item, same standing as rev k's original
|
||
finding: a real question for Bob, not something to guess at by
|
||
implementing a fix for either hypothesis speculatively.
|
||
|
||
**Net effect on Phase 3's data:** unchanged from rev k's conclusion.
|
||
The 180-run campaign already committed ran entirely under the frozen
|
||
seed value (`65536/3`), not this newly-discovered zero — the shared-
|
||
counter fix and this second defect were both found *after* that
|
||
campaign, during verification of the fix. Re-running Phase 3 now would
|
||
just freeze at a different, worse constant (`0` instead of `65536/3`)
|
||
until the model/precision question above is resolved.
|
||
|
||
## Rev m: rev l's zero-slope defect resolved — model mismatch, confirmed empirically, not truncation
|
||
|
||
Bob's question: "let's figure out which it is, a model mismatch or a
|
||
Q48.16 truncation. I'm thinking the former." Confirmed empirically —
|
||
it's model mismatch, and not a close call.
|
||
|
||
**Method:** temporary instrumentation (`TOUCHDBG` in
|
||
`vm_physics_touch()`, `SLOPEDBG` in `vm_physics_tick()`, printing
|
||
`traj_len`, distinct-value count, `sum_log_heat`, `sum_t_log_heat`,
|
||
the pre-`abs()` signed numerator, and `denominator`), a 20-run
|
||
amd64/QEMU verification boot, then full removal — confirmed via
|
||
`git diff --stat` showing zero residual change before this write-up.
|
||
|
||
**What the boot-time touches showed:** the seeded slope (`65536/3`)
|
||
does move heat during ordinary boot activity — `TOUCHDBG` recorded a
|
||
real, nonzero transfer (`elapsed_ns=3310198 amount=1103382`) — so the
|
||
transfer arithmetic itself is not the problem, and Q48.16 precision is
|
||
in no way starving the transfer step. But the fleet's own dynamics
|
||
(rev i's winner-take-all finding) drove that transfer to *completion*
|
||
before Phase 3 even started: by the first post-boot
|
||
`VM-PHYSICS-STATUS`, one VM already held the entire `65536` and the
|
||
others held `0`.
|
||
|
||
**What `SLOPEDBG` showed at the moment the regression fired:**
|
||
|
||
```
|
||
SLOPEDBG traj_len=64 distinct=2 n=64 sum_t=2016
|
||
sum_log_heat=0 sum_t_log_heat=0
|
||
numerator_signed=0 (pos-or-zero) denominator=1397760 slope=0
|
||
```
|
||
|
||
`sum_log_heat` and `sum_t_log_heat` are exact zero, not small values
|
||
that rounded to zero — and the reason is exact, not approximate:
|
||
`vm_physics_tick()` looks up each history sample's *current*
|
||
`execution_heat_q48`, and skips zero-heat samples
|
||
(`if (trajectory[i] == 0) continue`) before ever computing a log. With
|
||
heat fully concentrated in one VM, every *contributing* sample has the
|
||
identical value `Q48_ONE` (`65536`) — which is Q48.16 for `1.0`, and
|
||
`ln(1.0) = 0` exactly. So `sum_log_heat = 0` and `sum_t_log_heat = 0`
|
||
are not underflowed approximations of something small; they are the
|
||
sum of exact zeros. `numerator_signed` is then `0 - 0 = 0`, again
|
||
exactly, before any division ever happens. There is no division step
|
||
where a small nonzero value could have been truncated away — the
|
||
truncation hypothesis doesn't have anywhere left to hide once the
|
||
inputs to the division are already exactly zero.
|
||
|
||
This holds for **any** distribution the fleet ever reaches under the
|
||
current winner-take-all dynamics, not just the one observed: whichever
|
||
single VM holds all the heat holds exactly `Q48_ONE`, whose log is
|
||
always exactly zero, and every other live VM's zero-heat samples are
|
||
filtered out before contributing to the sums at all. The regression's
|
||
input is structurally incapable of being anything but a constant series
|
||
of exact zeros once the fleet has fully concentrated — which, per rev
|
||
i, it does quickly and by design (winner-take-all is the intended
|
||
touch semantics, not a bug).
|
||
|
||
**Conclusion: this is model mismatch, confirmed, not Q48.16
|
||
truncation.** The regression assumes smooth log-linear decay
|
||
(`heat(t) = h0·e^(-slope·t)`); the fleet's actual dynamics are a step
|
||
function between exactly `0` and exactly `Q48_ONE`. Fit that model to
|
||
that data and an exact algebraic zero is the *correct* output of the
|
||
math — infinite-precision floating point would produce the identical
|
||
zero, for the identical reason (`ln(1) = 0`, and OLS on a constant
|
||
series has zero slope by construction). No fix has been applied. Per
|
||
this session's practice, this is a finding to report and confirm, not
|
||
implement speculatively — the next question, not yet asked, is what
|
||
model *should* replace log-linear decay for a fleet whose real dynamics
|
||
are step functions, and that's Bob's call.
|
||
|
||
## Rev n: two real fixes applied ("continue the best path according to the architecture") — a third, distinct root cause found and left for Bob
|
||
|
||
Bob approved rev m's diagnosis as "the best resolution for the desired
|
||
architecture" and asked to continue. Two fixes landed; a third question
|
||
is still open.
|
||
|
||
**Fix 1 (implemented, committed `9806263c`):** `VMFleetWindow.touch_history`
|
||
renamed to `heat_history`, now storing `target->physics.execution_heat_q48`
|
||
at the moment of each touch (after any transfer has settled) instead of a
|
||
`vm_id` to re-resolve at fit time. This is the rev m fix itself — it makes
|
||
the trajectory an actual historical time series rather than N lookups of
|
||
"now." Correct and necessary, confirmed by rebuild across the touched file.
|
||
|
||
**Verification after fix 1 found a second, independent defect:** slope
|
||
still landed on exactly `0`. Root cause, confirmed by instrumentation
|
||
(added, checked, removed): `vm_physics_touch()` applied
|
||
`fleet_transfer_slope_q48` against raw `elapsed_ns`, but that slope's seed
|
||
value (`65536/3`) was explicitly chosen to mirror the word-level engine's
|
||
own bootstrap constant — and the word-level engine applies its slope
|
||
against elapsed **microseconds** (`physics_metadata.c:337`,
|
||
`elapsed_us = elapsed_ns / 1000`, commented there as a deliberate
|
||
overflow-avoidance choice). Using nanoseconds directly was a straight
|
||
1000x scale error, not a design choice.
|
||
|
||
**Fix 2 (implemented, committed `5784b5ac`):** convert to `elapsed_us`
|
||
before applying the slope, exactly mirroring `physics_metadata_apply_linear_decay`'s
|
||
own convention.
|
||
|
||
**Verification after fix 2 found a third, distinct defect — not yet
|
||
fixed:** `TOUCHDBG2` instrumentation (added, checked, removed) on a fresh
|
||
boot showed exactly why. The first two real touches after boot:
|
||
|
||
```
|
||
TOUCHDBG2 vm_id=2 elapsed_us=207021 amount=69005
|
||
TOUCHDBG2 vm_id=1 elapsed_us=361602 amount=120532
|
||
```
|
||
|
||
Both `amount` values already exceed the entire conserved heat pool
|
||
(`Q48_ONE = 65536`) — these are one-time boot-sequence pauses (capsule
|
||
loading / disk I/O), 207ms and 362ms respectively, and at the corrected
|
||
per-microsecond rate the full-drain threshold is only `65536*65536/21845 ≈
|
||
196,608 us ≈ 197ms`. So the very first real touch, before the fleet window
|
||
has a chance to accumulate any pre-transition samples, still fully drains
|
||
the pool in one step. Contrast with steady-state touches captured moments
|
||
later in the same boot, once the interpreter is in its normal dispatch
|
||
rhythm:
|
||
|
||
```
|
||
TOUCHDBG2 vm_id=0 elapsed_us=9367 amount=3122
|
||
TOUCHDBG2 vm_id=0 elapsed_us=9639 amount=3212
|
||
TOUCHDBG2 vm_id=0 elapsed_us=9194 amount=3064
|
||
```
|
||
|
||
These are genuinely gradual — roughly 5% of the pool per touch, exactly
|
||
the smooth-pull behavior the model was meant to produce. The unit fix is
|
||
real and working; it's just arriving too late relative to two anomalously
|
||
long one-time pauses baked into this specific boot sequence.
|
||
|
||
**Assessment, not yet acted on:** this reconfirms rev m's core finding
|
||
from a cleaner angle. The fleet's dynamics, once correctly scaled, *are*
|
||
capable of gradual, regressable transfer — but only during steady
|
||
interpreter execution. Boot-sequence pauses (disk I/O, capsule loading)
|
||
are structurally one-time and long relative to the corrected time
|
||
constant, so they will keep pre-empting the gradual regime before the
|
||
window fills, for any fleet whose boot includes comparable pauses. Fixing
|
||
this further means either (a) a maximum per-touch transfer cap (bound
|
||
`amount` to some fraction of the pool regardless of elapsed time), (b)
|
||
excluding known one-time boot pauses from the elapsed-time calculation
|
||
(treating first-touch-after-birth specially), or (c) accepting that
|
||
winner-take-all-after-a-long-gap is correct fleet semantics and it's the
|
||
regression model, not the transfer mechanics, that still needs to change
|
||
(rev m's original conclusion, now with the transfer math confirmed
|
||
correct on its own terms). Not chased further here — three fixes deep in
|
||
one sitting is enough to stop and let Bob pick the direction rather than
|
||
keep guessing at numeric knobs.
|
||
|
||
## Rev o: option (c) chosen — log-linear regression replaced with direct rate recovery, three-arch verified
|
||
|
||
Bob's call on rev n's three-way fork: option (c) — winner-take-all after
|
||
a real dormancy gap is correct fleet semantics; a cooperative
|
||
single-threaded fleet where exactly one VM executes at any instant
|
||
naturally produces discrete handoffs, not continuous diffusion, so the
|
||
regression model was fighting the actual physics rather than modeling
|
||
it. Options (a) and (b) would have been numeric knobs tuned to force
|
||
step dynamics to look smooth to a model that was never going to fit
|
||
them — the same speculative-tuning trap rev l/n already fell into twice.
|
||
|
||
**New model:** the transfer law (`amount = elapsed_us * slope >> 16`,
|
||
clamped to available heat) is already known exactly, so there's no need
|
||
to curve-fit a reconstructed trajectory against a synthetic time axis at
|
||
all. Each touch now records its own transfer-law inputs —
|
||
`{elapsed_us, amount, clamped}` — instead of a derived heat value.
|
||
`vm_physics_tick()` inverts the law directly for every **unclamped**
|
||
sample (`rate = amount / elapsed_us`, exact, no log transform) and takes
|
||
the **median** across the window. Clamped samples (where the actual
|
||
transfer was capped by what the rest of the fleet held, including the
|
||
`others_total == 0` fully-concentrated case) carry no rate information
|
||
and are excluded outright, not zero-filled. Fewer than 8 informative
|
||
samples in the 64-deep window: skip the update, keep the current slope
|
||
— the same "skip, don't substitute a degenerate value" philosophy as
|
||
the existing `is_warm` gate, now also closing the zero-freeze trap at
|
||
its source instead of hoping the input data avoids it.
|
||
|
||
`slope_fit_quality_q48` is now `informative_count / window_depth` — a
|
||
real number, replacing the fixed `0.8` placeholder that had been in
|
||
place since the original design.
|
||
|
||
**`VMFleetWindow.heat_history` renamed to `touch_samples`**, now an
|
||
array of `VMFleetTouchSample` instead of `uint64_t`. `q48_log_approx` /
|
||
`q48_mul` / `q48_add` / `q48_from_u64` are no longer used in this file
|
||
(still used elsewhere); `Q48_ONE` remains needed and the include stays.
|
||
|
||
**Verified, amd64, 60-run DOE campaign** (instrumentation added, checked
|
||
removed before commit, same methodology as every fix this session):
|
||
|
||
```
|
||
seed: fleet_transfer_slope_q48=21845 quality=0 (cold)
|
||
after warm-up: fleet_transfer_slope_q48=21845 quality=0 (< 8 informative, held)
|
||
mid-campaign: fleet_transfer_slope_q48=21840 quality=25600 (~39% informative)
|
||
post-campaign (60): fleet_transfer_slope_q48=21829 quality=14336 (~22% informative)
|
||
```
|
||
|
||
No freeze, no runaway, no exact zero. The small drift (21845→21840→21829)
|
||
is expected and benign: since `amount` for any given touch was itself
|
||
computed from whatever slope was in effect at that moment, an unclamped
|
||
touch's recovered rate is close to self-consistent by construction —
|
||
this estimator's job is to stay stable and quality-aware under real
|
||
fleet conditions, not to discover some independent ground truth the
|
||
transfer law doesn't already encode. That's a feature, not a
|
||
limitation: the failure modes being fixed here were freezing and
|
||
fighting the physics, not "insufficiently novel inference."
|
||
|
||
**Three-arch acceptance, full protocol** (`make -f Makefile.starkernel
|
||
ARCH={amd64,aarch64,riscv64} clean qemu`, one at a time, foreground):
|
||
all three booted to `ok>` clean, **976 PASS / 0 FAIL** identically on
|
||
every architecture, and identical `dict_hash`
|
||
(`0xd5c86b7db55efa58`) across all three — zero-deviation determinism
|
||
holds. Logs: `logs/20260706-061201/amd64/`,
|
||
`logs/20260706-061320/aarch64/`, `logs/20260706-061455/riscv64/`.
|
||
|
||
This closes the zero-slope investigation that ran rev k through rev o.
|
||
The fleet's transfer mechanics, the sampling that feeds its
|
||
self-tuning loop, and the estimator itself are now all consistent with
|
||
each other and with the fleet's actual winner-take-all dynamics.
|
||
|
||
## Rev p: Phase 3 re-run under the fixed model — richer dynamics, same null result
|
||
|
||
Identical protocol to rev i/j — seed `12345`, `180` runs, one boot per
|
||
architecture, `EXEC-FLEET-DOE` — re-run under rev o's fixed estimator to
|
||
see whether the earlier null result was an artifact of the frozen slope
|
||
or a real property of the fleet.
|
||
|
||
**Config-triple determinism: still holds, unchanged.** All three
|
||
architectures' `FLEETDOE` marker files are byte-identical after sorting
|
||
(`diff` on amd64 vs. aarch64 and amd64 vs. riscv64: 0 lines). Expected —
|
||
config draws are seeded PRNG arithmetic, untouched by anything in rev
|
||
n/o.
|
||
|
||
**Heat dynamics are now genuinely rich — and now architecture-
|
||
dependent, which they weren't before.** Distinct `(hera, hermes,
|
||
artemis)` heat-triple states across the 5781-row trace:
|
||
|
||
| Arch | Rev i (frozen model) | Rev p (fixed model) |
|
||
|------|----------------------|----------------------|
|
||
| amd64 | 5 | 612 |
|
||
| aarch64 | 5 | 2276 |
|
||
| riscv64 | 5 | 2273 |
|
||
|
||
The jump from 5 to hundreds/thousands of distinct states directly
|
||
confirms rev o's fix: the fleet is now spending real, observable time in
|
||
fractional-transfer states between winner-take-all snaps, not just
|
||
teleporting between three fixed corners. The new *cross-architecture
|
||
divergence* (612 vs. ~2275) is an expected, previously-invisible
|
||
consequence of the same fix: the old step-function model crossed its
|
||
full-drain threshold on almost any realistic gap, so the exact
|
||
elapsed-time value never mattered and all three architectures produced
|
||
byte-identical winner-take-all trajectories under TCG. The new model's
|
||
fractional transfers are directly proportional to real elapsed
|
||
microseconds between touches, so different emulation speeds now
|
||
produce genuinely different heat trajectories — config-sequence
|
||
determinism (what gets drawn) is preserved; heat-trajectory numeric
|
||
determinism (what happens with it) is not, and arguably should not be,
|
||
now that real timing is load-bearing. This doesn't threaten `BIRTH`/
|
||
parity determinism: the three-arch acceptance run in rev o already
|
||
confirmed identical `dict_hash` across all three architectures — that
|
||
guarantee lives in word-level dictionary state, which this doesn't
|
||
touch.
|
||
|
||
**Transition counts, per architecture** (destination distribution):
|
||
|
||
| Arch | Total transitions | hera | hermes | artemis |
|
||
|------|--------------------|------|--------|---------|
|
||
| amd64 | 301 | 134 | 91 | 76 |
|
||
| aarch64 | 403 | 167 | 162 | 74 |
|
||
| riscv64 | 432 | 179 | 173 | 80 |
|
||
| *rev j (frozen, single log)* | *318* | *140* | *98* | *80* |
|
||
|
||
Higher absolute counts than rev j across the board (expected — the
|
||
fixed model produces far more state changes, fractional and full,
|
||
rather than a handful of clean snaps), but the **relative ordering
|
||
(hera > hermes > artemis) holds in all three new logs**, matching rev
|
||
j's ordering exactly. A structural property of the fleet (Hera as root
|
||
absorbs on retire, per `vm_physics_retire`) surviving a complete
|
||
replacement of the transfer/estimator machinery underneath it.
|
||
|
||
**Permutation tests: same null result, now on a fully-corrected model.**
|
||
Re-ran rev j's exact test — 12 factor/VM combinations (`ent`, `cv`,
|
||
`tmp`, `stb` × hera/hermes/artemis), Wilson 95% CIs, 50,000-shuffle
|
||
permutation test — independently on all three new logs (36 tests
|
||
total). **No factor reached significance on any architecture**: p-values
|
||
ranged from 0.14 to 1.00, lowest at `artemis`/`ent` (amd64 0.49,
|
||
aarch64 0.14, riscv64 0.20) and `artemis`/`cv` (amd64 0.22, aarch64
|
||
0.42, riscv64 0.52). This is the same clean null rev j found, but it
|
||
now carries more weight: rev j's null could have been explained by the
|
||
frozen slope suppressing any real signal (the mechanism literally
|
||
couldn't respond to anything). That explanation is no longer available
|
||
— the mechanism now visibly responds to real touch timing (612-2276
|
||
distinct states), and the null persists anyway. **Confirms rev j's
|
||
architectural diagnosis rather than superseding it**: `ssm_apply_mode()`
|
||
still only ever writes `vm->ssm_config`; nothing in rev n or rev o gave
|
||
L8 config a causal path into `fleet_transfer_slope_q48` or
|
||
`vm_physics_touch()`, and this re-run is direct evidence that no
|
||
incidental side-channel (e.g. config-dependent timing shifts) creates
|
||
one either.
|
||
|
||
**One directionally-consistent, non-significant pattern worth
|
||
flagging:** `artemis`'s destination rate trends *lower* under high
|
||
`ent` and `cv` config bits in **all three** independent architecture
|
||
replicates (`ent`: amd64 −0.042, aarch64 −0.058, riscv64 −0.053; `cv`:
|
||
amd64 −0.070, aarch64 −0.036, riscv64 −0.029). No individual test is
|
||
significant, and three consistent signs out of three is not strong
|
||
evidence on its own — but it's the kind of pattern that would be worth
|
||
a purpose-built follow-up (larger n specifically on this factor/VM
|
||
pair) if this ever becomes a priority, rather than the untargeted full
|
||
16³ Phase 4 sweep.
|
||
|
||
**Recommendation: unchanged from rev j, now on firmer ground.**
|
||
Escalating to Phase 4 (~123,000 runs) is still not recommended without
|
||
first giving L8 config an actual causal pathway into the fleet
|
||
mechanism — that was true when the null might have been a broken-model
|
||
artifact, and it's still true now that the model is fixed and the null
|
||
held anyway. The zero-slope investigation (rev k–o) is fully resolved;
|
||
whether to build that causal pathway is a new, separate design
|
||
decision for Bob, not a continuation of this one.
|
||
|
||
Data: `experiments/bare_metal/runs/{doe,fleetdoe}-{amd64,aarch64,
|
||
riscv64}-20260706-*.csv`. Logs: `logs/20260706-062132/amd64/`,
|
||
`logs/20260706-062357/aarch64/`, `logs/20260706-062528/riscv64/`.
|
||
|
||
## Rev q: L8 given a real causal channel into the fleet — the pathway rev j/p found missing, now built
|
||
|
||
Bob's question after rev p closed the zero-slope investigation: "shouldn't
|
||
L8 be wired in now?" — pointing at `doe_metrics.c`'s `2^7` loop-enable
|
||
space (L1-L7) and the Jacquard selector's whole purpose being to
|
||
translate observed statistics into live physics adjustment. Confirmed:
|
||
`ssm_jacquard.c` already runs a full 128-config adaptive UCB bandit
|
||
(`SsmConfigTable`), not the 16-mode legacy selector rev i/j's campaign
|
||
actually drove — but three separate gaps kept that from mattering. All
|
||
three closed, three-arch verified at each step.
|
||
|
||
**Gap 1 — L1/L4/L7 silently discarded.** The 128-config bandit always
|
||
selected across the full 7-bit space, but `ssm_apply_mode_from_table()`
|
||
and the inline apply in `ssm_l8_trial_end()` only ever wrote 4 of those
|
||
7 bits (L2/L3/L5/L6) into `ssm_config_t` — L1 (hotwords cache), L4
|
||
(pipelining), L7 (adaptive heartrate) were computed by the bandit and
|
||
then thrown away before reaching any real gate, governed instead by
|
||
whichever compile-time macro the binary was built with. Fixed:
|
||
`ssm_config_t` extended to all 7 fields, both write-sites updated, real
|
||
runtime checks wired at the two live call sites reachable from the hot
|
||
path (L4 in `physics_execution_hooks.c`'s pipelining block; L1 by
|
||
syncing the existing `hotwords_cache_set_enabled()` toggle once per L8
|
||
tick). L7 left deliberately propagate-only: `HEARTBEAT_THREAD_ENABLED`
|
||
governs background-pthread-vs-synchronous-inline dispatch, a
|
||
concurrency-model choice the tick machinery runs under either way, not
|
||
an on/off feature — live-toggling it would mean spawning/joining a real
|
||
OS thread from inside a per-tick physics callback, a different order of
|
||
risk than L1/L4's "skip a block of computation," and L7's own DoE
|
||
characterization ("ALWAYS ON, beneficial in 71% of top configs") gives
|
||
no indication toggling it off would ever be the right move. Commit
|
||
`61ac7f5e`.
|
||
|
||
**Gap 2 — the Fleet DOE drove the wrong mechanism.** `L8-UPDATE`/
|
||
`L8-APPLY` (what rev i/j's campaign actually called) drives the legacy
|
||
16-mode path — but the VM's own heartbeat always uses the 128-config
|
||
adaptive table instead (unconditionally allocated at boot), and
|
||
whichever fires more recently wins the shared `ssm_config_t`. Over a
|
||
multi-run campaign with plenty of word executions between draws, the
|
||
bandit's own periodic tick could silently overwrite a DOE-injected
|
||
config before it had any chance to matter — a second, previously-
|
||
invisible confound on top of rev j's "no wired path" finding. Fixed:
|
||
new `ssm_l8_force_config()` sets the bandit's own `current_config`
|
||
directly (an externally-forced trial looks identical to a self-selected
|
||
one, so the next trial-end scores it coherently and the bandit's normal
|
||
reward loop continues from there), exposed as `L8-TABLE-FORCE
|
||
( config_idx -- )`. `capsules/init.4th`'s whole `FDOE-CFG-CMD-LO/HI`
|
||
CASE-dispatch table (translating a 0-15 mode index into four synthetic
|
||
Q48.16 factor values) is gone — a raw 0-127 index needs no per-value
|
||
translation, just a "N L8-TABLE-FORCE" command string built via `PAD` +
|
||
`<# #S #>` + `MOVE` for remote `VM-EXEC`. `EXEC-FLEET-DOE` now draws
|
||
0-127 per VM. Net six capsule blocks removed, one added. Verified via
|
||
direct REPL injection (`VM-EXEC: '17 L8-TABLE-FORCE' -> 'Hermes'`, no
|
||
errors) and a 20-run smoke test drawing across the full range. Commit
|
||
`a6206acc`.
|
||
|
||
**Gap 3 — even correctly applied, L8 config had no path to the fleet
|
||
mechanism at all.** `fleet_transfer_slope_q48` is computed entirely by
|
||
`vm_physics_tick()`'s own rate-recovery estimator (rev o), with zero
|
||
code reading anything L8-related — closing gaps 1 and 2 makes L8 config
|
||
*apply* correctly, but still can't *reach* the fleet. Fixed: new
|
||
`l8_regime_modulation_q16()` reads Hera's own adaptive table's
|
||
`current_regime` (the same 3-bit entropy/cv/temporal classification the
|
||
bandit already stratifies its own scores by) and scales the
|
||
empirically-recovered median rate by `(popcount(regime)+1) * 0.5` —
|
||
0.5x-2.0x, richer signal environments track change more aggressively,
|
||
quieter ones damp toward the raw recovered rate. No per-VM favoritism:
|
||
this scales whatever the estimator already computed uniformly, never
|
||
decides which VM receives heat, so "physics is a passive observer,
|
||
never a driver" (`VM-PHYSICS-DYNAMIC-FLEET-DESIGN-20260705.md`'s
|
||
Goal) holds. Floor is 0.5x, never zero, so this can't reintroduce a
|
||
freeze on its own. New coupling, named explicitly: `capsule_vm_physics.c`
|
||
previously had zero dependency on `vm.h`/`ssm_jacquard.h`; it now reads
|
||
Hera's L8 state read-only via `sk_get_mama_vm()`. Unlike
|
||
`SsmConfigTable`'s DoE-seeded priors, this modulation formula has no
|
||
empirical calibration behind it — a principled default (bounded,
|
||
symmetric, reuses an existing signal), not a proven-optimal one.
|
||
Verified via a 150-run campaign (instrumentation added, checked,
|
||
removed before commit): a window with 44/64 informative samples moved
|
||
`fleet_transfer_slope_q48` from the seed (21845) to 2705, with heat
|
||
genuinely three-way fractional (11972/31294/22270) rather than
|
||
winner-take-all. Commit `5e40808a`.
|
||
|
||
**Net effect:** L8 config now has an unbroken, verified path from a
|
||
VM's own observed workload statistics through to the fleet's transfer
|
||
rate — the exact channel rev j's null result identified as missing.
|
||
Re-running Phase 3/4 against this would, for the first time, be testing
|
||
a mechanism actually capable of showing an effect. Not done here — a
|
||
new campaign, not a continuation of this fix.
|
||
|
||
Three-arch acceptance at every step (976 PASS / 0 FAIL, identical
|
||
`dict_hash` within each step) confirms determinism holds throughout.
|
||
|
||
## Rev r/s: one clock only — collapsing the "trial" batching unit, then discovering (and fixing) why that broke determinism
|
||
|
||
Bob, after rev q: "yes. let's show it works" — a demonstration campaign
|
||
against the new L8-into-fleet wiring. Before running it, instrumentation
|
||
revealed the rev q modulation had never actually varied: `current_regime`
|
||
requires `SsmConfigTable.trial_tick_count` to reach `SSM_MIN_TRIAL_TICKS`
|
||
(2000) before a trial ends and the regime updates, and no VM in a
|
||
realistic campaign — Hera ~70 ticks, Hermes ~294, Artemis frozen at
|
||
1620 across a 40-run campaign — ever got there. The earlier "150-run
|
||
campaign confirms modulation" claim (rev q) was corrected: it confirmed
|
||
the multiplication executes without crashing, not that it responds to
|
||
anything.
|
||
|
||
Bob's diagnosis, verbatim: "`current_regime` sounds like some kind of
|
||
tight C99 coupling to a given FORTH executing experiment to me" — the
|
||
"trial" was an artificial batching unit with no principled tie to
|
||
anything, invented to smooth a bandit's reward signal but in practice
|
||
just preventing it from ever firing during a VM's real boot lifetime.
|
||
Directive, verbatim: "ONE clock only, the heartbeat tick... a tick is a
|
||
tick and a tick is our resolution." No batching, no special-casing
|
||
between VM ticks and word ticks.
|
||
|
||
**Rev r.** `ssm_l8_trial_end()` → `ssm_l8_tick_score_and_select()`,
|
||
called unconditionally every tick from `ssm_l8_update_table()` instead
|
||
of gated behind `trial_tick_count >= SSM_MIN_TRIAL_TICKS`. ANOVA
|
||
stability changed from an accumulated fraction
|
||
(`anova_exits_this_trial / inference_runs_this_trial`) to a per-tick
|
||
signal (this tick's own outcome, neutral when inference didn't run this
|
||
tick). `SsmConfigEntry.trial_count` → `tick_count`,
|
||
`SsmConfigTable.regime_trials`/`total_regime_trials` →
|
||
`regime_ticks`/`total_regime_ticks`, `SSM_MIN_TRIAL_TICKS` removed
|
||
entirely. Every "trial" name and comment scrubbed from touched code per
|
||
explicit instruction ("anything that says 'trial' must be scrubbed from
|
||
the codebase").
|
||
|
||
Bob, mid-implementation: "here's the thing I'm concerned about. You're
|
||
making assumptions based on existing DoE designs. only the DoE concept
|
||
exists, the experiment does not yet exist" — a correction against
|
||
letting the *old* Fleet DOE protocol's assumptions (a config "holding"
|
||
for a run) constrain the clock fix. Get the clock right first; design
|
||
whatever experiment fits it after, not before.
|
||
|
||
**The regression.** Three-arch acceptance under rev r showed `dict_hash`
|
||
diverging for the first time all session — amd64 `0x64bb90c39bb94b1f`
|
||
vs. aarch64 `0x41c2b512a2dc415b`. Root-caused via temporary
|
||
instrumentation (added, used, fully removed before commit) to two
|
||
wall-clock-tainted paths, both derived from `vm->decay_slope_q48`
|
||
(Loop #6's adaptively-inferred decay slope, itself indirectly
|
||
wall-clock-dependent via Loop #3's real-nanosecond heat decay): (1)
|
||
`metrics.temporal_decay`, one of L8's three regime-classification
|
||
inputs, dormant at the old once-per-2000-ticks cadence, now live and
|
||
differently influencing which physics loops are enabled per
|
||
architecture every single tick; (2) the reward's window/slope
|
||
joint-convergence signal, same root cause. Bob: "must be fixed before
|
||
moving forward." A proposed stopgap (drop the temporal dimension from
|
||
regime classification) was rejected outright: "accurate code. fix it.
|
||
never allow a workaround."
|
||
|
||
**Rev s: the generalized compudynamics module.** Mid-fix, Bob flagged
|
||
the coming shape of the problem: "sooner or later, more sooner, we will
|
||
be adding messages and blocks along with VMs and words" — any fix
|
||
needed to generalize across all four entity levels, not just patch
|
||
words. After a short language-constraint misunderstanding was cleared
|
||
up (TRIPOD.md's/ARTEMIS.md's "StarForth dialect ONLY" applies to a thin
|
||
administrative/observational word surface; the underlying mechanics are
|
||
meant to be C99, same as everything built earlier this session — Bob's
|
||
own words: "I think the confusion is you took the FORTH/C interpretation
|
||
too literally... My bad"), Bob proposed the concrete shape: "why not
|
||
just a generalized compudynamics C module and appropriate tuning knobs
|
||
for each of the 4 operational levels?" — confirmed as design-for-all-
|
||
four, instantiate-what-exists-now ("two levels now, blocks and messages
|
||
later... design for all 4"), then final go-ahead: "yes, build it. that
|
||
sounds exactly perfect."
|
||
|
||
New `include/compudynamics.h` / `src/compudynamics.c`:
|
||
|
||
- `cd_classify_ids()` — deterministic diversity/volatility/locality
|
||
classifier over a recent entity-ID touch sequence (word IDs for word
|
||
level today), purely execution-count-derived, zero wall-clock input.
|
||
Replaces `temporal_decay` with a `locality_q16` signal fed from
|
||
`rolling_window_get_recent_sequence()` (word level's existing
|
||
execution-history structure — reused, not duplicated) instead of
|
||
`decay_slope_q48`. The reward's joint-convergence signal now compares
|
||
tick-to-tick `locality_q16` deltas instead of slope deltas.
|
||
- A generic UCB1 config-space bandit (`CDConfigTable`/`CDConfigEntry`,
|
||
`cd_config_table_init/seed/tick/force`) lifted out of `ssm_jacquard.c`
|
||
verbatim in logic, malloc-sized to `CDTuning.num_configs`/
|
||
`num_regimes` rather than hardcoded 128/8.
|
||
- `CDTuning` knob struct per level: `cd_tuning_word()` (128 configs,
|
||
8 regimes, the exact constants `ssm_jacquard.c` used before this
|
||
module existed — a lift, not a retune) is the only level actually
|
||
wired; `cd_tuning_vm()` documents classifier-side parameters for a
|
||
possible future migration of VM-level regime classification off
|
||
Hera's proxy (`capsule_vm_physics.c`'s `l8_regime_modulation_q16()`,
|
||
unchanged — it transitively inherits determinism from reading Hera's
|
||
now-fixed `current_regime`, needing no direct edit); block and
|
||
message levels are not defined at all, deliberately, rather than
|
||
filled with invented numbers.
|
||
|
||
`ssm_jacquard.c` now delegates: `ssm_l8_state_t.table` is a
|
||
`CDConfigTable*` instead of a hand-rolled `SsmConfigTable*`;
|
||
`ssm_l8_update_table()`'s signature changed from
|
||
`(..., uint32_t current_window, uint64_t current_slope)` to
|
||
`(..., uint32_t current_window, const uint32_t *recent_word_ids,
|
||
uint32_t recent_word_ids_count)`. The legacy 16-mode threshold path
|
||
(`ssm_l8_update()`/`ssm_apply_mode()`, dormant except on table-malloc
|
||
failure) is untouched — it still reads `metrics.temporal_decay` exactly
|
||
as before, deliberately, to keep this fix's blast radius to the path
|
||
that was actually live and actually broken.
|
||
|
||
**A second bug, found chasing the first.** Wiring `cd_tuning_word()`'s
|
||
address across translation units (`&CD_TUNING_WORD` as an addressable
|
||
`extern const CDTuning`) crashed the kernel on a page fault at a
|
||
constant, reproducible wrong address (`CR2=0x0000000800000034`, always
|
||
base-plus-field-offset from the same wrong base) the instant
|
||
`ssm_l8_init_table()` reached across into `compudynamics.c`. Traced via
|
||
direct fault-address arithmetic (matching `CR2` offsets against
|
||
`CDTuning`'s field layout) to this freestanding kernel's boot-time
|
||
loader not reliably relocating the address of an extern const aggregate
|
||
referenced across translation units — confirmed by `readelf -r` on the
|
||
*final linked* kernel ELF showing thousands of un-applied relocation
|
||
entries (`.rela.text`/`.rela.rodata`/`.rela.data`), meaning this build
|
||
is not a normal fully-resolved EXEC image; something at boot is meant
|
||
to walk `.rela.dyn` and apply these, and whatever that mechanism is,
|
||
this particular relocation type wasn't landing correctly for a plain
|
||
data symbol's address. Cross-TU *function calls* — including struct
|
||
return-by-value, which uses a caller-local hidden pointer, never a
|
||
global's address — worked correctly throughout (confirmed: `malloc()`
|
||
and `cd_config_table_init()` both executed correctly across the same
|
||
TU boundary). Fix: `CD_TUNING_WORD`/`CD_TUNING_VM` are not addressable
|
||
globals at all — `cd_tuning_word()`/`cd_tuning_vm()` return `CDTuning`
|
||
by value. This is the aggregate-typed version of a constraint every
|
||
scalar tunable elsewhere in this codebase already satisfied by being a
|
||
`#define` immediate rather than a referenced global; not a workaround,
|
||
an accurate fix for a real, verified constraint of this build target.
|
||
|
||
**A process failure, corrected mid-fix.** Recovering from a stray log
|
||
directory (`rm -rf logs/*/amd64`, glob far too broad — deleted
|
||
committed historical audit logs back to June 21) was itself botched:
|
||
`git checkout -- .` was run to restore the deleted files, but that
|
||
command reverts *every* modified tracked file to HEAD, not just deleted
|
||
ones — discarding this session's in-progress `ssm_jacquard.c`/`.h`,
|
||
`vm_time.c`, `vm_runtime.c`, and `capsule_vm_physics.c` changes,
|
||
including the already-verified relocation fix above. `compudynamics.h`/
|
||
`.c` were untracked and survived; the reverted files were reconstructed
|
||
exactly from conversation record and reverified with a clean
|
||
zero-warning hosted rebuild before re-attempting kernel acceptance.
|
||
Both mistakes reported to Bob in full as they were discovered, per
|
||
standing instruction.
|
||
|
||
**Three-arch acceptance.** Hera and Hermes `dict_hash` are now
|
||
byte-identical across amd64/aarch64/riscv64
|
||
(`0x83c2c109100e2ed6` / `0x4159dcb326d79759`) — the fix holds. Artemis's
|
||
`dict_hash` matches on amd64/aarch64 (`0xc8cff7d7710a91fc`) but diverges
|
||
on riscv64 (`0x41a705a6b637d492`) — new since this change (rev q's own
|
||
riscv64 log shows Artemis at `0xc31f4c8e7ca40740`, matching amd64 at the
|
||
same commit, so this is not a pre-existing, already-accepted gap).
|
||
|
||
Root-cause hypothesis, not yet confirmed by a live instrumented run:
|
||
Hera and Hermes do no real I/O during boot — their capsule code is pure
|
||
computation, identical word sequence every time regardless of real
|
||
elapsed time. Artemis's capsule (`capsules/artemis/init.4th`) does real
|
||
`virtio-blk-pci` disk I/O at boot (header read, format-write, two full
|
||
alloc→fetch→persist→free round-trips), and
|
||
`src/starkernel/virtio/virtio_blk.c`'s completion wait
|
||
(`while (s->used->idx == s->last_used_idx) { ...; if (!--spin) return
|
||
BLKIO_EIO; }`, lines ~279-284) is a genuine busy-wait whose iteration
|
||
count depends on real QEMU device-emulation and TCG-translation timing
|
||
per architecture — meaning Artemis's real boot duration is
|
||
architecture-variable. Because rev r makes L8 rescore and potentially
|
||
reselect on *every* heartbeat tick, and the heartbeat thread fires on a
|
||
real wall-clock cadence (`HEARTBEAT_TICK_NS`) rather than an
|
||
execution-count cadence, any VM whose real boot duration varies by
|
||
architecture gets a correspondingly variable *number* of heartbeat
|
||
ticks during boot — which can flip which config is active by the time
|
||
boot finishes, independent of whether the classifier itself is
|
||
deterministic given fixed input. Under rev q, no VM ever lived long
|
||
enough in ticks to reach a trial boundary, so this tension was
|
||
structurally inert; rev r's correctness (scoring/reselecting on the
|
||
clock that actually exists) is what exposes it. Hera/Hermes apparently
|
||
don't hit this in practice — their config trajectory happens to stay
|
||
stable regardless of a few ticks' difference — while Artemis's heavier,
|
||
I/O-bound boot does. The accurate fix, if this hypothesis holds, is a
|
||
heartbeat-architecture decision (tying ticks to execution count rather
|
||
than wall-clock time, at least for config-reselection purposes) — out
|
||
of scope for this fix and explicitly deferred: Bob's direction was
|
||
"commit what's verified, log this as a known gap" rather than open-
|
||
ended investigation or a narrow Artemis-side patch.
|
||
|
||
Hosted build: clean, zero warnings, `-Wall -Werror`, throughout every
|
||
step above. Commit `07f72874`.
|
||
|
||
## Rev t: the Artemis/riscv64 gap was Loop #3, not Artemis — decay converted to tick-based, gap closed
|
||
|
||
Rev s's closing hypothesis (I/O-driven boot-duration variance flipping
|
||
which L8 config is active) was Bob's cue for a sharper correction:
|
||
"the Artemis riscv64 gap isn't really a 'riscv64' gap. it's a
|
||
compudynamic bug it sounds like. EVERYTHING runs off the adaptive
|
||
heartbeat period. there's no exceptions or hidden dependency other than
|
||
the adaptive heartbeat on the monotonic wall clock timer." Tracing
|
||
precisely rather than re-asserting the rev s framing: `dict_hash`
|
||
(`capsule_dict_hash_hook()`) hashes each dictionary entry's name and
|
||
`execution_heat` — nothing else. `execution_heat` is decayed by Loop #3
|
||
(`physics_metadata_apply_linear_decay()`), and Loop #3's decay amount
|
||
has always — since before this session — been computed from real
|
||
elapsed nanoseconds (`sf_monotonic_ns()` / `vm_monotonic_ns()` deltas
|
||
against `DictPhysics.last_active_ns`/`last_decay_ns`), not from any
|
||
tick count. That's the actual hidden dependency: not new, not
|
||
riscv64-specific, just newly *live* because rev r made L8 responsive
|
||
enough to real conditions to expose it (at rev q, no VM ever lived long
|
||
enough in ticks to reach the old trial boundary, so which config was
|
||
active — and therefore whether Loop #3 was even gated on — never
|
||
changed mid-boot, on any architecture; the wall-clock-tainted decay
|
||
path existed the whole time but had no opportunity to diverge).
|
||
|
||
Bob's diagnostic question, verbatim: "We need some sense of relative
|
||
time to compute slope? is that it?" — correctly identifying that decay
|
||
inherently needs *some* notion of elapsed time to express a rate, and
|
||
asking whether that's the whole story. Answer, confirmed by reading
|
||
`vm_tick()` (`vm_runtime.c`): `vm->heartbeat.tick_count` already *is*
|
||
a purely execution-driven virtual clock — incremented synchronously
|
||
from the interpreter's own dispatch path, gated by word-execution count
|
||
(`HEARTBEAT_CHECK_FREQUENCY` in the non-threaded/kernel case), never by
|
||
a timer interrupt. Artemis's `virtio_blk.c` completion busy-wait (a
|
||
real, timing-variable spin loop; see rev s) is plain C — it never calls
|
||
`vm_tick()`, so it advances `tick_count` by exactly zero regardless of
|
||
how long it spins in real time. The correct clock for Loop #3 already
|
||
existed in the codebase; Loop #3 simply wasn't using it. Loop #6
|
||
(decay-slope inference, `inference_engine.c`) was checked and confirmed
|
||
to have no direct wall-clock reads at all — the bug was isolated to
|
||
Loop #3's decay *application*, not also present in the slope's
|
||
*calibration*.
|
||
|
||
**Fix**, once confirmed ("yes... yup"): added `DictPhysics.last_decay_tick`
|
||
(`include/vm.h`) alongside the existing `last_active_ns`/`last_decay_ns`
|
||
(kept, but now diagnostics-only — `physics_diagnostic_words.c` prints
|
||
them for humans, a legitimate use this fix doesn't touch).
|
||
`physics_metadata_apply_linear_decay()`'s signature changed from
|
||
`(entry, elapsed_ns, vm)` to `(entry, elapsed_ticks, vm)`; internally,
|
||
`decay_amount = (elapsed_ticks * slope_q48) >> 16` replaces
|
||
`(elapsed_ns/1000 * slope_q48) >> 16` — no conversion constant between
|
||
"per microsecond" and "per tick" introduced, since Loop #6 adaptively
|
||
recalibrates `decay_slope_q48` against whatever units it's actually
|
||
applied in. The `elapsed_ns < DECAY_MIN_INTERVAL` insignificant-interval
|
||
gate became `elapsed_ticks == 0` — semantically identical purpose,
|
||
tick-native expression. Decay turned out to be applied at six call
|
||
sites across four files, not the one background-batch site originally
|
||
scoped: `physics_execution_hooks.c`'s `physics_pre_execute()` (every
|
||
word execution) and `physics_on_lookup()` (every word lookup by name),
|
||
the kernel's `vm_core.c` inner-interpreter loop with its own inline
|
||
duplicate of the same two sites, and the periodic background-sweep
|
||
decay in `vm_time.c`/`vm_runtime.c`. All six converted identically:
|
||
`elapsed_ticks = vm->heartbeat.tick_count - entry->physics.last_decay_tick`,
|
||
then `entry->physics.last_decay_tick = vm->heartbeat.tick_count` after.
|
||
A real behavioral consequence, not just a units change: tightly-looping
|
||
words (multiple touches within the same heartbeat-tick window) now
|
||
correctly see zero decay between touches instead of an accumulating
|
||
sub-tick real-time trickle — a word that hasn't gone idle for even one
|
||
full tick hasn't earned any decay, which is closer to the intended
|
||
"hot word stays hot" model than the old scheme's continuous real-time
|
||
erosion.
|
||
|
||
**Three-arch acceptance, decisive.** All four VM identities now
|
||
byte-identical across amd64/aarch64/riscv64: Hera `0x83c2c109100e2ed6`,
|
||
Artemis `0x5284ea5cd0f9983c` (the previously-diverging value — riscv64
|
||
now matches amd64/aarch64 exactly, closing the rev s gap completely),
|
||
Hermes `0x4159dcb326d79759` (both the original birth and the
|
||
kill/respawn birth), Mama `0xc88c3c1db6ef601b`. Hosted build clean,
|
||
zero warnings, throughout.
|
||
|
||
## Rev u: Kconfig migration, Phase 2 — real symbol tree + bridge, wired but inert
|
||
|
||
Separate track from the rev r/s/t decay work above: StarForth's ~25 live
|
||
build-time knobs were scattered across two Makefiles with drifted, conflicting
|
||
fallback defaults (e.g. `ADAPTIVE_SHRINK_RATE` = 50 in the Makefile, 75 in
|
||
`rolling_window_knobs.h`'s own independent fallback — never in sync, never
|
||
discoverable). Bob asked for a real Linux-style `menuconfig`/`xconfig`/
|
||
`kconfig` system, explicitly choosing to vendor the genuine Linux
|
||
`scripts/kconfig` tooling over a lighter Kconfiglib-style alternative, and to
|
||
cover both the hosted `Makefile` and `Makefile.starkernel` from the start.
|
||
Phase 0 (five dead knobs removed) and Phase 1 (`tools/kconfig/` vendored,
|
||
building all three frontends — `conf`/`mconf`/`qconf` — as standalone host
|
||
tools) are covered by prior commits. This entry covers Phase 2: the real
|
||
Kconfig symbol tree and the Makefile-side bridge that consumes it.
|
||
|
||
**Symbol tree** (`Kconfig` + `Kconfig.arch`/`Kconfig.variant`/
|
||
`Kconfig.physics`/`Kconfig.heartbeat`, four `source`-included submenus):
|
||
`ARCH` as a three-way choice (amd64/aarch64/riscv64) with a derived
|
||
`ARCH_STRING`; `STARFORTH_VARIANT_HOSTED`/`STARFORTH_VARIANT_KERNEL` as the
|
||
top-level choice gating the hosted-only `TARGET=` profile choice and
|
||
platform-mode choice inside an `if STARFORTH_VARIANT_HOSTED` block (the
|
||
kernel build has no equivalent of either); the physics/SSM knob family
|
||
(`STRICT_PTR`, `ENABLE_HOTWORDS_CACHE`, `ENABLE_PIPELINING`,
|
||
`TRANSITION_WINDOW_SIZE` depending on `ENABLE_PIPELINING`, the `ADAPTIVE_*`
|
||
window-shrink family, `INITIAL_DECAY_SLOPE_Q48`, `DECAY_RATE_PER_US_Q16`,
|
||
`HEARTBEAT_INFERENCE_FREQUENCY`); and the heartbeat family
|
||
(`HEARTBEAT_THREAD_ENABLED`, `HEARTBEAT_TICK_NS` depending on it,
|
||
`HEARTBEAT_CHECK_FREQUENCY`, `HEARTBEAT_WINDOW_TUNING_FREQUENCY`,
|
||
`HEARTBEAT_SLOPE_VALIDATION_FREQUENCY`, `EMERGENCY_CONSOLE_ENABLED`). Every
|
||
default was hand-diffed against the actual current Makefile/header default
|
||
for that symbol, not copied from either Makefile's stale comments — this
|
||
caught and preserved the same `ADAPTIVE_SHRINK_RATE`/`ENABLE_HOTWORDS_CACHE`/
|
||
`ENABLE_PIPELINING` drift-vs-comment mismatches already known from Phase 0/1
|
||
scoping, documented in-line in `Kconfig.physics` rather than silently
|
||
"corrected" (Phase 3's job, not Phase 2's).
|
||
|
||
**One known, deliberately-preserved discrepancy**, documented directly in
|
||
`Kconfig.heartbeat`'s help text: `Makefile.starkernel:249` unconditionally
|
||
appends `-DHEARTBEAT_THREAD_ENABLED=0` after knob forwarding, regardless of
|
||
what's requested — LithosAnanke has no pthreads. `HEARTBEAT_THREAD_ENABLED`
|
||
is modeled here as an ordinary hosted-side bool (default y), so a
|
||
kernel-variant `.config` will show it as `y` even though the actual kernel
|
||
build always forces it off. Resolving this (either a `depends on
|
||
!STARFORTH_VARIANT_KERNEL` restriction or a documented kept override) is
|
||
Phase 4's named verification target, not Phase 2's.
|
||
|
||
**Bridge** (`mk/Kconfig.mk`, `include`d by both Makefiles): per-architecture
|
||
config under `build/$(ARCH)/.config` rather than one root `.config` — the
|
||
project's own QEMU acceptance workflow builds all three kernel architectures
|
||
back-to-back in the same working directory, so a shared root config would
|
||
desync from whichever arch is actually being built. The two Makefiles
|
||
canonicalize `aarch64` differently for their own build-directory naming
|
||
(hosted `Makefile`: `aarch64` → `arm64`; `Makefile.starkernel`: `aarch64` →
|
||
`aarch64`), so `mk/Kconfig.mk` doesn't re-derive an architecture name at all
|
||
— it requires the including Makefile to set `KCONFIG_ARCH_DIR` to whatever
|
||
directory name *it* already builds objects under, immediately before the
|
||
`include`.
|
||
|
||
Two things confirmed only by direct observation while building this, not
|
||
foreseeable from reading `conf --help`:
|
||
|
||
1. **`conf`/`confdata.c` reads four separate env vars**, not the one
|
||
(`KCONFIG_CONFIG`) the Phase 1 plan had anticipated: `KCONFIG_CONFIG`
|
||
(`.config` itself), `KCONFIG_AUTOCONFIG` (`include/config/auto.conf`),
|
||
`KCONFIG_AUTOHEADER` (`include/generated/autoconf.h`), and
|
||
`KCONFIG_RUSTCCFG` (`include/generated/rustc_cfg`, written unconditionally
|
||
even though StarForth has no Rust code). Missing any one of the four lets
|
||
that file fall back to a bare `include/config`/`include/generated` path
|
||
relative to `$(CURDIR)` — i.e. leaking generated files straight into the
|
||
real source tree. Caught twice during validation (once via an interactive
|
||
`mconf` test run in Phase 1, once again here via a `%_defconfig` run that
|
||
only set three of the four vars) before all four were set together in
|
||
`mk/Kconfig.mk`.
|
||
2. **Including `mk/Kconfig.mk` before either Makefile's own `all:` target
|
||
silently hijacked Make's default goal.** `mk/Kconfig.mk`'s first rule
|
||
(`$(KCONFIG_CONF):`, building the `conf` binary) became the *first rule
|
||
Make had seen anywhere*, at the point `mk/Kconfig.mk` needed to be
|
||
included (immediately after `ARCH` is resolved, well before either
|
||
Makefile's own `all:` appears later in the file) — so a bare `make` with
|
||
no target built `tools/kconfig/conf` and stopped, never touching `all`.
|
||
Caught by literally running `make` after wiring the include and getting
|
||
`make: 'tools/kconfig/conf' is up to date.` as the *entire* output. Fixed
|
||
with an explicit `.DEFAULT_GOAL := all` in both Makefiles at the
|
||
`include mk/Kconfig.mk` site, immune to include order.
|
||
|
||
**Genuinely inert until opted into.** The `include $(KCONFIG_AUTOCONF_FILE)`
|
||
line — the one that would actually pull `CONFIG_*` variables into a Makefile
|
||
— is itself guarded behind `ifneq ($(wildcard $(KCONFIG_CONFIG_FILE)),)`. No
|
||
`build/$(ARCH)/.config` exists until a developer explicitly runs
|
||
`menuconfig`/`xconfig`/`config`/`oldconfig`/a `*_defconfig` target; until
|
||
then this guard is false and `mk/Kconfig.mk` contributes zero behavior
|
||
change to any existing invocation. This is what makes "wired but inert" true
|
||
in practice rather than true in name only — it was verified, not assumed,
|
||
by running `make help`/`make clean`/`make` against both Makefiles with no
|
||
`.config` present anywhere and confirming byte-identical output and behavior
|
||
to before this phase.
|
||
|
||
Four example defconfigs added (`configs/hosted_standard_defconfig`,
|
||
`configs/kernel_{amd64,aarch64,riscv64}_defconfig`), each just the minimal
|
||
choice selections (`ARCH_*`, `STARFORTH_VARIANT_*`, `TARGET_STANDARD`,
|
||
`PLATFORM_DEFAULT` where applicable) needed to steer `conf --defconfig`;
|
||
every other symbol resolves through its Kconfig `default` and was hand-diffed
|
||
against today's real Makefile/header defaults for all four combinations
|
||
before being trusted.
|
||
|
||
**No `-D` flag repointed.** Confirmed by grepping the actual link command a
|
||
plain hosted `make` emits — still bare `$(VAR)`-driven
|
||
(`-DSTRICT_PTR=1 -DENABLE_HOTWORDS_CACHE=0 ...`), no `CONFIG_` prefix
|
||
anywhere. Repointing specific knob families to `$(CONFIG_VAR)` is Phase 3+,
|
||
one family at a time.
|
||
|
||
**Three-arch acceptance.** Both Makefiles touched (the `.DEFAULT_GOAL` fix
|
||
and the bridge `include` land in `Makefile.starkernel` too), so the full
|
||
three-arch QEMU run was required before commit per this project's standing
|
||
rule. All four VM identities landed byte-identical across amd64/aarch64/
|
||
riscv64, matching rev t's values exactly: Hera `0x83c2c109100e2ed6`,
|
||
Artemis `0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births),
|
||
Mama `0xc88c3c1db6ef601b`. Hosted build clean, zero warnings.
|
||
|
||
## Rev v: Kconfig migration, Phase 3 — SSM/physics knob family cutover
|
||
|
||
First "live" phase: the ~18-symbol physics/SSM family (`STRICT_PTR`,
|
||
`ENABLE_HOTWORDS_CACHE`, `ENABLE_PIPELINING`, `ROLLING_WINDOW_SIZE`,
|
||
`TRANSITION_WINDOW_SIZE`, the four `ADAPTIVE_*` window-shrink knobs,
|
||
`INITIAL_DECAY_SLOPE_Q48`, `DECAY_MIN_INTERVAL`, `DECAY_RATE_PER_US_Q16`,
|
||
`HEARTBEAT_INFERENCE_FREQUENCY`, plus five previously-unwired `SSM_*` L8
|
||
thresholds) now actually reads from `$(CONFIG_VAR)` when a `.config`
|
||
exists, instead of Phase 2's "bridge present but nothing consumes it."
|
||
|
||
**Bridge macros** (`mk/Kconfig.mk`, `kconfig_bool`/`kconfig_int`): both
|
||
compile down to exactly `$(VAR) ?= $(default)` — today's behavior,
|
||
unchanged — whenever `KCONFIG_ACTIVE` is unset (no `.config` yet); once
|
||
active, `kconfig_bool` maps Kconfig's `y`/absent to `1`/`0` (the C side has
|
||
always taken 0/1 integers, never y/n), and `kconfig_int` takes
|
||
`$(or $(CONFIG_VAR),$(default))`, covering both "Kconfig inactive" and
|
||
"Kconfig active but this symbol's `depends on` was unmet so it never
|
||
appears in `auto.conf` at all" (e.g. `TRANSITION_WINDOW_SIZE` when
|
||
`ENABLE_PIPELINING=n`) with the same historical fallback rather than an
|
||
empty `-D`. Command-line/environment overrides of a knob always win
|
||
regardless of which branch fires — ordinary Make semantics, not something
|
||
this bridge has to implement itself. Verified directly: `make -n
|
||
ADAPTIVE_SHRINK_RATE=99` with an active `.config` requesting 77 still
|
||
produces `-DADAPTIVE_SHRINK_RATE=99`.
|
||
|
||
**Two build systems, two different existing philosophies, same bridge.**
|
||
The hosted `Makefile` blanket-forwards every knob unconditionally
|
||
(`?=` always emits a `-D`); each bare `?=` line became one
|
||
`$(eval $(call kconfig_bool/int,VAR,default))` call in place. The kernel
|
||
`Makefile.starkernel` only forwards a knob if a developer explicitly set
|
||
it (`origin($1) != undefined`) — Phase 3 preserves that asymmetry exactly:
|
||
a new `ifeq ($(KCONFIG_ACTIVE),1)` block pre-assigns each knob from its
|
||
`CONFIG_` value *before* the existing `add_vm_flag` foreach loop runs, so
|
||
`origin()` becomes `"file"` and the knob gets forwarded as if a developer
|
||
had typed it — but only when Kconfig is actually active. With no
|
||
`.config`, kernel-build behavior for this family is unchanged: nothing
|
||
forwarded unless explicitly requested, same as every kernel build before
|
||
this migration.
|
||
|
||
**A real, live bug found and fixed, not just a documented drift.**
|
||
`ADAPTIVE_SHRINK_RATE` (75 vs the Makefile's 50) and — newly discovered
|
||
here — `ADAPTIVE_GROWTH_THRESHOLD` (1 vs 5) had conflicting fallback
|
||
values in `rolling_window_knobs.h`, but both were confirmed *dormant* for
|
||
every consumer (`doe_metrics.c`, `rolling_window_of_truth.c`): both
|
||
transitively include `vm.h` before `rolling_window_knobs.h`, so
|
||
`starforth_config.h`'s `#ifndef` always wins the race regardless of
|
||
whether a `-D` flag was present. `TRANSITION_WINDOW_SIZE` (2 vs 8) in
|
||
`physics_pipelining_metrics.h` was not dormant: `physics_pipelining_metrics.c`
|
||
includes `physics_pipelining_metrics.h` *before* `vm.h`, and the kernel
|
||
Makefile's opt-in-only forwarding means no `-D` flag exists for this
|
||
symbol unless a developer explicitly sets one — so a stock kernel build
|
||
compiling that one translation unit was silently getting `2`, not the
|
||
intended `8`, whenever pipelining happened to be on. Never observed in
|
||
practice only because `ENABLE_PIPELINING` defaults off. Fix, applied
|
||
uniformly to all three duplicate-fallback headers (`rolling_window_knobs.h`,
|
||
`physics_pipelining_metrics.h`, and — extending the same treatment for
|
||
consistency, since starforth_config.h's own file-header comment declares
|
||
it "the single source of truth for VM build-time defaults" — `ssm_jacquard.h`,
|
||
which was not itself drifted but was a second independent copy of the same
|
||
five values): delete each header's own `#ifndef X #define X (own value)
|
||
#endif` block, add `#include "starforth_config.h"` at the top instead.
|
||
Confirmed safe: `starforth_config.h` has zero includes of its own, and
|
||
`vm.h` explicitly avoids including either of the other two headers
|
||
("avoids circular include" comments already in place), so no cycle risk.
|
||
|
||
**SSM_\* thresholds promoted, not just documented.** `ssm_jacquard.h`'s
|
||
five L8 mode-selector constants (`SSM_ENTROPY_HIGH_THRESHOLD`,
|
||
`SSM_CV_HIGH_THRESHOLD`, `SSM_TEMPORAL_DECAY_THRESHOLD`,
|
||
`SSM_TEMPORAL_DECAY_LOW_THRESHOLD`, `SSM_HYSTERESIS_TICKS`) had no Makefile
|
||
knob and no `-D` forwarding at all before this phase — genuinely unwired,
|
||
per the migration plan's own framing, "lowest risk since nothing currently
|
||
overrides them." Now wired identically to every other physics knob in both
|
||
Makefiles. The four threshold values are C `double`s compared against
|
||
`ssm_config_t` fields also typed `double`; Kconfig has no native float
|
||
symbol type, so they're modeled as Kconfig `string` symbols
|
||
(`default "0.75"` etc.) rather than `int`. Confirmed by direct
|
||
experimentation that this needs no quote-stripping on the Make side:
|
||
`conf`'s `.config` output keeps a string default quoted
|
||
(`CONFIG_FOO="0.75"`), but the `--syncconfig`-generated `auto.conf` that
|
||
Make actually `include`s writes string defaults unquoted
|
||
(`CONFIG_FOO=0.75`) — deliberately Make-syntax-safe. This means
|
||
`kconfig_int` (originally written only for true `int` symbols) already
|
||
handles these `string`-typed-but-numeric-literal symbols correctly with no
|
||
changes, and StarForth ended up not needing the separate `kconfig_str`
|
||
macro/quote-stripping logic originally sketched for this.
|
||
|
||
**Verification.** Every repointed default was hand-diffed against the
|
||
knob's actual pre-Phase-3 value (not the sometimes-stale Makefile comment
|
||
beside it — several comments described a different number than the `?=`
|
||
line actually set, e.g. `TRANSITION_WINDOW_SIZE`'s comment said "Default: 2"
|
||
next to a `?= 8` line; comments corrected in the same edit). End-to-end
|
||
round-trip confirmed by hand: generating a `.config` with
|
||
`ADAPTIVE_SHRINK_RATE=77`/`SSM_ENTROPY_HIGH_THRESHOLD="0.42"` and observing
|
||
`make -n` emit exactly those values in the link command; reverting and
|
||
confirming the inactive path emits the identical `-D` flag list, in the
|
||
identical order, as the pre-Phase-3 baseline. Both Makefiles touched
|
||
(`Makefile.starkernel`'s knob-forwarding block changed), so the full
|
||
three-arch QEMU acceptance ran again: all four VM identities landed
|
||
byte-identical to rev t/u's baseline — Hera `0x83c2c109100e2ed6`, Artemis
|
||
`0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births), Mama
|
||
`0xc88c3c1db6ef601b`. Hosted build clean, zero warnings, `-D` flag list
|
||
for a plain `make` byte-for-byte identical to before this phase aside from
|
||
the five newly-added `SSM_*` flags (new, not a repointing of anything that
|
||
existed before).
|
||
|
||
**One pre-existing, out-of-scope finding, reported rather than fixed per
|
||
standing instruction not to fix unrequested bugs:** `DECAY_MIN_INTERVAL`
|
||
has a `?=` default in the hosted Makefile but was never actually forwarded
|
||
as a `-D` flag in `BASE_CFLAGS` — the Make variable exists, but changing it
|
||
via `make DECAY_MIN_INTERVAL=999` has always had zero effect on
|
||
compilation, independent of and predating the tick-based decay conversion
|
||
(rev t) that made the value non-load-bearing at the C-logic level too.
|
||
Left exactly as found; not this phase's job to fix.
|
||
|
||
## Rev w: Kconfig migration, Phase 4 — heartbeat family cutover, HEARTBEAT_THREAD_ENABLED gate resolved
|
||
|
||
Repointed `EMERGENCY_CONSOLE_ENABLED`, `HEARTBEAT_THREAD_ENABLED`,
|
||
`HEARTBEAT_TICK_NS`, `HEARTBEAT_CHECK_FREQUENCY`,
|
||
`HEARTBEAT_WINDOW_TUNING_FREQUENCY`, `HEARTBEAT_SLOPE_VALIDATION_FREQUENCY`
|
||
through the same `kconfig_bool`/`kconfig_int` bridge as Phase 3's physics
|
||
family, in both Makefiles. Promoted the three `HEARTBEAT_CHECK_FREQUENCY`/
|
||
`HEARTBEAT_WINDOW_TUNING_FREQUENCY`/`HEARTBEAT_SLOPE_VALIDATION_FREQUENCY`
|
||
knobs into the hosted Makefile for the first time -- they had opt-in
|
||
forwarding in `Makefile.starkernel` already, but no `?=` and no `-D`
|
||
forwarding at all in the hosted `Makefile` before this phase, so hosted
|
||
builds always got `starforth_config.h`'s bare defaults with no override
|
||
mechanism. Same "promote a previously-unwired knob" treatment as Phase 3's
|
||
`SSM_*` constants.
|
||
|
||
**The named verification gate.** `Makefile.starkernel`'s
|
||
`VM_FEATURE_OVERRIDES += -DHEARTBEAT_THREAD_ENABLED=0` unconditionally
|
||
forces the kernel build's heartbeat thread off after knob forwarding,
|
||
regardless of what was requested -- LithosAnanke is freestanding, no
|
||
pthreads. The plan asked this phase to decide explicitly between modeling
|
||
it as a Kconfig `depends on` restriction or keeping the hardcoded
|
||
post-override. Chose both, deliberately: added
|
||
`depends on STARFORTH_VARIANT_HOSTED` to `HEARTBEAT_THREAD_ENABLED` in
|
||
`Kconfig.heartbeat` (so a kernel-variant `.config` never shows or sets it
|
||
-- `CONFIG_HEARTBEAT_THREAD_ENABLED` simply doesn't appear in that config's
|
||
`auto.conf`, and the `kconfig_bool` bridge correctly resolves the absence
|
||
to `0` with no special-casing needed), *and* kept
|
||
`Makefile.starkernel`'s own hardcoded override exactly as it was. The
|
||
`depends on` gets the Kconfig model right (menuconfig can't offer a choice
|
||
the kernel has no way to honor); the override guarantees the fact holds
|
||
even if the Kconfig tree, a hand-edited `auto.conf`, or a future refactor
|
||
ever disagrees. Belt-and-suspenders on purpose -- a freestanding kernel
|
||
image linking against pthreads that don't exist is a failure mode that's
|
||
silent until boot and expensive to debug on real hardware, not something
|
||
to trust to a single layer.
|
||
|
||
Verified directly, both branches: with no `.config`, the kernel build's
|
||
`-D` list contains only `-DHEARTBEAT_THREAD_ENABLED=0` (from the hardcoded
|
||
override) and no other heartbeat flags — identical to before this phase.
|
||
With a `kernel_amd64_defconfig`-derived `.config` active,
|
||
`-DHEARTBEAT_THREAD_ENABLED=0` still appears (now doubly-sourced: the
|
||
`depends on`-driven absence resolving to `0` via `kconfig_bool`, *and* the
|
||
hardcoded override — both agree, no redefinition conflict), while
|
||
`HEARTBEAT_TICK_NS`/`CHECK_FREQUENCY`/`WINDOW_TUNING_FREQUENCY`/
|
||
`SLOPE_VALIDATION_FREQUENCY` all forward correctly with their Kconfig
|
||
defaults. The hosted build's active path was checked too, confirming
|
||
`HEARTBEAT_THREAD_ENABLED=1` there (unaffected by the kernel-only
|
||
`depends on`, since `STARFORTH_VARIANT_HOSTED` config properly makes the
|
||
symbol visible again). `HEARTBEAT_TICK_NS` forwarding a fallback value for
|
||
a kernel build even though its own `depends on HEARTBEAT_THREAD_ENABLED`
|
||
also makes it Kconfig-invisible there is intentional, not a bug -- same
|
||
precedent as Phase 3's `TRANSITION_WINDOW_SIZE`-under-disabled-pipelining
|
||
case: `kconfig_int`'s `$(or ...)` always falls back to the historical
|
||
literal default rather than an empty `-D`, and the C side never reads
|
||
`HEARTBEAT_TICK_NS` when `HEARTBEAT_THREAD_ENABLED` is `0` regardless, so
|
||
this is a harmless no-op, not a behavior change.
|
||
|
||
**Repeated the same operational mistake from Phase 1/2, twice more, while
|
||
writing this phase.** A direct ad-hoc `tools/kconfig/conf --defconfig=...`
|
||
call used only to re-verify the new `depends on` relationship (not routed
|
||
through `make`) set `KCONFIG_CONFIG` but not the other three required env
|
||
vars, leaking `include/config/`+`include/generated/` into the real source
|
||
tree again -- the exact class of mistake already documented as a lesson in
|
||
rev u. Caught and cleaned up both times via `git status` before staging
|
||
anything. Recording again, more bluntly this time: there is no safe
|
||
shorthand for a one-off `conf` invocation outside `make` -- all four
|
||
(`KCONFIG_CONFIG`, `KCONFIG_AUTOCONFIG`, `KCONFIG_AUTOHEADER`,
|
||
`KCONFIG_RUSTCCFG`) or none of the ad-hoc-testing convenience is worth it.
|
||
|
||
**Verification.** Both Makefiles touched, so the full three-arch QEMU
|
||
acceptance ran again: all four VM identities byte-identical to the
|
||
established baseline -- Hera `0x83c2c109100e2ed6`, Artemis
|
||
`0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births), Mama
|
||
`0xc88c3c1db6ef601b`. Hosted build clean, zero warnings.
|
||
|
||
## Rev x: Kconfig migration, Phase 5 — platform-mode choice + remaining pipelining constants
|
||
|
||
Two independent pieces of work, both scoped to this phase by the plan.
|
||
|
||
**Platform-mode choice wired.** `Kconfig.variant`'s `PLATFORM_DEFAULT`/
|
||
`PLATFORM_MINIMAL`/`PLATFORM_L4RE` choice (written inert in Phase 2) now
|
||
actually drives the hosted `Makefile`'s `MINIMAL`/`L4RE` variables when a
|
||
`.config` selects something other than default. This family can't reuse
|
||
`kconfig_bool`/`kconfig_int`: `MINIMAL`/`L4RE` are tested with `ifdef`
|
||
downstream, which cares about *definedness*, not value, so unconditionally
|
||
defining either to 0 the way `kconfig_bool` does would make `ifdef` true
|
||
regardless of which platform was actually selected. Wrote the bridge
|
||
directly instead: only assign `MINIMAL := 1` or `L4RE := 1` (never both;
|
||
never at all for `PLATFORM_DEFAULT`) when Kconfig is active *and* neither
|
||
variable already has an explicit command-line/environment origin -- so
|
||
`make L4RE=1` always wins even if `.config` says `PLATFORM_MINIMAL`, a
|
||
case ordinary Make command-line precedence doesn't cover on its own since
|
||
it protects a variable from being overridden by an assignment to *itself*,
|
||
not from a *different* variable also becoming defined and winning an
|
||
`ifdef`/`else ifdef` chain.
|
||
|
||
**A real ordering bug, caught by testing, not by inspection.** First
|
||
placement of this bridge was directly above the `ifdef MINIMAL` platform
|
||
block near the bottom of the file -- seemed natural, right next to the
|
||
code it feeds. Testing a `PLATFORM_L4RE` config immediately showed
|
||
`HEARTBEAT_THREAD_ENABLED=1` and `-pthread` still present in the link
|
||
command, when both should be suppressed once `L4RE` is set. Root cause:
|
||
Make evaluates a file top-to-bottom, and `L4RE`'s existing
|
||
`ifeq ($(L4RE),1) HEARTBEAT_THREAD_ENABLED := 0 endif` heartbeat-override
|
||
check sits much earlier in the file (right after the heartbeat knobs) --
|
||
so at the point that check ran, `L4RE` hadn't been set yet by a bridge
|
||
placed near the bottom. Fixed by moving the whole platform-mode bridge to
|
||
immediately after `include mk/Kconfig.mk`, before anything else in the
|
||
file branches on `MINIMAL`/`L4RE`. Verified all three cases directly by
|
||
generating a `.config` for each and inspecting `make -n`'s actual command
|
||
line: `PLATFORM_MINIMAL` produces `-DSTARFORTH_MINIMAL=1 -nostdlib
|
||
-ffreestanding`; `PLATFORM_L4RE` produces `-D__l4__=1`,
|
||
`-DHEARTBEAT_THREAD_ENABLED=0`, and no `-pthread`; `PLATFORM_DEFAULT`
|
||
produces neither; and `make L4RE=1` against a `PLATFORM_MINIMAL` config
|
||
produces `-D__l4__=1` with no `STARFORTH_MINIMAL`/`ffreestanding` at
|
||
all -- command-line intent wins cleanly.
|
||
|
||
**Remaining `physics_pipelining_metrics.h` constants promoted.**
|
||
`SPECULATION_THRESHOLD_Q48`, `SPECULATION_DEPTH`,
|
||
`MIN_SAMPLES_FOR_SPECULATION`, `MISPREDICTION_COST_Q48`, and
|
||
`MINIMUM_PREFETCH_ROI` were plain, unconditional `#define`s with no
|
||
override mechanism at all before this phase (unlike `TRANSITION_WINDOW_SIZE`
|
||
and `ENABLE_PIPELINING` in the same file, already `#ifndef`-guarded and
|
||
cut over in Phase 3). Converted all five to the established pattern:
|
||
`#ifndef`-guarded, default sourced from `starforth_config.h`, wired through
|
||
`kconfig_int` in both Makefiles as Kconfig `hex` (the three Q48.16-encoded
|
||
values) or `int` (the two plain counts) symbols, gated `depends on
|
||
ENABLE_PIPELINING` alongside `TRANSITION_WINDOW_SIZE`. Verified the header
|
||
change alone (no Kconfig involved) with a clean hosted build at both
|
||
`ENABLE_PIPELINING=0` (today's default) and `=1`, before touching either
|
||
Makefile -- zero warnings both ways, confirming the `-D`-flag-always-wins
|
||
`#ifndef` pattern holds regardless of pipelining state.
|
||
|
||
**A second real bug found, not fixed.** Computing the exact Q48.16
|
||
encodings to transcribe into Kconfig defaults (rather than trusting the
|
||
existing C comments) surfaced a second drift bug, independent of anything
|
||
found in Phase 3: `MINIMUM_PREFETCH_ROI`'s own comment says
|
||
`1.10 = 1.10 * (1 << 16) = 0x11999AL`, but `1.10 * 65536 = 72089.6 ≈
|
||
0x1199A` (5 hex digits) -- the shipped constant `0x11999AL` (6 hex digits)
|
||
is `1153434`, which is `≈17.6` in Q48.16, not `1.10`. An order-of-magnitude
|
||
error that's been silently shipping in every build with pipelining enabled
|
||
since this constant was introduced, since ROI comparisons like
|
||
`(prefetch_latency_saved / total_attempts) > MINIMUM_PREFETCH_ROI` would
|
||
essentially never trigger speculation at 17.6 as the bar instead of 1.10.
|
||
Per standing instruction not to fix bugs found while doing unrelated work,
|
||
preserved the exact shipped value (`0x11999A`) as the Kconfig default and
|
||
`starforth_config.h` fallback, with an explicit comment at both sites
|
||
documenting the discrepancy and pointing back here. Flagging directly:
|
||
**Bob may want to fix `MINIMUM_PREFETCH_ROI` to `0x1199A` in a future
|
||
session** -- not done as part of this migration.
|
||
|
||
**Verification.** Both Makefiles touched, so the full three-arch QEMU
|
||
acceptance ran again: all four VM identities byte-identical to the
|
||
established baseline -- Hera `0x83c2c109100e2ed6`, Artemis
|
||
`0x5284ea5cd0f9983c`, Hermes `0x4159dcb326d79759` (both births), Mama
|
||
`0xc88c3c1db6ef601b`. Hosted build clean, zero warnings.
|
||
|
||
## Rev y: MINIMUM_PREFETCH_ROI fixed on Bob's explicit instruction
|
||
|
||
Rev x reported, but deliberately did not fix, that `MINIMUM_PREFETCH_ROI`
|
||
had been shipping as `0x11999A` (~17.6 in Q48.16) against a documented
|
||
intent of `1.10` (correct encoding `0x1199A`) -- an order-of-magnitude
|
||
error that's been silently suppressing prefetch speculation (the ROI bar
|
||
was ~16x higher than intended) in every build with `ENABLE_PIPELINING=1`
|
||
since the constant was introduced. Bob's instruction: fix it, then
|
||
continue to Phase 6.
|
||
|
||
Corrected the value at all four sites that had inherited the wrong
|
||
default while it was merely being made overridable (not yet fixed) in
|
||
Phase 5: `include/starforth_config.h`'s
|
||
`STARFORTH_CONFIG_MINIMUM_PREFETCH_ROI_DEFAULT`, `Kconfig.physics`'s
|
||
`MINIMUM_PREFETCH_ROI` symbol default, and the `kconfig_int` fallback
|
||
literal in both `Makefile` and `Makefile.starkernel`. Updated each site's
|
||
comment from "preserved as shipped, not fixed" to a plain statement of the
|
||
correct derivation, removing the now-stale "flagged, not fixed" language.
|
||
`physics_pipelining_metrics.h`'s own doc comment already stated the
|
||
correct `0x1199A` derivation (only its two upstream defaults were wrong),
|
||
so needed no correction, just removal of the discrepancy note pointing at
|
||
`starforth_config.h`.
|
||
|
||
Verified: a `.config` with `ENABLE_PIPELINING=y` now reports
|
||
`CONFIG_MINIMUM_PREFETCH_ROI=0x1199A`; a hosted build with
|
||
`ENABLE_PIPELINING=1` emits `-DMINIMUM_PREFETCH_ROI=0x1199A` in the actual
|
||
link command, zero warnings. `grep -rn 11999A` across the tree returns
|
||
only the historical-note comments explaining the fix, no live default.
|
||
Both Makefiles touched (again), so the full three-arch QEMU acceptance ran
|
||
once more: byte-identical to the established baseline -- Hera
|
||
`0x83c2c109100e2ed6`, Artemis `0x5284ea5cd0f9983c`, Hermes
|
||
`0x4159dcb326d79759` (both births), Mama `0xc88c3c1db6ef601b`. Unsurprising
|
||
that dict_hash is unaffected either way -- `MINIMUM_PREFETCH_ROI` governs
|
||
pipelining speculation decisions, not dictionary word heat, and the
|
||
kernel's default build has `ENABLE_PIPELINING=0` regardless -- but the
|
||
acceptance bar is unconditional per project rule, so it ran anyway.
|
||
|
||
## Rev z: Kconfig migration, Phase 6 — kernel-only flags, plan complete
|
||
|
||
Final phase. `PARITY_MODE`, `STARFORTH_ENABLE_VM`, `SK_PARITY_DEBUG`,
|
||
`HEARTBEAT_DOE_LOG` -- all kernel-only, no hosted equivalent -- repointed
|
||
through the Kconfig bridge in a new `Kconfig.kernel` (sourced from the
|
||
root `Kconfig`, wrapped in `if STARFORTH_VARIANT_KERNEL ... endif` so
|
||
these four symbols simply don't exist for a hosted-variant `.config`).
|
||
`STARFORTH_ENABLE_VM` and `PARITY_MODE` had bare `?=` defaults already
|
||
(blanket-forwarded); repointed via `kconfig_bool` in place, same as every
|
||
other blanket-forwarded knob this migration. `HEARTBEAT_DOE_LOG` likewise.
|
||
`SK_PARITY_DEBUG` had no bare default at all -- opt-in-only via
|
||
`VM_FEATURE_FLAG_VARS`, and (per this project's convention for that
|
||
family) pre-assigned from its `CONFIG_` value inside the existing
|
||
`ifeq ($(KCONFIG_ACTIVE),1)` block, same treatment as the `SSM_*`/
|
||
pipelining constants in Phases 3/5.
|
||
|
||
**Doesn't touch the hosted `Makefile` at all**, exactly as the plan
|
||
specified -- confirmed directly, not just by omission: `git status` after
|
||
this phase's edits shows only `Kconfig`, `Kconfig.kernel` (new), and
|
||
`Makefile.starkernel` changed, and a full hosted `make clean && make` ran
|
||
clean with zero warnings and a byte-identical link command to before this
|
||
phase. This doubles as the regression check the plan asked for: five
|
||
phases of Kconfig work landing entirely inside `Makefile.starkernel`,
|
||
`mk/Kconfig.mk`, and the `Kconfig*` tree, with the hosted build path
|
||
provably undisturbed.
|
||
|
||
Verified both branches on the kernel side directly, not just by
|
||
inspection: with no `.config`, the kernel `-D` list is
|
||
`PARITY_MODE=0 STARFORTH_ENABLE_VM=1 HEARTBEAT_DOE_LOG=1` and no
|
||
`SK_PARITY_DEBUG` flag at all -- byte-identical to pre-Phase-6 behavior.
|
||
With a `kernel_amd64_defconfig`-derived `.config`, the same defaults
|
||
appear plus `-DSK_PARITY_DEBUG=0` (now forwarded, correctly at its
|
||
default). A hand-built `.config` flipping all four
|
||
(`PARITY_MODE=y SK_PARITY_DEBUG=y HEARTBEAT_DOE_LOG=n`) produced exactly
|
||
`-DPARITY_MODE=1 -DSK_PARITY_DEBUG=1 -DHEARTBEAT_DOE_LOG=0
|
||
-DSTARFORTH_ENABLE_VM=1` in the actual `make -n` command line.
|
||
|
||
Both Makefiles touched by the migration as a whole across all six phases,
|
||
kernel-only this phase specifically, so the full three-arch QEMU
|
||
acceptance ran one final time: all four VM identities byte-identical to
|
||
the baseline established at rev t and held through every phase since --
|
||
Hera `0x83c2c109100e2ed6`, Artemis `0x5284ea5cd0f9983c`, Hermes
|
||
`0x4159dcb326d79759` (both births), Mama `0xc88c3c1db6ef601b`. Hosted
|
||
build clean, zero warnings.
|
||
|
||
**Kconfig migration plan complete.** Six phases, six commits (plus one
|
||
out-of-band fix commit at Bob's request between Phase 5 and Phase 6),
|
||
zero behavioral regressions at any step, two real pre-existing bugs found
|
||
and reported (`TRANSITION_WINDOW_SIZE`'s live-not-latent kernel-build
|
||
drift, Phase 3; `MINIMUM_PREFETCH_ROI`'s order-of-magnitude Q48.16
|
||
encoding error, Phase 5), one of the two fixed on explicit instruction
|
||
(rev y). Every physics/SSM/heartbeat/pipelining/kernel-only knob that had
|
||
a Makefile presence before this migration still has one, now sourced from
|
||
a single discoverable Kconfig tree (`Kconfig` + `Kconfig.arch`/
|
||
`Kconfig.variant`/`Kconfig.physics`/`Kconfig.heartbeat`/`Kconfig.kernel`)
|
||
when a developer opts in via `menuconfig`/`xconfig`/`config`/`oldconfig`/
|
||
a `*_defconfig` target, and unchanged when they don't. `CDTuning` values
|
||
(`compudynamics.h`) and runtime-only DoE/profiling parameters remain
|
||
explicit non-goals, as scoped from the start.
|