Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run

Third stage of the preemptive context-switching plan. The real
save/restore switch mechanism now exists -- the first time anything has
ever executed on a VM's own native stack (Stage 1 allocated them,
unused).

New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function
call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already
covers every caller-saved register; only the callee-saved set needs
explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV
has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64:
s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for
both halves within one call, since FS is genuine global CPU state, not
part of what's switched). A sibling sk_vm_switch_prime() in the same file
builds the synthetic first-entry frame, kept in assembly so the layout
can never drift out of sync with sk_vm_switch_to() itself.

New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry
priming vs. resuming a parked context, and updates registry state (new
VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no
live frame, this means the opposite). sk_vm_switch_entry() is the minimal
permanent trampoline every freshly-entered VM lands in: no production
behavior defined yet, so it just yields straight back to whoever switched
to it, forever.

Closes the confirmed unguarded-KILL UAF found during planning:
capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama()
all now refuse (or silently leak rather than free, on the cold-restart
path where arch_cold_reset() wipes everything immediately after anyway)
tearing down a switched-out VM. Side effect found, not built on purpose:
the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it
automatically stopped dispatching into a switched-out VM with zero
changes needed there.

Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing
can type interactively into a foreground-only QEMU session) that
round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3
architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe
fully reverted after capture; kernel_main.c shows zero diff.

Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new
files added explicitly (this project's "loader" PE binary is the full
running kernel, not a thin bootstrap stage), and aarch64's switch.S
needed the same #ifndef _WIN32 guard around .hidden that isr.S already
carries (aarch64's loader assembles via clang targeting a PE/COFF
target with no .hidden equivalent) -- caught by a build failure, fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
This commit is contained in:
Robert Allan James
2026-09-13 15:25:39 -04:00
co-authored by Claude Sonnet 5
parent 57ac3fc304
commit f790d0995e
25 changed files with 54894 additions and 7 deletions
+56
View File
@@ -3579,3 +3579,59 @@ All 3 architectures re-verified clean boot to `ok>`, no native-stack allocation
for any Tripod-fleet VM on any arch. Stage 2 (the actual save/restore switch primitive,
cooperative only, no timer) is next.
**Stage 2 CLOSED, same day.** The real save/restore switch primitive now exists and is proven
correct on all 3 architectures -- the first time anything has ever executed on a VM's own
native stack (Stage 1's stacks were allocated but unused until now).
Deliberately NOT built the same way as Stage 0's interrupt trap frame -- `sk_vm_switch_to()`
(new `switch.S` beside each arch's `isr.S`) is an ordinary function call, so the standard ABI
already guarantees every caller-saved register is the caller's own problem; only the
callee-saved set needs explicit handling (amd64: rbx/rbp/r12-r15, no FP at all since SysV has
no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, the FP
half still correctly FS-gated like Stage 0's trap frame, but read once and reused for both the
save and restore decision within one call rather than persisted -- FS is genuine single global
CPU state, not part of what's "switched"). A sibling `sk_vm_switch_prime()` in each same file
builds the synthetic first-entry frame (zeroed callee-saved slots + a trampoline address in the
return-address slot) so a VM's first-ever switch-in lands somewhere real instead of into
whatever real prior caller the layout would otherwise imply -- kept in assembly, not C, so the
frame layout can never drift out of sync between the two functions.
New C glue (`vm/switch.c`, `vm/switch.h`): `sk_vm_context_switch(from, to)` -- handles first-entry
priming vs. resuming a previously-parked context, and updates registry state around the switch
(new `VM_STATE_SWITCHED_OUT`, distinct from `VM_STATE_STOPPED` per the Stage 1 correction --
STOPPED means no live frame, SWITCHED_OUT means the opposite). `sk_vm_switch_entry()` is the
minimal permanent trampoline every freshly-entered VM lands in: no production behavior is
defined yet (that's later-stage work), so it just yields straight back to whoever switched to
it, forever -- must never fall off the end, since there is no legitimate return address below a
synthetic first-entry frame.
**Closed the confirmed unguarded-KILL UAF** found during planning: `capsule_vm_kill()`,
`mama_word_kill()`, and `capsule_vm_kill_all_nonmama()` all now refuse (or, for the
cold-restart-only path, silently leak rather than free -- harmless there specifically, since
`arch_cold_reset()` always follows immediately and wipes everything regardless) tearing down a
VM in `VM_STATE_SWITCHED_OUT`. A nice side effect discovered while verifying, not something
built deliberately: the existing MSG-TICK idle-pump (`repl.c`) already filters on
`VM_STATE_LIVE`, so it automatically stopped dispatching into Hermes the moment she was
switched out, with zero code changes needed there.
**Verification, per this project's write-probe-capture-revert discipline:** a temporary
`SWITCH-TEST` word (registered in Hera's own vocabulary, triggered once at boot right after
Hermes's birth, since there's no way to type interactively into a foreground-only QEMU session)
round-tripped a distinctive sentinel through 5 real Hera<->Hermes context switches. All 3
architectures: 5/5 rounds, 0 failures, PASS, clean continuation of boot to `ok>` afterward. The
probe word, its boot-time trigger, and its helper were then fully reverted -- re-verified clean
on all 3 architectures again with the probe gone, confirming no residual effects (Hermes back to
normal `VM_STATE_LIVE` MSG-TICK dispatch once no longer switched out).
Also fixed along the way: `Makefile.starkernel`'s `LOADER_EXTRA_SRCS`/`LOADER_ASM` lists needed
`switch.c`/`switch.S` added explicitly (this project's "loader" PE binary is effectively the
full running kernel, not a thin bootstrap stage -- confirmed by `isr.S` already being in
`LOADER_ASM`); and aarch64's `switch.S` needed the same `#ifndef _WIN32` guard around `.hidden`
that `isr.S` already carries, since aarch64's loader is assembled via clang targeting
`aarch64-pc-windows-msvc` (PE/COFF, no `.hidden` equivalent) -- missed on the first pass, caught
by the aarch64 build failing outright.
Stage 3 (timer-driven preemption, Tripod fleet only, the new run-readiness signal) is next --
the biggest remaining stage, and the first one that actually amends the §21.1
nothing-is-concurrent-on-one-hart ruling for real.