Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Third stage of the preemptive context-switching plan. The real save/restore switch mechanism now exists -- the first time anything has ever executed on a VM's own native stack (Stage 1 allocated them, unused). New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already covers every caller-saved register; only the callee-saved set needs explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for both halves within one call, since FS is genuine global CPU state, not part of what's switched). A sibling sk_vm_switch_prime() in the same file builds the synthetic first-entry frame, kept in assembly so the layout can never drift out of sync with sk_vm_switch_to() itself. New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry priming vs. resuming a parked context, and updates registry state (new VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no live frame, this means the opposite). sk_vm_switch_entry() is the minimal permanent trampoline every freshly-entered VM lands in: no production behavior defined yet, so it just yields straight back to whoever switched to it, forever. Closes the confirmed unguarded-KILL UAF found during planning: capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama() all now refuse (or silently leak rather than free, on the cold-restart path where arch_cold_reset() wipes everything immediately after anyway) tearing down a switched-out VM. Side effect found, not built on purpose: the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it automatically stopped dispatching into a switched-out VM with zero changes needed there. Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing can type interactively into a foreground-only QEMU session) that round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3 architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe fully reverted after capture; kernel_main.c shows zero diff. Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new files added explicitly (this project's "loader" PE binary is the full running kernel, not a thin bootstrap stage), and aarch64's switch.S needed the same #ifndef _WIN32 guard around .hidden that isr.S already carries (aarch64's loader assembles via clang targeting a PE/COFF target with no .hidden equivalent) -- caught by a build failure, fixed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
This commit is contained in:
co-authored by
Claude Sonnet 5
parent
57ac3fc304
commit
f790d0995e
+56
@@ -3579,3 +3579,59 @@ All 3 architectures re-verified clean boot to `ok>`, no native-stack allocation
|
||||
for any Tripod-fleet VM on any arch. Stage 2 (the actual save/restore switch primitive,
|
||||
cooperative only, no timer) is next.
|
||||
|
||||
**Stage 2 CLOSED, same day.** The real save/restore switch primitive now exists and is proven
|
||||
correct on all 3 architectures -- the first time anything has ever executed on a VM's own
|
||||
native stack (Stage 1's stacks were allocated but unused until now).
|
||||
|
||||
Deliberately NOT built the same way as Stage 0's interrupt trap frame -- `sk_vm_switch_to()`
|
||||
(new `switch.S` beside each arch's `isr.S`) is an ordinary function call, so the standard ABI
|
||||
already guarantees every caller-saved register is the caller's own problem; only the
|
||||
callee-saved set needs explicit handling (amd64: rbx/rbp/r12-r15, no FP at all since SysV has
|
||||
no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64: s0-s11/ra + fs0-fs11, the FP
|
||||
half still correctly FS-gated like Stage 0's trap frame, but read once and reused for both the
|
||||
save and restore decision within one call rather than persisted -- FS is genuine single global
|
||||
CPU state, not part of what's "switched"). A sibling `sk_vm_switch_prime()` in each same file
|
||||
builds the synthetic first-entry frame (zeroed callee-saved slots + a trampoline address in the
|
||||
return-address slot) so a VM's first-ever switch-in lands somewhere real instead of into
|
||||
whatever real prior caller the layout would otherwise imply -- kept in assembly, not C, so the
|
||||
frame layout can never drift out of sync between the two functions.
|
||||
|
||||
New C glue (`vm/switch.c`, `vm/switch.h`): `sk_vm_context_switch(from, to)` -- handles first-entry
|
||||
priming vs. resuming a previously-parked context, and updates registry state around the switch
|
||||
(new `VM_STATE_SWITCHED_OUT`, distinct from `VM_STATE_STOPPED` per the Stage 1 correction --
|
||||
STOPPED means no live frame, SWITCHED_OUT means the opposite). `sk_vm_switch_entry()` is the
|
||||
minimal permanent trampoline every freshly-entered VM lands in: no production behavior is
|
||||
defined yet (that's later-stage work), so it just yields straight back to whoever switched to
|
||||
it, forever -- must never fall off the end, since there is no legitimate return address below a
|
||||
synthetic first-entry frame.
|
||||
|
||||
**Closed the confirmed unguarded-KILL UAF** found during planning: `capsule_vm_kill()`,
|
||||
`mama_word_kill()`, and `capsule_vm_kill_all_nonmama()` all now refuse (or, for the
|
||||
cold-restart-only path, silently leak rather than free -- harmless there specifically, since
|
||||
`arch_cold_reset()` always follows immediately and wipes everything regardless) tearing down a
|
||||
VM in `VM_STATE_SWITCHED_OUT`. A nice side effect discovered while verifying, not something
|
||||
built deliberately: the existing MSG-TICK idle-pump (`repl.c`) already filters on
|
||||
`VM_STATE_LIVE`, so it automatically stopped dispatching into Hermes the moment she was
|
||||
switched out, with zero code changes needed there.
|
||||
|
||||
**Verification, per this project's write-probe-capture-revert discipline:** a temporary
|
||||
`SWITCH-TEST` word (registered in Hera's own vocabulary, triggered once at boot right after
|
||||
Hermes's birth, since there's no way to type interactively into a foreground-only QEMU session)
|
||||
round-tripped a distinctive sentinel through 5 real Hera<->Hermes context switches. All 3
|
||||
architectures: 5/5 rounds, 0 failures, PASS, clean continuation of boot to `ok>` afterward. The
|
||||
probe word, its boot-time trigger, and its helper were then fully reverted -- re-verified clean
|
||||
on all 3 architectures again with the probe gone, confirming no residual effects (Hermes back to
|
||||
normal `VM_STATE_LIVE` MSG-TICK dispatch once no longer switched out).
|
||||
|
||||
Also fixed along the way: `Makefile.starkernel`'s `LOADER_EXTRA_SRCS`/`LOADER_ASM` lists needed
|
||||
`switch.c`/`switch.S` added explicitly (this project's "loader" PE binary is effectively the
|
||||
full running kernel, not a thin bootstrap stage -- confirmed by `isr.S` already being in
|
||||
`LOADER_ASM`); and aarch64's `switch.S` needed the same `#ifndef _WIN32` guard around `.hidden`
|
||||
that `isr.S` already carries, since aarch64's loader is assembled via clang targeting
|
||||
`aarch64-pc-windows-msvc` (PE/COFF, no `.hidden` equivalent) -- missed on the first pass, caught
|
||||
by the aarch64 build failing outright.
|
||||
|
||||
Stage 3 (timer-driven preemption, Tripod fleet only, the new run-readiness signal) is next --
|
||||
the biggest remaining stage, and the first one that actually amends the §21.1
|
||||
nothing-is-concurrent-on-one-hart ruling for real.
|
||||
|
||||
|
||||
Reference in New Issue
Block a user