Stage 2: cooperative VM context switch primitive, proven on all 3 arches (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run

Third stage of the preemptive context-switching plan. The real
save/restore switch mechanism now exists -- the first time anything has
ever executed on a VM's own native stack (Stage 1 allocated them,
unused).

New sk_vm_switch_to() (switch.S, one per arch) is an ordinary function
call, not an interrupt -- so unlike Stage 0's trap frame, the ABI already
covers every caller-saved register; only the callee-saved set needs
explicit save/restore (amd64: rbx/rbp/r12-r15, no FP at all since SysV
has no callee-saved XMM; aarch64: x19-x28/x29/x30 + d8-d15; riscv64:
s0-s11/ra + fs0-fs11, FS-gated like Stage 0 but read once and reused for
both halves within one call, since FS is genuine global CPU state, not
part of what's switched). A sibling sk_vm_switch_prime() in the same file
builds the synthetic first-entry frame, kept in assembly so the layout
can never drift out of sync with sk_vm_switch_to() itself.

New switch.c/switch.h: sk_vm_context_switch(from, to) handles first-entry
priming vs. resuming a parked context, and updates registry state (new
VM_STATE_SWITCHED_OUT, distinct from VM_STATE_STOPPED -- STOPPED means no
live frame, this means the opposite). sk_vm_switch_entry() is the minimal
permanent trampoline every freshly-entered VM lands in: no production
behavior defined yet, so it just yields straight back to whoever switched
to it, forever.

Closes the confirmed unguarded-KILL UAF found during planning:
capsule_vm_kill(), mama_word_kill(), and capsule_vm_kill_all_nonmama()
all now refuse (or silently leak rather than free, on the cold-restart
path where arch_cold_reset() wipes everything immediately after anyway)
tearing down a switched-out VM. Side effect found, not built on purpose:
the existing MSG-TICK idle-pump already filters on VM_STATE_LIVE, so it
automatically stopped dispatching into a switched-out VM with zero
changes needed there.

Verified via a temporary SWITCH-TEST probe (boot-triggered, since nothing
can type interactively into a foreground-only QEMU session) that
round-tripped a sentinel through 5 real Hera<->Hermes switches on all 3
architectures: 5/5 rounds, 0 failures, clean continuation to ok>. Probe
fully reverted after capture; kernel_main.c shows zero diff.

Also: Makefile.starkernel's LOADER_EXTRA_SRCS/LOADER_ASM needed the new
files added explicitly (this project's "loader" PE binary is the full
running kernel, not a thin bootstrap stage), and aarch64's switch.S
needed the same #ifndef _WIN32 guard around .hidden that isr.S already
carries (aarch64's loader assembles via clang targeting a PE/COFF
target with no .hidden equivalent) -- caught by a build failure, fixed.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
This commit is contained in:
Robert Allan James
2026-09-13 15:25:39 -04:00
co-authored by Claude Sonnet 5
parent 57ac3fc304
commit f790d0995e
25 changed files with 54894 additions and 7 deletions
+98
View File
@@ -0,0 +1,98 @@
/*
StarForth — Steady-State Virtual Machine Runtime
Copyright (c) 20232025 Robert A. James
All rights reserved.
This file is part of the StarForth project.
Licensed under the StarForth License, Version 1.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at:
https://github.com/star.4th@proton.me/StarForth/LICENSE.txt
This software is provided "AS IS", WITHOUT WARRANTY OF ANY KIND,
express or implied, including but not limited to the warranties of
merchantability, fitness for a particular purpose, and noninfringement.
See the License for the specific language governing permissions and
limitations under the License.
*/
/**
* switch.c - Cooperative VM context switch glue (FABRIC-3.md §XXVIII,
* Stage 2, 2026-09-13). See switch.h for the contract.
*/
#ifdef __STARKERNEL__
#include "starkernel/vm/switch.h"
#include "vm.h"
#include "starkernel/capsule_birth.h" /* capsule_vm_set_state */
#include <stdint.h>
/* Arch-specific (switch.S in each arch dir). sk_vm_switch_prime() builds
* the exact synthetic first-entry frame sk_vm_switch_to() expects -- kept
* in assembly, not duplicated here, so the layout can never drift out of
* sync between the two. */
extern void sk_vm_switch_to(uint64_t *save_sp_out, uint64_t new_sp);
extern uint64_t sk_vm_switch_prime(uint64_t stack_top, void *entry);
/* Single-writer-mainline globals, safe under this stage's cooperative-
* only, one-VM-runs-at-a-time-on-one-hart discipline (same class of
* bookkeeping as vm_core.c's g_log_attrib_vm) -- set immediately before a
* first-ever switch into a VM, read exactly once, synchronously, at the
* very top of sk_vm_switch_entry() before any other switch could occur.
* Preemption (Stage 3) will need to revisit this; flagged there, not here. */
static VM *g_switch_entry_vm;
static VM *g_switch_back_to;
/**
* sk_vm_switch_entry - trampoline for a VM's first-ever entry.
*
* A freshly-entered VM has no production behavior defined yet -- what a
* switched-to VM should actually DO once running is later-stage work
* (message dispatch, an idle loop, whatever Stage 3+ needs). For now this
* is a minimal, permanent placeholder: immediately yield back to whoever
* switched to it, forever. Must never fall off the end -- there is no
* legitimate return address below a synthetic first-entry frame, unlike
* an ordinary function.
*/
void sk_vm_switch_entry(void) {
VM *self = g_switch_entry_vm;
VM *back = g_switch_back_to;
for (;;) {
sk_vm_context_switch(self, back);
}
}
int sk_vm_context_switch(VM *from, VM *to) {
uint64_t new_sp;
if (!from || !to || to->native_stack_top == 0) {
return -1;
}
if (to->native_stack_saved_sp == 0) {
/* First-ever entry: synthesize the initial frame. */
g_switch_entry_vm = to;
g_switch_back_to = from;
new_sp = sk_vm_switch_prime(to->native_stack_top, (void *)&sk_vm_switch_entry);
} else {
new_sp = to->native_stack_saved_sp;
to->native_stack_saved_sp = 0; /* about to be running, not parked */
}
capsule_vm_set_state(to->stadium_vm_id, VM_STATE_LIVE);
capsule_vm_set_state(from->stadium_vm_id, VM_STATE_SWITCHED_OUT);
sk_vm_switch_to(&from->native_stack_saved_sp, new_sp);
/* Resumes here only once something later switches back to `from` --
* from `from`'s own point of view, this call simply took a while. */
capsule_vm_set_state(from->stadium_vm_id, VM_STATE_LIVE);
return 0;
}
#endif /* __STARKERNEL__ */