Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run

Second stage of the preemptive context-switching plan. Every VM (Hera,
every capsule_birth_baby()-born VM including WIREBIND identities) now
gets its own dedicated 2 MiB native C stack at birth -- but nothing runs
on it yet, that's Stage 2. Pure allocation-machinery proof.

Design correction made before writing code: the plan called for cloning
sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to
be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a
plain kmalloc() block for its dictionary arena, not a real guarded PMM
allocation. Stacks get the real treatment instead (new
sk_vm_native_stack_alloc()/_free() in arena.c): independent
pmm_alloc_contiguous() + guard pages for every VM without exception, no
singleton, no kmalloc fallback -- a stack overflow is exactly the
failure mode guard pages exist for, and a corrupted stack could corrupt
whatever saved context Stage 2 trusts.

2 MiB size matches this project's own established kernel-stack
convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one
shared 2 MiB stack today already carries all VMs' combined nested
VM-EXEC recursion.

Three new VM struct fields, freed in vm_cleanup() alongside the existing
call_stack free. Allocation failure is non-fatal to birth.

All 3 architectures re-verified clean boot to ok>, no native-stack
allocation failures for any Tripod-fleet VM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
This commit is contained in:
Robert Allan James
2026-09-13 14:30:46 -04:00
co-authored by Claude Sonnet 5
parent 15672ce17c
commit 57ac3fc304
15 changed files with 27337 additions and 1 deletions
+26
View File
@@ -3553,3 +3553,29 @@ No FORTH-visible behavior changed, as intended -- this stage was purely "make th
complete," nothing yet reads or writes any of it beyond the ISR's own entry/exit. Stage 1 (per-VM
native stacks) is next.
**Stage 1 CLOSED, same day.** Every VM now gets its own dedicated 2 MiB native (C) stack at
birth, allocated but not yet executed on. Design correction made before writing any code: the
plan's own text said to clone `sk_vm_arena_alloc()`'s guard-page pattern, but investigation
showed that pattern is actually Mama-only -- `host_services.c`'s `kernel_alloc()` gives every
baby VM a plain `kmalloc()` block for its 5 MB dictionary arena, not a real guarded PMM
allocation; only Hera gets the true singleton. Stacks are different: a stack overflow is
precisely the failure mode guard pages exist for, and a corrupted stack can corrupt whatever
saved context Stage 2 later trusts, so the new `sk_vm_native_stack_alloc()`/`_free()`
(`arena.c`) gives **every** VM without exception a real, independent `pmm_alloc_contiguous()`
allocation with guard pages -- no singleton, no kmalloc fallback. Size (2 MiB) is not a guess:
it matches this project's own established kernel-stack convention (`g_kernel_stack`/
`g_rpi5_native_stack` in the arch entry files), and that single 2 MiB stack today already
carries the entire shared kernel C stack's worth of nested VM-EXEC/`execute_colon_word`
recursion across every VM combined -- generous headroom for one VM alone.
Wired at both VM-creation sites: `sk_vm_bootstrap_parity()` (Hera) and `capsule_birth_baby()`
(every other VM, WIREBIND-birthed identities included -- purely additive there, no attach/detach
behavior change). Three new fields on `VM` (`native_stack_paddr`/`native_stack_guard_vaddr`/
`native_stack_top`), freed in `vm_cleanup()` alongside the existing `call_stack` free.
Allocation failure is deliberately non-fatal to birth (degrades to "not yet switchable," not
"stillborn," since nothing executes on these stacks yet regardless).
All 3 architectures re-verified clean boot to `ok>`, no native-stack allocation failures logged
for any Tripod-fleet VM on any arch. Stage 2 (the actual save/restore switch primitive,
cooperative only, no timer) is next.