Stage 1: per-VM native stacks, allocated but not yet executed on (FABRIC-3.md §XXVIII)
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run

Second stage of the preemptive context-switching plan. Every VM (Hera,
every capsule_birth_baby()-born VM including WIREBIND identities) now
gets its own dedicated 2 MiB native C stack at birth -- but nothing runs
on it yet, that's Stage 2. Pure allocation-machinery proof.

Design correction made before writing code: the plan called for cloning
sk_vm_arena_alloc()'s guard-page pattern, but that pattern turns out to
be Mama-only -- host_services.c's kernel_alloc() gives every baby VM a
plain kmalloc() block for its dictionary arena, not a real guarded PMM
allocation. Stacks get the real treatment instead (new
sk_vm_native_stack_alloc()/_free() in arena.c): independent
pmm_alloc_contiguous() + guard pages for every VM without exception, no
singleton, no kmalloc fallback -- a stack overflow is exactly the
failure mode guard pages exist for, and a corrupted stack could corrupt
whatever saved context Stage 2 trusts.

2 MiB size matches this project's own established kernel-stack
convention (g_kernel_stack/g_rpi5_native_stack), not a guess -- that one
shared 2 MiB stack today already carries all VMs' combined nested
VM-EXEC recursion.

Three new VM struct fields, freed in vm_cleanup() alongside the existing
call_stack free. Allocation failure is non-fatal to birth.

All 3 architectures re-verified clean boot to ok>, no native-stack
allocation failures for any Tripod-fleet VM.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016UNhH1mhi52i6Qihh7ZV5S
This commit is contained in:
Robert Allan James
2026-09-13 14:30:46 -04:00
co-authored by Claude Sonnet 5
parent 15672ce17c
commit 57ac3fc304
15 changed files with 27337 additions and 1 deletions
+29
View File
@@ -62,6 +62,35 @@ size_t sk_vm_arena_size(void);
int sk_vm_arena_is_initialized(void);
void sk_vm_arena_assert_guards(const char *tag);
/**
* sk_vm_native_stack_t - one VM's own dedicated native (C) stack.
*
* FABRIC-3.md §XXVIII (Stage 1, preemptive-context-switch per-VM native
* stacks, 2026-09-13). Unlike sk_vm_arena_alloc() -- whose PMM+guard-page
* path is exercised only for Hera; every baby VM's "arena" is actually a
* plain kmalloc block, per host_services.c's kernel_alloc() -- every native
* stack, for every VM without exception, gets its own real, independent
* pmm_alloc_contiguous() allocation with guard pages. A stack overflow is
* exactly the failure mode guard pages exist for, and unlike the dictionary
* arena, a corrupted stack can also corrupt whatever saved context Stage 2
* later trusts -- worth the extra PMM pages every VM, not just Hera.
*
* Caller owns this struct (stored directly on the VM, mirroring how
* vm->memory already holds its own arena pointer directly rather than going
* through any module-level registry) and passes it back unchanged to
* sk_vm_native_stack_free().
*/
typedef struct {
uint64_t paddr; /* physical base (for pmm_free_contiguous) */
uint64_t guard_vaddr; /* virtual base of the whole guarded region */
uint64_t stack_top; /* initial SP value -- stacks grow down on all
* 3 arches, so this is guard_vaddr + one guard
* page + the full stack size */
} sk_vm_native_stack_t;
int sk_vm_native_stack_alloc(sk_vm_native_stack_t *out);
void sk_vm_native_stack_free(const sk_vm_native_stack_t *stack);
#endif /* __STARKERNEL__ */
#endif /* STARKERNEL_VM_ARENA_H */