FABRIC.md did its job: §1-24's design argument is settled and every
implementation item through 4.5/4.4ac either landed or was explicitly
deferred with a reason. At 7,595 lines it was no longer a good place
to find what's actually still open, so it's now archival -- header
rewritten to say so, pointing to FABRIC-2.md.
Before closing it, read the entire document end to end (not sampled)
looking for anything unresolved: punch-list checkboxes, the nine
"### N.N Open" architectural subsections in §1-24, the §25.7
"reported, not scheduled" list, and any other "not yet"/"deferred"
language. Found and fixed four stale bookkeeping spots where later
work had actually resolved something but the note was never updated:
§19.6 #3 (resolved by item 2.1), §21.5 #4 (resolved by §20.5 #4), the
§25.7 stadium_owner[idx] bullet (resolved by item 4.2), and item 4.5's
own parent checkbox (all six sub-items 4.5a-4.5f were already [x]).
FABRIC-2.md carries forward everything genuinely still open: the
blocked/scoped punch-list items (1.11, 4.3, 4.4s, 4.6, 5.1-5.3, plus a
specific pending TRIPOD.md edit found within 5.3), two regressions
that were invisible with Tripod pruned to Hera-alone and are now live
since item 4.2 restored Hermes (the fleet heat leak in
vm_physics_touch(), and multi-VM heartbeat ownership), nine dead-code/
cruft reports, three open design questions (§12 Q5, §17.4, §23.4 #2),
and two documentation-debt items (the taxonomy/glossary Captain Bob
flagged 2026-08-04, and re-measuring ACL-RWT DoE overhead now that
real compiler optimization is enabled).
Also includes BLOCK_MAP.md/artemis.img/amd64.csv regenerated by builds
during this session, and a qemu boot log/DoE run that weren't from any
command in this session -- kept per repo convention, logs are audit
artifacts, not deleted.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Extends the amd64 headless screendump verification (99999 SCROLL-BACK
injected over the serial chardev socket, captured via QEMU's HMP
screendump over its monitor socket) to aarch64 (-device ramfb) and
riscv64 (-device ramfb) -- same script shape, no GUI or physical
typing needed on either. Both show the same deep-POST recovery
(Init: Mama birth..., HADES DoE rows 58-73, boot banner) and the same
pre-existing 4.4q live-cursor-draw glitch already confirmed on amd64,
present identically -- boot-mode scrollback is real and reachable on
every architecture, not just amd64.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirms 4.4ac's boot-mode scrollback ring actually works: injected
'99999 SCROLL-BACK' over the serial chardev socket (same mechanism the
old DOE_INJECT automation used to drive EXEC-DOE -- SCROLL-BACK is a
plain FORTH word, and sk_repl_step()'s input comes from console_getc()
polling the UART), then captured the framebuffer with QEMU's own HMP
screendump command over its monitor socket. No GTK session or physical
typing needed, correcting this item's own earlier assumption that it
would.
Result (logs/screendump-4.4ac/amd64-scrollback.png) shows PARITY:
MAMA_INIT, Init: Mama birth OK, ACL: CAPSULE-BIRTH pinned STRICT,
Starting heartbeat..., and HADES DoE rows from steps 58-73 all on
screen at once -- genuinely early POST content, not just the pre-TTF
tail. Also confirms a small pre-existing 4.4q limitation (live cursor
draw corrupting a scrolled-back view) is unaffected by this item, not
a new regression.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
vt100.c: font_8x16/bitmap-mode boot output previously had no scrollback
at all -- g_shadow/g_ring were only allocated in vt100_enable_ttf(), so
POST/self-test/heartbeat text was gone the instant it scrolled off,
recoverable only from the serial log. Gives boot mode its own ring/
shadow pair (bitmap cell geometry, 4096-line capacity), frozen as a
snapshot the moment vt100_enable_ttf() switches to the TTF-geometry
pair, per the two-independent-rings design scoped with Captain Bob.
scrollback_line_at()/scrollback_redraw()/vt100_scroll_back() now walk
all four segments (boot ring, boot shadow, TTF ring, TTF shadow) as one
continuous history, so PgUp from the REPL reaches back through POST.
Three-arch QEMU boot + logs clean (amd64/aarch64/riscv64, no faults, no
dictionary/parity regressions). Visual verification that PgUp actually
recalls POST text still needs an interactive GTK screendump -- noted as
open in FABRIC.md, same pattern as 4.4ab's screendump.
Also includes BLOCK_MAP.md/artemis.img regenerated by these builds, and
the acceptance-boot logs (plus stray logs from an earlier QEMU-instance
collision during testing -- kept per repo convention, logs are audit
artifacts, not deleted).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The qemu target used to poll the serial log for ok>/zuse)ok>, then
unconditionally kill the VM (optionally injecting EXEC-DOE first) — a
DoE-campaign automation shape that also fired during plain interactive
use, cutting the session out from under you the moment the prompt
appeared. All three arch branches (amd64/aarch64/riscv64) now just run
qemu-system-* in the foreground and block until it's closed manually;
serial logging to logs/ and DoE CSV extraction on exit are unchanged.
Also includes BLOCK_MAP.md/artemis.img/amd64.csv regenerated by the
qemu-esp test run, and that run's log/CSV artifacts.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Brings in the cursor indicator + HB-ON/HB-OFF runtime DoE toggle work.
Diverged from master's own three-arch verification commit (0d8fff3,
logs only, no code overlap) since that verification was made directly
on master rather than merged back to stadium-step-one first.
Cursor (Captain Bob: "the only thing we need is a cursor"):
vt100_draw_cursor() draws a solid block at the terminal's current
position, called from repl.c after the prompt prints and after every
keystroke/backspace. vt100_erase_cursor() cleans up the one gap a static
cursor has -- Enter/newline moves away from the cursor cell without a
character draw ever overwriting it, which left a stray block behind
until this fix.
HB-ON/HB-OFF (Captain Bob: run a program with or without instrumentation
without rebuilding):
Converted per-tick DoE logging from a build-time flag (HEARTBEAT_DOE_LOG)
to a runtime one. doe_log_tick_row() now self-gates on g_doe_log_enabled
(default 1, matching the old default) instead of being compiled out
entirely; the call site in vm_runtime.c is unconditional. Two new FORTH
words, HB-ON and HB-OFF, flip the flag live. Removed the now-dead
HEARTBEAT_DOE_LOG plumbing: the Kconfig symbol, and the -D forwarding in
both LOADER_CFLAGS and KERNEL_CFLAGS.
Verified: three-arch clean QEMU boot + logs; dictionary word count 466
(463 baseline + ALT+TAB + HB-ON + HB-OFF, exactly the three words added
across this session); amd64 screendump confirms the cursor renders
correctly after real interactive typing.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Confirms all three architectures boot clean on master after the
fast-forward merge from stadium-step-one (af20efa), identical to the
behavior verified on that branch: UEFI -> POST -> Mama birth -> Hermes
self-test -> heartbeat -> ok>, no DoE/ECW noise.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
full-screen vt100 terminal
4.4v -- keyboard-to-REPL bridge, real and tested:
Refactored KEY-EVENT's per-arch translation logic (keyboard_words.c) into
a shared C function, sk_key_event_poll(), so the REPL bridge reuses item
4.3.5f's already-converged Linux-keycode-namespace event stream instead
of building separate amd64/aarch64/riscv64 tables. repl.c's sk_kbd_getc()
decodes the standard US-QWERTY printable range plus Enter/Backspace/Shift
against that stream; sk_readline() polls it as a second source alongside
console_getc(). Verified via QEMU monitor sendkey injection, and by
Captain Bob typing directly into the live QEMU window over real emulated
PS/2 hardware mid-session (1 1 + . -> 2 ok, then a clean BYE shutdown).
4.4ab -- simplify to a full-screen terminal:
Captain Bob's call, reverting the 640x480 CANVAS box + independent REPL
strip (4.4o/4.4t/4.4x/4.4z) in favor of the simplest shape: the entire
framebuffer is one vt100 terminal, g_vt.cols/rows = fb_width()/fb_height()
divided by cell size, no origin offset, no box, no strip, no border
drawing. The REPL prompt is just the terminal's last scrolling line.
Scrollback, TTF rendering, and SGR color are all box-agnostic and keep
working unmodified.
4.4r -- reframed as a text/graphics mode toggle:
"Hide/show the scroll box" stopped meaning anything once the box was
removed; the underlying need survives as a whole-screen mode switch.
vt100_toggle_graphics() is a two-state machine (VISIBLE/HIDDEN) -- hidden
mode stops the terminal from touching the framebuffer while its logical
state keeps advancing, so direct framebuffer/TTF-TEXT drawing can use the
whole screen; showing again wipes and reuses scrollback_redraw() to
restore the terminal exactly. Reachable two ways, one transition function:
physically via Alt+TAB (4.4y revised from Ctrl+TAB) and programmatically
via the new ALT+TAB FORTH word.
Verified: three-arch clean QEMU boot + logs; amd64 screendump confirms
full-width text with no box/strip artifacts.
Punch list §25 items 4.4v/4.4r/4.4ab complete; 4.4y revised.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
draw_box_border() (vt100.c) strokes four 1px edges around the 640x480
CANVAS box, reusing the same border-gray constant the REPL strip's
border lines use (renamed VT100_STRIP_BORDER_GRAY -> VT100_BORDER_GRAY
since it's now shared -- one pinned color decision, 4.4w, not two).
Called from erase_display()'s box-scoped branch so the border survives
every box clear (the initial one and any later ESC[2J), not just the
first.
Verified: three-arch clean QEMU boot + logs, amd64 screendump showing a
full rectangle outline around the box, visually distinct from the strip
below it.
Punch list §25 item 4.4z complete.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scope expanded from pure CANVAS-rectangle arithmetic (as originally
scoped) to also splitting the REPL prompt/input line out of the
scrollback box into an independent single-line strip, per Captain Bob's
explicit fold-in after the gap was reported (§25.0 rule 3) rather than
silently expanded.
vt100.c: VT100_BOX_ORIGIN_X/Y are no longer hardcoded per-arch literals --
both are now derived from fb_width()/fb_height() at vt100_enable_ttf()
time. New vt100_strip_draw() renders the bottom strip (gray border lines,
bright-white text) directly via the existing ttf_draw_glyph_cell()
rasterizer, independent of the box's own grid/cursor state. Border lines
are drawn after the glyph loop so an oversized cell can only be clipped
by them, never erase them.
console.c/console.h: console_fb_strip_draw() thin wrapper, matching the
existing console_fb_enable_ttf()/console_fb_scroll_*() pattern.
repl.c: builds a plain-text "[VMName] ok> <input>" mirror in
g_strip_prompt/strip_refresh(), refreshed on every keystroke (including
backspace) from sk_readline() -- already wired for item 4.4v, since
keyboard-typed characters will flow through the same console_getc() path
once that lands. Also widened sk_repl_step()/sk_repl_run()'s local input
buffer from a second, smaller 256-byte buffer to INPUT_BUFFER_SIZE
(1025), per 4.4w's decision.
Verified: three-arch clean QEMU boot + logs, amd64 screendump showing
the box and strip as two visually distinct regions with no visible
glyph/border clipping.
Punch list §25 item 4.4x complete.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
4.4w resolved: border gray = FB_ANSI_PALETTE[7] (0xAAAAAA); REPL input
line uses the full INPUT_BUFFER_SIZE=1025 buffer with left/right
horizontal scroll on a single line, not a smaller practical limit; the
15px/15px REPL-strip gaps and 8px box-to-strip gap from the second
mockup pass confirmed as-is. 4.4y resolved: toggle meta-key is Ctrl+TAB.
4.4u's own done-when (three open gaps resolved with Bob) is satisfied by
4.4w's resolution, so it's checked off too. Design-only, no code changes
in this commit -- 4.4x/4.4z pick up the drawing work these numbers
unblock.
Punch list §25 items 4.4u/4.4w/4.4y complete.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
vm_core.c: demote the per-word "ECW: w=... func=... 'NAME'" dispatch trace
from LOG_INFO to LOG_DEBUG. POST forces the logger to LOG_TEST for the
duration of the self-test run, and LOG_TEST includes LOG_INFO, so every
single word execution during POST was echoing this trace -- hundreds of
lines burying the actual module summaries and pass/fail tally. Still
available via --log-level=debug.
Combined with HEARTBEAT_DOE_LOG=0 (command-line Kconfig override, no
default change -- experiments/bare_metal/'s own DoE tooling still gets
HEARTBEAT_DOE_LOG=1 by default), all three architectures now boot clean:
UEFI -> POST summary -> Mama birth -> Hermes self-test -> heartbeat ->
ok>, with no [HADES][DOE] rows and no ECW flood. Verified by three-arch
QEMU boot; logs attached.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Default log level dropped from info to warn so the per-word ECW dispatch
trace doesn't flood REPL output after POST (--log-level=info/debug still
re-enables it). qemu target gains QEMU_DISPLAY (default gtk) so the
framebuffer window shows by default; serial log is tee'd live via
`tail -f` instead of dumped with `cat` at the end. Includes regenerated
BLOCK_MAP.md/amd64.csv/artemis.img and this morning's boot logs/DoE runs
from the sessions that produced this WIP.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Captain Bob flagged the document (7,300+ lines) has grown cluttered,
with punch-list items scattered throughout rather than collected in one
place. Recorded as the explicit first task for next session, before any
4.4-series work resumes: a full reorg pass collecting all punch-list
items to the end of the document, preserving every item's content
(including Done blockquote history) exactly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Captain Bob was explicit: don't consult docs/lithosananke/ROADMAP.md at
all anymore, for anything, not just the REPL/M8 section already marked
obsolete inside it. Added an explicit statement to the preamble so this
isn't scoped too narrowly by a future reader.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
The prior commit's resume note was a single blockquote blob with a
numbered list buried inside item 4.4v -- not the document's own
established format. Replaced with four real standalone items matching
every other entry in this section (own checkbox, description, Done
when, Refs): 4.4w (pin 4.4u's open numbers), 4.4x (redo 4.4n's CANVAS
math for the real strip height), 4.4y (decide the toggle meta-key),
4.4z (draw the scroll box's visible border). Updated 4.4r's dependency
list to include 4.4v and 4.4z.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
TAB was recorded too specifically in 4.4u/4.4v -- Bob deferred the exact
toggle key (TAB/Ctrl+TAB/other) to decide later; both items now say so
and 4.4v's interception scope is keyed off "whatever key gets chosen"
instead of hardcoded to TAB.
Added an explicit six-step breakpoint/punch list after 4.4v recording
the dependency-ordered resume sequence for tomorrow: pin 4.4u's open
numbers, redo 4.4n's CANVAS math, decide the toggle key, build 4.4v,
draw the scroll box's border, then build 4.4r.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
4.4u records the full bottom-up console layout Captain Bob walked through
tonight (REPL strip geometry, single-line horizontal-scroll input, drawn
scroll box, TAB-toggle, colors), including a mockup review round that
grew the REPL-strip gaps by 3px and added scroll-box content soft-wrap.
Flags that it revises 4.4m/4.4n's REPL-strip sizing without editing those
items in place. 4.4v scopes the keyboard-to-REPL bridge precisely: the
i8042/virtio_input interrupt-driven keyboard layer is real and working,
but nothing connects it to sk_readline() yet -- that gap is what's left,
not a from-scratch keyboard subsystem.
Neither item has code yet -- design/scope documentation only, per this
document's own discipline of capturing design before implementation.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
vt100's TTF-mode text grid now operates within the per-arch 640x480 box
(computed in 4.4o, pixel-verified in 4.4p) instead of the full framebuffer:
box-origin offset in px_of()/py_of(), box-derived cols/rows (53x20) set
before the 4.4q scrollback allocation depends on them, mode-aware
erase_display()/reverse-index fill, and a new box-scoped fb_scroll_rect()
alongside the existing whole-framebuffer fb_scroll_rows() (bitmap/boot mode
unaffected either way). Also clears the full framebuffer once at the
bitmap-to-TTF switch so leftover boot debris doesn't sit frozen outside the
box now that erase_display(2) is box-scoped afterward.
Three-arch QEMU boot + pixel-scanned screendumps confirm zero non-background
pixels land outside the box on amd64, aarch64, and riscv64.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scope decided with Captain Bob before implementation: ring buffer +
recall on today's full-screen vt100 grid, not also confining REPL text
to the 4.4o 640x480 box (that confinement stays open as its own future
item, not a third silent deferral). No keyboard input path exists yet
(M8 unstarted), so the trigger is two new FORTH words, SCROLL-BACK
( n -- ) / SCROLL-FWD ( n -- ), exercised via serial injection.
vt100.c gains a text-only 1000-line ring buffer (kmalloc'd, tens of KB
-- not pixel snapshots, which would be ~1000x larger for no benefit)
plus a shadow buffer mirroring the current screen. scroll_up() now
pushes evicted rows into the ring before the pixel scroll. History is
one continuous sequence (ring then shadow); scrolling always redraws
from that sequence -- no separate pixel-scroll path for scrollback,
decided up front to avoid retrofitting later.
New src/word_source/scroll_words.c (Module 31), thin wrappers over
console_fb_scroll_back()/_fwd() -> vt100_scroll_back()/_fwd(). Bug
caught during live testing: both words initially used an off-by-one
underflow check (dsp < 1) copied from a different, older dsp
convention elsewhere in this codebase; vm_pop() (which these words
actually call) uses dsp as a 0-based top-of-stack index, so the check
rejected every legitimate single-argument call. Fixed by removing the
separate precheck and relying on vm_pop()'s own guard.
Live-verified on all three architectures (exceeds this item's
amd64-minimum bar): generated 50+ lines via a FORTH loop, confirmed
SCROLL-BACK recovers correctly older content, and on amd64 confirmed
SCROLL-FWD returns to genuinely live state (not a frozen snapshot) by
showing the injected commands' own echo. Known limitation confirmed by
direct pixel measurement: redrawn lines lose their original SGR color
(not stored per-cell) -- text recovers exactly, color does not.
Three-arch verified: Failed: 0, dict-hashes identical across all
three (values changed correctly from prior items -- two new words
were added).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Verification only, no permanent code (matches this item's own title
and 4.4i's same-shape precedent -- 4.4l-4.4o were geometry decisions,
not drawing code). One-shot diagnostic probe in sk_repl() drew 4.4o's
box outline plus a strip-top marker, then was reverted after capture
per this document's write/run-once/capture/revert discipline.
Measured pixel bounds directly from each screendump (not eyeballed)
and confirmed exact matches against 4.4o's computed coordinates on
all three architectures: amd64 box x:[320,959] y:[104,583], strip
marker y:704; aarch64/riscv64 box x:[80,719] y:[4,483], strip marker
y:504.
Observed, not fixed here: on the small architectures the REPL banner
text visibly overlaps the box's top edge, because vt100's cursor grid
still spans the whole screen rather than being confined to the 96px
strip -- a real gap between the landed REPL path and the mockup,
belongs to 4.4q/4.4r's wiring work, not this item.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pure geometry, no code change. Blocking question checked first: both
4.4n CANVAS heights (688px amd64, 488px aarch64/riscv64) clear the
480px minimum, so no architecture fails the fit.
amd64: box raster top-left (320,104), 320px side margins, 104px
top/bottom margins.
aarch64/riscv64: box raster top-left (80,4), 80px side margins, 4px
top/bottom margins (confirms 4.4n's 16px gap choice lands exactly
where predicted).
Also translated both into TTF-TEXT's Cartesian bottom-left-origin
convention for whatever later item issues the actual TTF-TEXT calls:
amd64 bottom-left (320,216), aarch64/riscv64 bottom-left (80,116).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Pure geometry, no code change. Gap chosen at 16px, within 4.4m's
flagged <=24px ceiling for the 600px-tall architectures, leaving 8px
of margin rather than cutting to the exact limit.
amd64 (1280x800): CANVAS top-left (0,0), 1280x688, strip at (0,704)
96px tall.
aarch64/riscv64 (800x600): CANVAS top-left (0,0), 800x488, strip at
(0,504) 96px tall.
Checked against 4.4o's 480px scroll-box requirement: 488px clears it
with exactly the 8px margin the gap choice was picked to preserve.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Design decision, no code change (4.4j's constants are already the
final values). Text size 20px (unchanged from 4.4j), 4 visible lines,
strip height 96px (4 x the existing 24px cell height, no extra
padding -- leading is already baked into that cell height).
Flagged the constraint this feeds into before deciding: 4.4o needs a
640x480 scroll-box to fit inside CANVAS on the 600px-tall
aarch64/riscv64 screens, which only leaves 24px of slack over the
480px minimum once the 96px strip is subtracted. 4.4n's gap choice
must stay <=24px on those architectures or 4.4o's fit check fails --
recorded now so it's not a surprise two items later.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Investigation only, no code change. Answered incidentally by 4.4k's
screendumps rather than requiring a separate run: amd64 1280x800
(QEMU/OVMF GOP default), aarch64 and riscv64 both 800x600 (-device
ramfb's default). Two different resolutions across the fleet, not
three uniform ones -- flagged for 4.4n/4.4o's CANVAS geometry work,
which needs to branch per architecture rather than assume one shared
screen size.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
No code changes -- investigation/verification only. Injected a raw SGR
escape sequence over the serial socket at the ok> prompt (composed from
EMIT + ." since no FORTH word emits a literal ESC byte + text
directly): 27 EMIT ." [95mCOLOR-TEST" 27 EMIT ." [39m", bright magenta
(the p>=90&&p<=97 branch in apply_sgr()), a third SGR code path
distinct from 4.4h's already-verified truecolor prompt.
Screendump on all three architectures confirms COLOR-TEST renders in
bright magenta via the 4.4j-retargeted TTF draw call, proving the SGR
parser and the new glyph backend work together end-to-end, not just
independently. First screendump evidence ever captured for
aarch64/riscv64 in this document -- every prior screendump item was
amd64-only with that gap explicitly accepted; incidentally closed here
via the QEMU monitor screendump command against each arch's own
display device (ramfb for aarch64/riscv64). Both confirmed booting at
800x600.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
font_8x16.c keeps rendering everything through and including POST;
TTF-TEXT's rasterizer takes over at the interactive REPL boundary
(sk_repl()) via a new runtime mode switch, vt100_enable_ttf()
(console_fb_enable_ttf() wrapper), not a compile-time swap -- both
backends coexist in the same binary since boot/POST must stay
font_8x16.c per this item's own done-when.
TTF-TEXT (the FORTH word) isn't directly callable from vt100.c -- VM
stack arguments, different call shape than a one-glyph cell draw. Used
hal/ttf.c's VM-independent primitives directly instead (same rasterizer
TTF-TEXT itself calls underneath), added as a native C helper in
vt100.c. Lazily loads fonts:JetBrainsMono-Regular.ttf and kmallocs a
96-slot raster cache (covers all 95 printable ASCII, no eviction
thrash) on first switch.
Cell geometry changes at the switch (mode-aware cell_w()/cell_h()):
provisional 12x24 TTF cell (600/1000em * 20px = 12px exactly, using
4.4i's confirmed-uniform hmtx advance width) vs font_8x16's fixed 8x16
-- cols/rows re-derived and screen cleared at the switch point, same as
vt100_init() itself does. Final REPL text size is 4.4m's decision, not
this item's.
Also fixes the second call site 4.4i flagged: erase_line_range() now
uses one fb_fill_rect() instead of a per-cell font_8x16-specific blank
glyph draw, consistent with erase_display(2)'s full-screen case.
Three-arch verified: amd64/aarch64/riscv64 all reach ok>, POST
Failed: 0, identical dict-hashes. amd64 screendump shows real
proportional JetBrains Mono letterforms on the REPL tail, visibly
distinct from every prior font_8x16 screenshot.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Investigation only, no code change. draw_cursor_glyph() itself is
single-call-site as claimed, but found two things the item's original
framing missed:
1. A second, independent fb_draw_glyph() call site in
erase_line_range() (partial-line/partial-screen erase), inconsistent
with full-screen erase which already uses a pixel fb_fill_rect().
2. fb_cell_w()/fb_cell_h() are hardcoded 8*scale/16*scale literals
matching font_8x16.c specifically, not derived from any generic
font-metric abstraction; UNDERLINE_ROW hardcodes "row 14 of 16" of
that same fixed grid.
So 4.4j has three things to retarget/generalize, not one. Font data
verified directly from JetBrainsMono-Regular.ttf's hmtx table (parsed
by hand, no fontTools available): all 95 printable ASCII glyphs share
one advance width (600/1000 em units) -- genuinely monospace for the
glyphs in use, not just by filename.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
console.c's emit_prefix() now wraps [VMName] (brackets included) in
FABRIC.md 4.4's locked orange (0xFFA500), and repl.c's two "ok> "
call sites send FABRIC.md 4.4's locked cyan (0x55FFFF), both as real
SGR escape sequences through the existing font_8x16.c/vt100.c pipeline
-- 4.4b already established this needs no dependency on TTF-TEXT/4.4j.
Sent through both raw_putc() (serial) and vt100_putc() (framebuffer),
matching the existing dual-path pattern, so an ANSI-aware serial
terminal renders the same colors as the framebuffer.
Three-arch verified: amd64/aarch64/riscv64 all reach ok>, POST
Failed: 0, identical dict-hashes. Color applies correctly to any VM
name (confirmed via the [Hermes]-prefixed PARITY:BIRTH line in all
three logs, not just [Hera]). amd64 screendump confirms the rendered
colors directly.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Moves the console_fb_init() call in kernel_main.c from after
capsule_birth_mama() to before it, so the fleet-birth/self-test
transcript (Hermes x2, Artemis births, Stadium self-tests -- currently
serial-only) is also framebuffer-visible, not just the small post-birth
tail.
4.5f's -O2 experiment already showed this doesn't hang under
optimization, just costs roughly 12x more boot-time heartbeat ticks
(one-shot, paid only during fleet birth, never repeated at runtime).
Captain Bob's call: worth it, since the serial log was never the
problem -- this is about the same transcript also reaching a real
screen.
Three-arch verified: amd64/aarch64/riscv64 all reach ok>, POST
Failed: 0, identical dict-hashes across all three. amd64 screendump
confirms the framebuffer now carries the full transcript.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Uncommitted experiment (code reverted after capture, per Captain Bob):
moved console_fb_init() before capsule_birth_mama() in kernel_main.c,
amd64 only. At -O2 the boot completes cleanly and reaches ok> well
inside a 300s bound, versus the indefinite stall previously seen at
-O0. Real cost: ~12x more heartbeat ticks during boot from the extra
framebuffer scroll volume -- not free, but not a hang.
Screendump confirms the actual point of 4.4g: the framebuffer now
carries the full HADES/ECW/Stadium/self-test transcript, not just the
small post-birth tail.
This informs 4.4g's open reorder decision, it doesn't make it -- 4.5f
stays unchecked (single-arch feasibility check only, not the three-arch
acceptance pass its own done-when requires).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
FABRIC.md item 4.5d Finding 4: lidt()'s inline asm used a register-only
("r") constraint on the idtr pointer, never telling GCC the asm
dereferences the pointee. At -O2 this let the compiler treat the
256-entry idt[] population loop and idtr_desc's field writes as dead
stores and eliminate them entirely, loading IDTR from uninitialized
stack instead of the real table -- a #GP on the first APIC timer tick
that happened to land on garbage. Same bug class as the earlier
muldiv64 fix (b43e51a): an inline-asm constraint too weak for what the
asm actually touches, invisible at -O0, live at -O2.
Fixed by switching to a memory operand ("m"(*idtr_desc)), matching how
Linux's own load_idt() is written. aarch64/riscv64 checked for the same
pattern -- neither has it, both install their vector/trap tables
entirely in hand-written .S.
Verified: all three architectures boot clean to ok>, POST Failed: 0,
identical dict-hashes across all three under -O2, zero new warnings
vs an -O0 baseline (amd64 3040/3040, aarch64 3041/3041 serial,
riscv64 3037/3037). -O2/-U_FORTIFY_SOURCE landed permanently in
COMMON_CFLAGS.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Computed the loader's true runtime relocation delta to correctly correlate
fault addresses (the running image is starkernel_loader.efi, a relocated
PE, not the separately-linked starkernel_kernel.elf assumed at first).
Fault RIP decodes to log_message()'s entry -- coincidental, not causal,
since it's called on nearly every HADES dispatch during word registration.
Used QEMU's monitor for -d exec,int tracing. Late-start tracing (stop right
before the danger zone to keep trace size down) failed twice -- the window
between a detectable checkpoint and the crash is shorter than host-side
reaction latency. Fell back to full-boot tracing from -S (~2.7GB per
attempt, not committed). That trace shows an unremarkable, normal-looking
repeating three-block loop immediately before the fault, then "Servicing
hardware INT=0x20" (APIC_TIMER_VECTOR) with IDT already showing limit=0 at
that instant.
Ruled out a second illegitimate lidt call (only one call site exists
anywhere, one-time M4 boot setup; searched the trace for any later
execution of that address range and found none). Not yet established: the
actual corrupting write. Documented two remaining explanations (earlier
silent corruption vs. a genuine TCG artifact) and that pinpointing the
exact instruction needs GDB-level single-stepping, a bigger tooling step
than attempted this session.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Used QEMU-side instrumentation (-d int tracing, plus a working chardev-based
monitor -- the older bareword -monitor syntax silently fails on QEMU
10.2.1) as directed. The PM Timer "stall" was never a hang: it was a #DE
divide error cascading to a triple fault, which -no-reboot converts into a
silent clean QEMU exit -- indistinguishable from a hang without tracing,
and why arch_relax() (solving a hang that didn't exist) had no effect.
Root cause and fix documented in full (also see commit b43e51a).
With the fix, amd64 boot at -O2 now proceeds far past the original stall --
through capsule birth and into Hermes's word registration -- before hitting
a second, different fault (Finding 4): a direct #GP at IDT index 32
(APIC_TIMER_VECTOR), not yet root-caused. Same silent-triple-fault-exit
shape, so this was very likely bundled into "the hang" before tracing
distinguished the two separate bugs.
-O2 reverted again (uncommitted); muldiv64's fix is kept, real and
independently verified at unchanged -O0 on all three architectures.
Routine three-arch artifacts from this session's verification runs
included.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
FABRIC.md item 4.5d: root-caused the -O2 boot stall via QEMU '-d int'
tracing -- not a hang. It's a genuine divide-error (#DE) cascading through
a double fault into a triple fault, which -no-reboot converts into a
silent, clean QEMU exit (indistinguishable from a hang without tracing).
muldiv64()'s inline asm declared RDX as a plain output ("=d"(hi)), which
only tells GCC "I want RDX's value after this block" -- nothing told it
that mulq writes RDX *before* divq needs to read a *different* value (the
divisor c) out of it. Nothing stopped the register allocator from placing
c itself in RDX, which mulq then overwrites with the product's high 64
bits before divq ever reads it. Confirmed via the fault's register state:
RAX=0xe8d4a51000 (=1000*1e9 exactly, product fits in the low 64 bits, so
mulq's high-word output is 0) -- if c got allocated to RDX, divq then
divides by that corrupted 0, exactly matching #DE. Worked by accident at
-O0 (different, more conservative allocation); -O2 actually hit it.
A unsigned __int128 rewrite was tried first but needs libgcc's __udivti3
for the general 128-bit case, undefined in this freestanding build -- not
viable, same class of problem as the earlier putc/getc finding. Fixed
instead by declaring rdx a pure clobber rather than an output, the same
pattern the Linux kernel's own mul_u64_u64_div_u64 uses -- a clobber tells
GCC the register is used internally for the whole block and must never be
allocated to any operand, which is the guarantee the previous constraint
list was missing.
Also added a defensive end_tsc<start_tsc guard in the caller
(calibrate_tsc_with_pmtimer): this file's own comments already flag TSC
non-monotonicity as a real risk under TCG hypervisor mode, and an
underflowed delta_tsc would hit the same class of quotient-overflow #DE.
Not the bug that was found, but a real latent risk given what this
function's own documentation already says about the environment.
Verified: with -O2 (uncommitted, not yet reintroduced) this specific stall
is gone -- boot now proceeds far past this point, through capsule birth
and into Hermes's word registration, before hitting a second, different,
not-yet-root-caused fault (write-up follows). Three-arch acceptance boot
clean at unchanged -O0 (this fix doesn't change -O0 behavior, only
prevents UB that only manifested under optimization).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Follow-up probing narrowed Finding 3's behavior (inconsistent stall point
run-to-run: sometimes reaches iters=1000/delta=951 before stopping,
sometimes never gets past iters=0 even after a 400-second bounded wait) but
didn't pin the mechanism.
Tested the strongest available hypothesis: single-threaded TCG scheduling
starvation from an -O2-tightened spin loop, based on a real precedent --
calibrate_apic_timer() (apic.c) already calls arch_relax() every iteration
of its own spin-wait; calibrate_tsc_with_pmtimer() never had it. Added the
same call, matching that precedent exactly. Result: no change, same exact
stall point on a fresh 60-second bounded wait. Reverted.
Stopping here per this document's own §25.0 rule 5 -- root-causing further
needs either deeper TCG/QEMU internals knowledge or a different diagnostic
approach (host-side instrumentation) than another guess-and-check pass.
Makefile.starkernel and timer.c both back to committed -O0 state, confirmed
booting clean.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Finding 2 (putc/getc) fixed via shim.c backfill (previous commit). With
both Finding 1 (-U_FORTIFY_SOURCE) and Finding 2 fixed, amd64 links clean
at -O2 -- rigorously confirmed zero new warnings by diffing normalized
warning text between -O0 and -O2 builds (empty diff, ~3040 pre-existing
warnings in vendored code unchanged).
New Finding 3, not fixed, now blocking: amd64 boot stalls for minutes
inside calibrate_tsc_with_pmtimer() (timer.c:560-589) at -O2 -- a
mainline busy-wait loop reading the ACPI PM Timer via genuinely-volatile
inl() until 1000 real ticks elapse, bounded by a 5M-iteration timeout. A
one-shot diagnostic probe (reverted after capture) caught exactly one
checkpoint in 55 seconds of observation, then nothing -- not yet
root-caused whether this is the 5M-iteration timeout itself taking several
minutes to exhaust under -O2, or a genuine non-terminating condition.
Makefile.starkernel reverted to -O0 again; nothing broken landed in
history. Routine artifacts from this session's boot attempts included.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
FABRIC.md item 4.5d: GCC's -O2 folds putchar(c)/fputc(c,stdout) call sites
into putc(c,stdout), and getchar() into getc(stdin) -- neither symbol was
ever needed at -O0 because that fold pass is inactive there. Both are thin
wrappers reusing the existing putchar()/getchar() implementations exactly
(putc -> console_putc via putchar; getc -> the existing "no stdin in
kernel" -1/EOF stub via getchar), not new behavior.
Harmless at the current -O0 build (nothing calls these symbols directly at
-O0); this is prep for enabling optimization, verified against a clean
amd64 boot at unchanged -O0.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Attempting -O2 on amd64 surfaced two link-time findings, neither of which
was warnings the way 4.5d's own text anticipated:
1. Fixed: -O2 broke the link with __printf_chk/__memset_chk/__fread_chk/
__snprintf_chk undefined references across five vendored word_source
files. Root cause: those files unconditionally #include real glibc
headers (no __STARKERNEL__ guard), which under Ubuntu's default
_FORTIFY_SOURCE and -O2's __OPTIMIZE__ rewrite printf/etc. call sites to
_chk variants that shim.c's freestanding backfill never provided --
worked by accident at -O0 where the macros stay dormant. Fixed by adding
-U_FORTIFY_SOURCE, the standard fix for this exact situation in
freestanding/kernel builds.
2. Not fixed, blocking: with fortification disabled, plain putc/getc turned
up as genuinely undefined -- shim.c never backfilled those, only printf/
snprintf. -O0 doesn't need these symbols for reasons not fully
established (GCC's -O2 function-splitting visibly clones at least one
call site); this needs a real design decision (backfill shim.c vs. guard
the call sites out of __STARKERNEL__ builds), not a mechanical flag.
Makefile.starkernel reverted to the committed -O0 state (the -O2 change was
uncommitted, so a plain git restore -- nothing broken landed in history).
4.5d stays open pending Finding 2's resolution.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
-O2 over -O1/-Og: no optimization level has a verified track record here,
so pick the one that actually solves the motivating problem (4.4g's
fb_scroll_rows() stall needs real loop batching -Og doesn't reliably do)
and matches what both sibling Makefiles already default to.
No LTO yet: deliberately deferred, not rejected -- stacking a second risky
change (LTO, which CLAUDE.md already documents has broken this exact
codebase before) on top of a first-ever optimization pass would make any
4.5e failure ambiguous between two causes. Isolate them.
No -DNDEBUG: checked, not assumed. Zero runtime assert() calls exist
anywhere in the kernel build -- the two "assert(" hits found are a
_Static_assert pair (compile-time, ungated regardless) and a comment. The
freestanding assert.h shim is itself unconditional and ignores NDEBUG too.
Nothing for the flag to affect; copying it by habit would have been exactly
the cargo-culting this item warned against.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Marks 4.5b done with the verification record (three log dirs, all reaching
[Hera] ok> at unchanged -O0). Routine artifacts from this session's
three-arch runs: capsules/BLOCK_MAP.md, disk/artemis.img, DOE CSVs, QEMU
serial logs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Written directly in ISR context on all three architectures
(heartbeat_tick(), heartbeat.c:163) and read directly by mainline
(heartbeat_ticks(), including the busy-wait at kernel_main.c:880) without
being volatile -- worked by accident at -O0, would be a real bug once the
kernel builds with optimization (item 4.5). Every other field in this
struct is mainline-only (heartbeat_service()'s deferred window/variance/
trust processing), so only this one field needed the qualifier.
Punch list item 4.5b complete.
Three-arch acceptance boot clean at unchanged -O0 (no behavior change
intended yet -- this is prep for enabling optimization, not the switch
itself).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Read every interrupt/exception vector handler on amd64, aarch64, and
riscv64, and every global or static variable each one touches directly or
through a called function, checked against its volatile declaration.
One confirmed hazard, matching what 4.5 already reported: TimeTrustState.
ticks is written in ISR context on all three architectures and read by
mainline (including a busy-wait) without being volatile.
Everything else checked out for one of three reasons, each verified by
reading the actual read/write sites rather than assumed: already correctly
volatile (g_sk_fault_word, g_spurious_count, g_plic_claim_count, the
heartbeat.c top/bottom-half handoff, i8042.c's ring buffer, virtio_input.c's
diagnostic counters); write-once during init then single-context for the
rest of boot, so never actually concurrent (each arch's timer-calibration
state, virtio_input.c's device-routing globals); or ISR-reachable only on
the fatal exception path, which halts the core permanently afterward so
there's no return to mainline to race with (console/framebuffer state).
Full findings recorded in FABRIC.md as 4.5a's inventory. No code changed --
investigation only, per the item's own scope.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Captain Bob 2026-08-11: no other item gets worked until 4.5a-4.5f (the
-O0-kernel fix) are done.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Captain Bob: the -O0-kernel finding (with its 4.5a-4.5f scoping) takes
priority over Artemis-last, so it moves up to 4.5 and Artemis moves down
to 4.6. Updated the three other places in this document that referenced
"the 4.5 Artemis boundary" by number to point at 4.6 instead, with a note
on the renumbering history for anyone reading those passages later.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Scoping only, per Captain Bob's explicit instruction -- nothing implemented.
Breaks the -O0-kernel finding into six sequential, individually-committable
items matching this document's §25.0 rule-1 discipline (same treatment 4.4
itself got before 4.4a onward):
4.6a full ISR-global volatile audit (investigation only), 4.6b fix whatever
4.6a finds (starting from the one already-confirmed TimeTrustState.ticks
hazard), 4.6c decide the actual -O flags and record the reasoning before
touching the Makefile, 4.6d apply them and get a clean three-arch build,
4.6e the real three-arch acceptance boot (a kernel that has only ever run
at -O0 has no track record at any other level), 4.6f retry 4.4g's reorder
now that optimization exists to make it viable.
Also flagged, not scoped: the ACL-RWT DoE campaign's overhead numbers were
all measured at -O0; nobody has asked whether they still hold once the
kernel builds differently.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Attempted 4.4g's console_fb_init() reorder twice this session: once bare,
once with the fb_scroll_rows() volatile fix applied. Both stalled boot
indefinitely (12,000+ lines logged, still running after 3 minutes vs. a
normal few-second boot) instead of completing. Root cause traced past the
scroll fix to Makefile.starkernel building the kernel at -O0 -- no
optimization flag has ever been configured there, verified against the
file's full git history (17 commits, only ever one unrelated host-tool -O2
line). Both the kernel's own vendored hosted Makefile and the standalone
StarForth repo's Makefile default to -O2 (up to -O3/-flto on faster
targets); Makefile.starkernel was written fresh for the bare-metal target
and never got that ladder.
Documented as new item 4.6: enabling optimization is not a safe drop-in
change on its own. Found one confirmed, isolated correctness hazard first --
TimeTrustState.ticks (timer.h:90) is written directly in ISR context on all
three architectures (heartbeat.c:163) and read directly by mainline
(heartbeat.c:200-202, including a busy-wait in kernel_main.c:880) without
being volatile, unlike every other ISR-shared global checked
(g_spurious_count, g_plic_claim_count, g_pending_counter/g_pending_valid/
g_adaptive_period_ns are all correctly volatile already). Reverted the
console_fb_init() reorder itself (uncommitted, so a plain git restore) --
4.4g stays open pending 4.6.
Both the reorder attempts' logs (stalled, never reached ok>) and this
session's routine three-arch artifacts (capsules/BLOCK_MAP.md, disk/
artemis.img, DOE CSV) are committed as audit trail per CLAUDE.md.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Found while investigating item 4.4g's boot stall (moving console_fb_init()
earlier caused a >3-minute hang scrolling the fleet-birth transcript). The
GOP framebuffer is mapped write-back RAM (vmm.c:350-363), not
cache-disabled MMIO with side effects, so there's no correctness reason for
the scroll/fill loops to force one un-batchable volatile access per pixel.
g_fb.base stays volatile for other call sites; this function now casts to a
plain pointer for its bulk copy only.
Confirmed this alone does not fix the 4.4g stall -- the kernel builds at
-O0 (no optimization flag anywhere in Makefile.starkernel), so nothing
here gets vectorized regardless of the qualifier. That's now documented as
its own item, FABRIC.md 4.6. Keeping this fix regardless: it's correct on
its own terms independent of 4.4g's outcome.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>