Files
LithosAnanake/disk
Robert Allan JamesandClaude Sonnet 5 065ab50240 FABRIC.md: item 4.5d Finding 4 -- deep-traced, not yet root-caused
Computed the loader's true runtime relocation delta to correctly correlate
fault addresses (the running image is starkernel_loader.efi, a relocated
PE, not the separately-linked starkernel_kernel.elf assumed at first).
Fault RIP decodes to log_message()'s entry -- coincidental, not causal,
since it's called on nearly every HADES dispatch during word registration.

Used QEMU's monitor for -d exec,int tracing. Late-start tracing (stop right
before the danger zone to keep trace size down) failed twice -- the window
between a detectable checkpoint and the crash is shorter than host-side
reaction latency. Fell back to full-boot tracing from -S (~2.7GB per
attempt, not committed). That trace shows an unremarkable, normal-looking
repeating three-block loop immediately before the fault, then "Servicing
hardware INT=0x20" (APIC_TIMER_VECTOR) with IDT already showing limit=0 at
that instant.

Ruled out a second illegitimate lidt call (only one call site exists
anywhere, one-time M4 boot setup; searched the trace for any later
execution of that address range and found none). Not yet established: the
actual corrupting write. Documented two remaining explanations (earlier
silent corruption vs. a genuine TCG artifact) and that pinpointing the
exact instruction needs GDB-level single-stepping, a bigger tooling step
than attempted this session.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-08-11 12:40:30 -04:00
..

disk/

QEMU disk images used for Artemis (block-storage VM) persistence testing. Mounted via Makefile.starkernel's ARTDISK variable (default disk/artemis.img) as a virtio-blk-pci device on all three architectures' kernel QEMU boots — not scripts/rundisk.sh, which targets a separate, currently-unused disks/ (plural) directory for the hosted VM's --disk-img= flag instead.

  • artemis.img — standard Artemis persistence test disk. Reformatted 2026-08-02: this image had been stuck in a corrupted state (valid LithosAnanke magic header, but data not matching what ART-READ-TEST expects) since before this repo's own git history begins (git log shows it already broken at the initial commit, carried over from the pre-split monorepo). Chasing down the resulting persistent FAIL: persist-read traced to the data, not the code — the write→reboot→read round trip works correctly on a fresh image (see artemis-debug-roundtrip.img below). Reformatted by blanking the file and letting a normal boot format+write-test it; verified PASS: persist-read on amd64, aarch64, and riscv64 against the same image afterward (cross-arch resume, matching the arch-neutral on-disk format .claude/ARTEMIS.md specifies).
  • artemis-debug-roundtrip.img — round-trip regression fixture created during that investigation. Known-good: format → self-test → write-test → reboot → resume → PASS: persist-read, confirmed 3 times in a row. Keep this in a passing state; if a future change breaks it, that's a real regression, not a stale-fixture artifact like artemis.img was.
  • artemis-persist-test.img — persistence round-trip test image (pre-existing; history/state not re-verified during the above investigation).
  • artemis-poison.img — a separate test image (exact scenario not documented elsewhere in the repo as of this writing; name suggests an adversarial/corruption test, not confirmed).
  • artemis-unrecognized-test.img — exercises ART-HALT-UNRECOG (.claude/ARTEMIS.md acceptance criterion #6). Regenerated 2026-08-02: the previous copy of this file had itself been silently reformatted by a since-fixed bug in the generic block subsystem (src/block_subsystem.c) — it carried a valid low-level 'STFR'/v2 header despite being meant to represent foreign disk content, direct forensic evidence of the bug described in .claude/ARTEMIS.md's Build Status item 6. Regenerated as 30MB of a repeating POISON-UNRECOGNIZED-DISK-TEST-FIXTURE--NOT-BLANK-NOT-STFR-NOT-ARTEMIS-- ASCII pattern — deliberately neither blank, nor the block subsystem's own 'STFR' magic, nor Artemis's "ARTEMIS\0" marker. Verified on amd64 and riscv64 post-fix: boot correctly halts (ARTEMIS HALT: unrecognized disk content) and the file's sha256 is now byte-for-byte identical before and after boot. Keep this fixture in this poisoned state — if a future change makes its sha256 change across a boot, that is exactly the regression this fixture exists to catch.

Note on incidental header churn: artemis.img picks up a few changed header bytes on every ordinary boot even though no user data changes — blk_subsys_attach_device() always records a fresh mounted_time on a successfully recognized disk, which gets flushed at shutdown. This is expected bookkeeping, not a bug; revert it before committing rather than carrying timestamp noise in git history.

These are regenerable QEMU raw disk images, not source — see .claude/ARTEMIS.md for the storage model they exercise.