Files
LithosAnanake/disk/README.md
T
Robert Allan JamesandClaude Sonnet 5 36d832ff47 Single-block relocation: RELOCATE-BLOCK, resolve_lbn(), persisted exception table
Implements the full design from the prior commit in one pass. resolve_lbn()
is the single choke point threaded through the ten public LBN-consuming
entry points (blk_get_buffer, blk_update, blk_flush, blk_is_allocated,
blk_mark_allocated, blk_mark_free, blk_is_valid, blk_get_meta, blk_set_meta,
plus blk_get_empty_buffer covered via delegation) -- an LBN->LBN redirect,
not a new storage allocator, since the LBN space is already unified across
every attached blkio_dev backend. VM window cache staleness across a
relocation reuses the existing blk_vm_check_epoch() mechanism from
Milestone 2h's hot-detach fix for free -- g.epoch bumps on relocation too.

Persistence lands in the same pass: two new uint32_t fields
(reloc_start/reloc_devblocks) appended after hdr_crc in blk_volume_meta_t,
carved from existing padding without moving any earlier field's byte
offset -- an old formatted volume's zeroed padding reads back as
reloc_devblocks=0 ("no reloc capacity"), gracefully, not a format-breaking
change. compute_totals_from_B() generalized to account for the new
reserved region. reloc_flush_to_disk()/reloc_load_from_disk() mirror the
BAM I/O functions' own absolute-devblock-addressing shape; the persisted
copy's owner is first_disk_slot() (already existed, already used for this
exact "which device is canonical" question by blk_get_volume_meta()).
blk_subsys_relocate_block() is a mechanical primitive only -- copies
content (staged through a local buffer, since obtaining the target's
blk_get_buffer() result can evict and invalidate the source's cache
pointer if they share a device), frees the source BAM entry, appends the
exception entry, bumps the epoch, flushes to disk. RELOCATE-BLOCK exposes
it to FORTH, no policy of its own (ACL's job, per this session's direction).

A first live-test attempt gave a false negative against disk/artemis.img
(predates reloc capacity, so relocation only ever existed in memory that
boot) -- traced to the test's own setup before being mistaken for a bug,
then re-verified correctly against a fresh volume (new fixture,
disk/artemis-reloc-test.img): relocated a RAMDRIVE block to the fresh
disk, confirmed live resolution through the redirect, then confirmed both
the redirect and the relocated content survived an abrupt QEMU kill and
full reboot. Also fixed three lingering "glibc" doc-comment
misattributions from Milestone 2h (the actual allocator is this kernel's
own kmalloc) that survived an earlier FABRIC-2.md-only correction. All
three architectures re-verified clean.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CXjAPTEKrgY2Mrk25KoLDn
2026-08-25 19:14:01 -04:00

95 lines
5.8 KiB
Markdown

# disk/
QEMU disk images used for Artemis (block-storage VM) persistence testing.
Mounted via `Makefile.starkernel`'s `ARTDISK` variable (default
`disk/artemis.img`) as a `virtio-blk-pci` device on all three
architectures' kernel QEMU boots — not `scripts/rundisk.sh`, which targets
a separate, currently-unused `disks/` (plural) directory for the *hosted*
VM's `--disk-img=` flag instead.
- `artemis.img` — standard Artemis persistence test disk. **Reformatted
2026-08-02**: this image had been stuck in a corrupted state (valid
`LithosAnanke` magic header, but data not matching what
`ART-READ-TEST` expects) since before this repo's own git history
begins (`git log` shows it already broken at the initial commit,
carried over from the pre-split monorepo). Chasing down the resulting
persistent `FAIL: persist-read` traced to the *data*, not the code —
the write→reboot→read round trip works correctly on a fresh image
(see `artemis-debug-roundtrip.img` below). Reformatted by blanking the
file and letting a normal boot format+write-test it; verified
`PASS: persist-read` on amd64, aarch64, and riscv64 against the same
image afterward (cross-arch resume, matching the arch-neutral on-disk
format `.claude/ARTEMIS.md` specifies).
- `artemis-debug-roundtrip.img` — round-trip regression fixture created
during that investigation. Known-good: format → self-test → write-test
→ reboot → resume → `PASS: persist-read`, confirmed 3 times in a row.
Keep this in a passing state; if a future change breaks it, that's a
real regression, not a stale-fixture artifact like `artemis.img` was.
- `artemis-persist-test.img` — persistence round-trip test image
(pre-existing; history/state not re-verified during the above
investigation).
- `artemis-poison.img` — a separate test image (exact scenario not
documented elsewhere in the repo as of this writing; name suggests an
adversarial/corruption test, not confirmed).
- `artemis-unrecognized-test.img` — exercises `ART-HALT-UNRECOG`
(`.claude/ARTEMIS.md` acceptance criterion #6). **Regenerated
2026-08-02**: the previous copy of this file had itself been silently
reformatted by a since-fixed bug in the *generic* block subsystem
(`src/block_subsystem.c`) — it carried a valid low-level `'STFR'`/v2
header despite being meant to represent foreign disk content, direct
forensic evidence of the bug described in `.claude/ARTEMIS.md`'s
Build Status item 6. Regenerated as 30MB of a repeating
`POISON-UNRECOGNIZED-DISK-TEST-FIXTURE--NOT-BLANK-NOT-STFR-NOT-ARTEMIS--`
ASCII pattern — deliberately neither blank, nor the block subsystem's
own `'STFR'` magic, nor Artemis's `"ARTEMIS\0"` marker. Verified on
amd64 and riscv64 post-fix: boot correctly halts
(`ARTEMIS HALT: unrecognized disk content`) and the file's sha256 is
now byte-for-byte identical before and after boot. Keep this fixture
in this poisoned state — if a future change makes its sha256 change
across a boot, that is exactly the regression this fixture exists to
catch.
Note on incidental header churn: `artemis.img` picks up a few changed
header bytes on every ordinary boot even though no user data changes —
`blk_subsys_attach_device()` always records a fresh `mounted_time` on a
successfully recognized disk, which gets flushed at shutdown. This is
expected bookkeeping, not a bug; revert it before committing rather than
carrying timestamp noise in git history.
- `usb-thumbdrive-test.img` — 64MB raw image backing a QEMU `usb-storage`
device attached to the xHCI controller's bus (`xhci0.0`) for Milestone 2e/
2h hotplug testing, added 2026-08-22. Blank (all zero) — `blkio_usb.c` +
`blk_subsys_attach_device()` wiring (Milestone 2h, done 2026-08-25) attach
it as `BLK_FMT_PROVISIONAL` every time, which is the intended, exercised
state; not yet `BLK_FMT_FORMATTED` via `BLK-CONFIRM-FORMAT`.
- `usb-thumbdrive-test2.img` — 64MB raw image, added 2026-08-25 for
Milestone 2h hot-detach/re-attach verification. Filled with a repeating
`HOTDETACH-REATTACH-FIXTURE-2026-08-25--` ASCII pattern, deliberately
distinguishable from `usb-thumbdrive-test.img`'s all-zero content — the
point is proving a block read *after* detaching `usb-thumbdrive-test.img`
and re-attaching this one actually returns this pattern rather than
silently replaying the old device's cached (all-zero) content, which is
exactly the class of bug a same-LBN-range device swap can cause if the
VM block window cache (`vm->blk_vm_cbuf[]`) isn't re-validated on a hit.
- `artemis-reloc-test.img` — 64MB raw image, added 2026-08-25 for single-block
relocation (`RELOCATE-BLOCK`) persistence verification. Blank at creation;
a genuinely *fresh* volume was required (not `artemis.img`) because
relocation-table capacity (`reloc_start`/`reloc_devblocks` in
`blk_volume_meta_t`) is only reserved by `blk_compute_fresh_geometry()` on
a fresh format — `artemis.img` predates the feature and correctly reads
back `reloc_devblocks=0` from what was previously unused header padding.
Attach via `make -f Makefile.starkernel ARTDISK=disk/artemis-reloc-test.img
...` (the Makefile's `ARTDISK` var is `?=`-overridable). Not yet
`BLK-CONFIRM-FORMAT`-committed in the repo copy — commit that step live if
reusing this fixture for further reloc-table testing.
**Convention, standing as of 2026-08-22: every virtual disk/thumb-drive image
used for testing — Artemis persistence disks above, and USB Mass Storage
backing images alike — lives in this directory and is a tracked, committed
file, never scratchpad.** This was already `artemis.img`'s convention;
`usb-thumbdrive-test.img` and any future USB test images follow the same
rule. Confirmed no `.gitignore` in this repo excludes `disk/*.img`.
These are regenerable QEMU raw disk images, not source — see
`.claude/ARTEMIS.md` for the storage model they exercise.