Fix M*: use __int128 for a genuine 128-bit double-cell product (FABRIC-3.md §XII.4)
Build / build-amd64-iso (push) Waiting to run
Build / build-aarch64-iso (push) Waiting to run
Build / build-riscv64-img (push) Waiting to run

mixed_math_word_m_star() confused "double" (two full cell_t-width cells,
128 bits total on this 64-bit build -- what D+/D-/D./etc. all actually
expect) with "the low/high 32-bit halves of a single 64-bit product" --
code clearly written assuming cell_t is 32-bit. It computed an ordinary
64-bit `long long` product (already wrong for any true product exceeding
64 bits, since long long is the same width as one cell here) and split
that into 32-bit halves via `result & 0xFFFFFFFF` / `result >> 32`. For a
small negative product like -56088, this produced a positive, zero-
extended low cell paired with a correctly-looking dhigh=-1 -- D.'s
overflow check (correctly) rejected the resulting malformed double, on
every architecture, every time (this bug was never architecture-specific,
unlike the D+/D-/DNEGATE/d_compare family already fixed in bea8d74/
1a716c8).

Fixed via __int128 for a genuine 64x64->128-bit multiply, no truncation --
an already-established safe pattern in this codebase for exactly this
operation (src/starkernel/crypto/fe25519.c/scalar25519.c already use it;
src/starkernel/arch/amd64/timer.c's own doc comment confirms __int128
multiply/shift-by-constant compile cleanly with zero undefined symbols on
all three target toolchains -- only division needs unavailable libgcc
support, not used here).

Verified: rebuilt and booted all three architectures clean. T13/T14 both
correct everywhere (56088/-56088). Manual cases confirm genuine 128-bit
precision: -1 -1 M* -> 1; 1000000000000 1000000000000 M* -> dhigh=54210
dlow=2003764205206896640 (10^12 x 10^12 = 10^24, correctly exceeding 64
bits). Full 24-case exerciser reran clean on all three, no regressions.

Found while verifying, NOT fixed (report only, out of scope of this
request): mixed_math_word_m_slash_mod() (M/MOD, the very next function in
the same file) has the identical bug class -- confirmed genuinely broken
for a true wide double (feeding this fix's own correct large-magnitude M*
output into M/MOD produces a result wrong by many orders of magnitude and
the wrong sign). Flagged in FABRIC-3.md for a future fix request.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXieurDfDSsDFdnSyusuWo
This commit is contained in:
Robert Allan James
2026-09-11 08:23:51 -04:00
co-authored by Claude Sonnet 5
parent 1a716c8048
commit 9a09949c69
10 changed files with 27136 additions and 13 deletions
+46 -6
View File
@@ -2142,8 +2142,9 @@ against the other two architectures once the campaign completes; not yet root-ca
as a bug on its own. as a bug on its own.
### XII.4 — Campaign completed: full 27-leg run (9 identities × 3 architectures), two real ### XII.4 — Campaign completed: full 27-leg run (9 identities × 3 architectures), two real
FORTH-79 engine bugs found — bug 2 (`D+`/`D-`/`DNEGATE`/`d_compare`) root-caused and FIXED FORTH-79 engine bugs found — BOTH root-caused and FIXED 2026-09-11 (bug 1 `M*`; bug 2 `D+`/`D-`/
2026-09-11; bug 1 (`M*`) still open, report only (2026-09-10, after §XIII's WIREBIND fix) `DNEGATE`/`d_compare`); a third, `M/MOD`, found while fixing bug 1, reported not fixed
(2026-09-10 campaign, after §XIII's WIREBIND fix)
All 27 legs run: `zuse` (auto-attached, exercised directly on Hera's own console) plus `rajames` All 27 legs run: `zuse` (auto-attached, exercised directly on Hera's own console) plus `rajames`
(the `bob` thumbdrive's actual registered identity -- see the naming-mismatch note below) and (the `bob` thumbdrive's actual registered identity -- see the naming-mismatch note below) and
@@ -2169,10 +2170,49 @@ which identity is running them -- expected, and a useful negative result on its
project's standing rule (report, don't fix without being asked):** project's standing rule (report, don't fix without being asked):**
1. **`M*` on a negative operand → `D.` reports `DOUBLE-OVERFLOW`, on all three architectures 1. **`M*` on a negative operand → `D.` reports `DOUBLE-OVERFLOW`, on all three architectures
identically.** `T14: -123 456 M* SWAP D. CR` should print `-56088`; it prints `DOUBLE-OVERFLOW` identically — root-caused and FIXED 2026-09-11.** `T14: -123 456 M* SWAP D. CR` should print
on amd64, aarch64, *and* riscv64 -- an engine-level bug (in `M*`'s double-cell result, or in `-56088`; it printed `DOUBLE-OVERFLOW` on amd64, aarch64, *and* riscv64 -- an engine-level bug,
`D.`'s own overflow check, or both), not architecture-specific. Universal, 100% reproducible universal, 100% reproducible across all 9 identities × 3 architectures (unlike bug 2 below,
across all 9 identities × 3 architectures. this one really is architecture-independent, since it doesn't involve `long`'s width at all).
**Root cause:** `mixed_math_word_m_star()` (`mixed_arithmetic_words.c`) confused "double" (two
full `cell_t`-width cells, 128 bits total on this 64-bit build -- what `D+`/`D-`/`D.`/etc. all
actually expect) with "the low/high 32-bit halves of a single 64-bit product" -- code clearly
written assuming `cell_t` is 32-bit. It computed an ordinary 64-bit `long long` product
(`(long long) n1 * (long long) n2`, already wrong/truncated for any true product not fitting
in 64 bits, since `long long` is only 64-bit here, the same width as one cell already) and
then split *that* into 32-bit low/high halves via `result & 0xFFFFFFFF` / `result >> 32`. For
a small negative product like `-56088` (fits entirely in one 64-bit cell with sign extension),
this produced a *positive*, zero-extended low cell paired with a correctly-looking `dhigh=-1`
-- the exact same "malformed double, `D.` correctly rejects it" shape as bug 2's `D+` defect,
just from an entirely different, unrelated piece of code.
**Fix:** switched to `__int128` for a genuine 64×64→128-bit multiply (no truncation), an
already-established safe pattern in this codebase for exactly this operation --
`src/starkernel/crypto/fe25519.c`/`scalar25519.c` already use it, and
`src/starkernel/arch/amd64/timer.c`'s own doc comment confirms `__int128`
multiply/shift-by-constant compile cleanly with zero undefined symbols on all three target
toolchains (only *division* needs libgcc support unavailable in this freestanding build --
not needed for `M*`). `dlow` = low 64 bits of the 128-bit product, `dhigh` = high 64 bits
(arithmetic shift, sign-preserving).
**Verified:** rebuilt and booted all three architectures clean. `T13`/`T14` both correct
everywhere now (`56088`/`-56088`). Additional manual cases run live on all three, identical
results: `-1 -1 M*``1` (small negative×negative); `1000000000000 1000000000000 M*`
`dhigh=54210 dlow=2003764205206896640` (10^12 × 10^12 = 10^24, correctly exceeding 64 bits --
proof the multiply itself is now genuinely 128-bit, not just sign-extension-correct for small
values). Full 24-case exerciser reran clean on all three, no regressions.
**Found while verifying, NOT fixed (report only, out of scope of what was requested):**
`mixed_math_word_m_slash_mod()` (`M/MOD`, the very next function in the same file) has the
*identical* bug class -- it reconstructs the dividend from `(dhigh << 32) | (dlow &
0xFFFFFFFF)`, the same wrong "32-bit halves of a 64-bit value" assumption `M*` had. Latent for
small inputs (`T15: 100000 S>D SWAP 7 M/MOD . .` -> `14285 5`, correct, because `100000` fits
entirely within 32 bits so the reconstruction coincidentally works), but confirmed genuinely
broken for a true wide double: feeding the now-fixed `M*`'s own correct large-magnitude output
into `M/MOD` (`1000000000000 1000000000000 M* 1000000 M/MOD . .`, expected quotient
`10^18`) produces `-6845471433603` remainder, `-99710` quotient -- wrong by many orders of
magnitude and the wrong sign. Not touched here; flag for a future fix request.
2. **`D+` on two negative doubles → `D.` reports `DOUBLE-OVERFLOW`, aarch64 only — root-caused 2. **`D+` on two negative doubles → `D.` reports `DOUBLE-OVERFLOW`, aarch64 only — root-caused
2026-09-11, FIXED 2026-09-11 (`D+`/`D-`/`DNEGATE`).** 2026-09-11, FIXED 2026-09-11 (`D+`/`D-`/`DNEGATE`).**
+1 -1
View File
@@ -1,5 +1,5 @@
# Capsule Block Manifest — Auto-generated # Capsule Block Manifest — Auto-generated
<!-- Generated by mkcapsule --manifest 2026-09-11T11:42:38Z --> <!-- Generated by mkcapsule --manifest 2026-09-11T12:21:21Z -->
<!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. --> <!-- DO NOT EDIT — re-run mkcapsule --manifest to refresh. -->
<!-- Hand-written justifications and immutability notes live --> <!-- Hand-written justifications and immutability notes live -->
<!-- in MANIFEST.md alongside this auto-generated index. --> <!-- in MANIFEST.md alongside this auto-generated index. -->
BIN
View File
Binary file not shown.
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
File diff suppressed because it is too large Load Diff
+23 -6
View File
@@ -121,9 +121,26 @@ void mixed_math_word_m_minus(VM *vm) {
* @brief M* ( n1 n2 -- d ) * @brief M* ( n1 n2 -- d )
* *
* Multiplies two single-cell signed integers, producing a double-cell result. * Multiplies two single-cell signed integers, producing a double-cell result.
* On 64-bit builds the product is computed via @c long long and the low 32 bits * A "double" here is two full @c cell_t-width cells (128 bits total on a
* go to the deeper stack cell while the high 32 bits go to TOS. On 32-bit builds * 64-bit build), matching what D+/D-/D./etc. all expect -- NOT the low/high
* the result fits in a single cell so the high cell is pushed as 0. * 32-bit halves of a single 64-bit product. The previous implementation
* confused the two: it computed an ordinary 64-bit @c long long product
* (already truncated/overflowed for any operand pair whose true product
* doesn't fit in 64 bits) and then split THAT into 32-bit low/high halves,
* as if @c cell_t were 32-bit. For a small negative product like
* `-123 456 M*` (-56088, which fits entirely in one 64-bit cell with sign
* extension), that produced a positive, zero-extended @c dlow paired with a
* correctly-looking @c dhigh=-1 -- a malformed double that D.'s overflow
* check (correctly) rejected as DOUBLE-OVERFLOW on every architecture,
* every time a negative product was printed (FABRIC-3.md §XII.4 bug 1).
*
* Fixed via @c __int128 (a genuine 64x64->128-bit multiply, no truncation)
* -- an established, safe pattern in this codebase for exactly this
* operation: see src/starkernel/crypto/fe25519.c and scalar25519.c, and
* src/starkernel/arch/amd64/timer.c's own doc comment confirming
* __int128 multiply/shift-by-constant compile cleanly with zero undefined
* symbols on all three target toolchains (only *division* needs libgcc
* support unavailable in this freestanding build -- not used here).
* *
* Stack effect: ( n1 n2 -- d_low d_high ) TOS = high * Stack effect: ( n1 n2 -- d_low d_high ) TOS = high
* *
@@ -139,9 +156,9 @@ void mixed_math_word_m_star(VM *vm) {
cell_t n1 = vm_pop(vm); cell_t n1 = vm_pop(vm);
if (sizeof(cell_t) == 8) { if (sizeof(cell_t) == 8) {
long long result = (long long) n1 * (long long) n2; __int128 result = (__int128) n1 * (__int128) n2;
cell_t dlow = (cell_t)(result & 0xFFFFFFFFLL); cell_t dlow = (cell_t)(unsigned __int128) result; // low 64 bits
cell_t dhigh = (cell_t)(result >> 32); cell_t dhigh = (cell_t)(result >> 64); // high 64 bits, sign-preserving
vm_push(vm, dlow); // lo first vm_push(vm, dlow); // lo first
vm_push(vm, dhigh); // hi last (TOS) vm_push(vm, dhigh); // hi last (TOS)
} else { } else {