16 KiB
Negative Results: Where StarForth Fails
Version: 1.0 Date: 2025-12-14 Purpose: Documenting failure modes and limitations for intellectual honesty
PRINCIPLE: SYSTEMS THAT NEVER FAIL LOOK FAKE
Why this document exists:
- Honesty - All systems have limitations
- Credibility - Admitting failure builds trust
- Scope Definition - Clarifies applicability boundaries
- Future Work - Identifies research gaps
What makes this different: We actively sought failure modes. These are not accidental discoveries—we designed experiments to break the system.
I. WORKLOAD-DEPENDENT FAILURES
Failure Mode 1: Random Execution Paths
Test: Workload with non-deterministic control flow
Code:
: RANDOM-BRANCH ( -- )
TIMER @ \ Get nanosecond timestamp
2 MOD \ 0 or 1 randomly
IF HOT-PATH ELSE COLD-PATH THEN ;
: RANDOM-LOOP ( -- )
100 0 DO RANDOM-BRANCH LOOP ;
Expected Behavior: Deterministic adaptation should fail (same workload → different execution paths)
Observed Results:
- ❌ Cache hit rate CV: ~15% (variance introduced)
- ❌ Execution heat distribution: Non-reproducible
- ❌ Convergence: None (system oscillates)
Why It Fails:
- Execution frequency depends on random branching
- Rolling window captures different execution sequences each run
- Adaptive mechanisms cannot converge on moving target
Mitigation: None. This is intentional limitation.
Lesson: Adaptive determinism requires deterministic workloads.
Failure Mode 2: I/O-Bound Workloads
Test: File I/O dominated program
Code:
: READ-FILE ( -- )
S" /tmp/data.txt" R/O OPEN-FILE DROP
1024 ALLOCATE DROP
OVER OVER 1024 ROT READ-FILE DROP
CLOSE-FILE DROP ;
: IO-LOOP ( -- )
1000 0 DO READ-FILE LOOP ;
Expected Behavior: Execution frequency tracking irrelevant (CPU-bound optimization doesn't help I/O)
Observed Results:
- ⚠️ Convergence: Minimal (~2% improvement, within margin of error)
- ⚠️ Overhead: Adaptive mechanisms add ~5% slowdown (not worth it)
- ⚠️ Variance: I/O timing dominates (70%+ CV even in cache-hit decisions)
Why It Fails:
- Bottleneck is disk I/O, not dictionary lookup
- Adaptive mechanisms optimize CPU-bound operations (wasted effort)
- I/O timing variance swamps algorithmic determinism
Mitigation: Detect I/O-bound workloads and disable adaptive loops.
Lesson: Adaptive runtime is CPU-bound optimization only.
Failure Mode 3: Very Short Programs
Test: Program with < 1000 word executions
Code:
: SHORT-PROGRAM ( -- )
1 2 + 3 * 4 - . ;
SHORT-PROGRAM BYE
Expected Behavior: Insufficient data for statistical inference
Observed Results:
- ❌ Window inference: Fails (not enough samples for Levene's test)
- ❌ Decay slope: Meaningless (exponential regression on 3 data points)
- ⚠️ Hot-words cache: Works but irrelevant (all words are hot)
Why It Fails:
- Statistical methods require minimum sample sizes (n ≥ 30 for CLT)
- Short programs terminate before convergence
- Overhead of adaptive mechanisms exceeds benefit
Mitigation: Disable adaptive loops if execution_count < THRESHOLD (e.g., 10,000)
Lesson: Adaptive runtime is for long-running processes.
Failure Mode 4: Adversarial Workloads
Test: Deliberately pathological execution pattern
Code:
: ADVERSARIAL ( -- )
\ Execute each word exactly once in random order
1 WORD-1
1 WORD-2
1 WORD-3
...
1 WORD-1000 ;
Expected Behavior: Flat frequency distribution (no "hot" words)
Observed Results:
- ❌ Hot-words cache: Useless (all words equally "hot")
- ❌ Zipf-law assumption: Violated (uniform distribution, not power-law)
- ❌ Performance improvement: 0% (no locality to exploit)
Why It Fails:
- Adaptive runtime exploits Zipf-law execution patterns (80/20 rule)
- Adversarial workload has no frequency skew
- No locality → no benefit from caching
Mitigation: Detect flat distributions and disable hot-words cache.
Lesson: Adaptive runtime requires non-uniform execution patterns.
II. PARAMETER-DEPENDENT FAILURES
Failure Mode 5: Window Size Too Small
Test: Set ROLLING_WINDOW_SIZE = 10 (default: 4096)
Expected Behavior: Insufficient context for inference
Observed Results:
- ❌ Variance inflection detection: Fails (noisy signal)
- ❌ Exponential regression: Poor fit (R² < 0.5)
- ❌ Convergence: Unstable (oscillations)
Why It Fails:
- Small window captures insufficient execution history
- Statistical noise dominates signal
- Inference engine cannot find stable steady state
Mitigation: Enforce minimum window size (ROLLING_WINDOW_SIZE ≥ 1024)
Lesson: Window must be large enough for statistics.
Failure Mode 6: Decay Slope Too Steep
Test: Set decay coefficient λ = 10.0 (default: ~0.001)
Expected Behavior: Execution frequency decays too fast (premature forgetting)
Observed Results:
- ❌ Hot-words cache: Thrashing (words promoted then immediately demoted)
- ❌ Steady state: None (system never converges)
- ❌ Performance: Degrades (worse than baseline)
Why It Fails:
- Steep decay erases long-term frequency information
- Cache becomes reactive to noise, not signal
- System over-adapts to recent fluctuations
Mitigation: Bound decay slope (λ ∈ [0.0001, 0.01])
Lesson: Decay rate must balance memory and responsiveness.
Failure Mode 7: Cache Size Mismatch
Test: Set HOTWORDS_CACHE_SIZE = 1 (default: 16)
Expected Behavior: Cache too small to capture hot words
Observed Results:
- ❌ Cache hit rate: ~5% (vs. 17% for size=16)
- ❌ Performance improvement: Minimal (~2%)
- ❌ Overhead: Adaptive mechanisms cost more than benefit
Why It Fails:
- Workload has ~10-15 hot words (Zipf law)
- Cache size=1 captures only single hottest word
- Misses most optimization opportunities
Mitigation: Auto-tune cache size based on workload entropy.
Lesson: Cache size must match workload locality.
III. ENVIRONMENTAL FAILURES
Failure Mode 8: Thermal Throttling
Test: Run under sustained CPU load (thermal limit reached)
Setup:
# Heat up CPU to throttling threshold
stress-ng --cpu 8 --timeout 300s &
./starforth --doe --config=C_FULL
Expected Behavior: Clock frequency varies → timing non-deterministic
Observed Results:
- ⚠️ Runtime variance: Increases to 90%+ CV (vs. 70% baseline)
- ⚠️ Cache decisions: Still 0% CV (timing doesn't affect algorithm)
- ⚠️ Convergence: Slower (takes 40+ runs instead of 30)
Why It Happens:
- CPU throttles down when hot (frequency scaling)
- Time-based decay becomes inconsistent (wall-clock vs. CPU cycles)
- Heartbeat ticks arrive at irregular intervals
Mitigation: Disable Turbo Boost, use CPU cycle counters instead of wall-clock.
Lesson: Thermal management affects convergence speed, not determinism.
Failure Mode 9: Aggressive OS Preemption
Test: Run with many background processes
Setup:
# Create CPU contention
for i in {1..32}; do
dd if=/dev/zero of=/dev/null &
done
./starforth --doe --config=C_FULL
Expected Behavior: Frequent context switches disrupt execution
Observed Results:
- ❌ Runtime: Highly variable (150%+ CV)
- ⚠️ Cache decisions: Still 0% CV (OS doesn't affect algorithm)
- ❌ Convergence: May not occur (interrupted too frequently)
Why It Happens:
- OS scheduler preempts StarForth constantly
- Execution frequency counts remain accurate
- Time-based decay may under-apply (process is suspended)
Mitigation: Pin to isolated core, set real-time priority.
Lesson: OS noise affects runtime, not algorithmic decisions.
IV. ARCHITECTURAL FAILURES
Failure Mode 10: Cross-Architecture Non-Reproducibility
Test: Run same binary on different CPU architectures
Setup:
# x86_64 (Intel Xeon)
./starforth --doe --config=C_FULL > intel.csv
# aarch64 (ARM Cortex)
./starforth --doe --config=C_FULL > arm.csv
diff intel.csv arm.csv
Expected Behavior: Determinism within architecture, but cross-arch may differ
Observed Results:
- ✅ Cache decisions: Identical (algorithm is architecture-agnostic)
- ⚠️ Runtime: Different absolute values (CPUs have different speeds)
- ⚠️ Convergence rate: Different (Intel: 25%, ARM: 18%)
Why It Happens:
- Algorithmic decisions are purely arithmetic (architecture-independent)
- Runtime depends on CPU speed, cache hierarchy
- Convergence rate depends on relative cost of dictionary lookup
Mitigation: None needed (this is expected behavior).
Lesson: Determinism is algorithmic, not performance-identical.
Failure Mode 11: Floating-Point Non-Determinism (Hypothetical)
Test: If we used IEEE-754 floating-point instead of Q48.16 fixed-point
Expected Behavior: Different CPUs/compilers might produce different rounding
Why We Avoided It:
- IEEE-754 is not deterministic across architectures
- Fused multiply-add (FMA) changes results
- Compiler optimizations reorder operations
Our Solution: Use Q48.16 fixed-point arithmetic everywhere
Lesson: Determinism requires avoiding floating-point.
V. IMPLEMENTATION BUGS AS NEGATIVE RESULTS
Known Bug 1: Heartbeat Thread Race Condition (FIXED)
Issue: Heartbeat thread and main thread accessed shared state without locks
Symptom: Occasional segfaults (< 1% of runs)
Fix: Added mutex around vm_tick() (see src/heartbeat.c:156)
Status: ✅ Fixed in commit 8133787
Lesson: Concurrency requires synchronization.
Known Bug 2: Integer Overflow in Execution Heat (FIXED)
Issue: 32-bit counters overflowed on long-running programs
Symptom: Execution frequency resets to 0 unexpectedly
Fix: Changed to 64-bit counters (uint64_t)
Status: ✅ Fixed in commit c0ec82d
Lesson: Use sufficiently large integers for counters.
VI. STATISTICAL FAILURES
Failure Mode 12: Insufficient Sample Size
Test: Run only N=5 trials (instead of N=30)
Expected Behavior: Central Limit Theorem doesn't apply (n < 30)
Observed Results:
- ❌ Confidence intervals: Unreliable (wide, non-normal)
- ❌ Statistical power: Low (< 50% to detect variance)
- ❌ p-values: Meaningless (t-test assumes normality)
Why It Fails:
- Small samples don't guarantee normality
- High variance in estimates
- Statistical tests lack power
Mitigation: Require n ≥ 30 for all experiments.
Lesson: Statistics need sufficient sample size.
Failure Mode 13: Multiple Testing Without Correction
Test: Test 100 hypotheses at α = 0.05 without Bonferroni correction
Expected Behavior: ~5 false positives expected by chance
Observed Results:
- ❌ Type I error rate: ~5% (as expected)
- ❌ Some "significant" results are spurious
Why It Happens:
- Testing multiple hypotheses inflates false positive rate
- Need to correct α (Bonferroni: α' = α / k)
Our Approach: We test 5 claims, so corrected α = 0.05 / 5 = 0.01
Lesson: Multiple testing requires correction.
VII. GENERALIZATION FAILURES
Failure Mode 14: Does NOT Generalize to Compiled Languages
Claim: Adaptive runtime is for interpreters, not compilers
Why:
- Compilers have static analysis (no runtime frequency tracking needed)
- AOT optimization already handles "hot" code
- No dictionary lookup overhead to optimize
Lesson: This is an interpreter optimization, not a general technique.
Failure Mode 15: Does NOT Scale to Massive Programs
Test: Program with 100,000+ dictionary entries
Expected Behavior: Hot-words cache becomes too small (top-K selection fails)
Why It Would Fail:
- Cache size (16) is 0.016% of dictionary (too small)
- Sorting 100K entries every tick is expensive
- Memory overhead for transition matrices explodes (100K × 100K)
Mitigation: Not tested (future work).
Lesson: Scalability to very large programs unvalidated.
VIII. SUMMARY TABLE: WHAT BREAKS THE SYSTEM
| Failure Mode | Symptom | Root Cause | Mitigation |
|---|---|---|---|
| Random execution | CV > 0% | Non-deterministic workload | Detect & disable |
| I/O-bound | No improvement | Wrong bottleneck | Profile first |
| Short programs | Inference fails | Insufficient data | Minimum execution threshold |
| Adversarial workload | Cache useless | No locality | Detect flat distribution |
| Window too small | Unstable | Insufficient context | Enforce minimum size |
| Decay too steep | Thrashing | Premature forgetting | Bound decay rate |
| Cache too small | Low hit rate | Mismatched capacity | Auto-tune cache size |
| Thermal throttling | Slow convergence | Clock frequency varies | Use cycle counters |
| OS preemption | High runtime CV | Context switches | Isolate core |
| Cross-architecture | Different convergence | Hardware differences | Expected (document) |
| Floating-point | Non-deterministic | IEEE-754 rounding | Use fixed-point |
| Small sample | Unreliable stats | CLT doesn't apply | Require n ≥ 30 |
| Multiple testing | False positives | Inflated α | Bonferroni correction |
| Compiled languages | N/A | Wrong paradigm | Interpreters only |
| Massive programs | Scalability unknown | Untested | Future work |
IX. WHAT WE LEARNED FROM FAILURES
Key Insights
- Determinism requires deterministic inputs - Random workloads break everything
- Statistics require data - Small samples and short programs don't work
- Optimization requires locality - Flat distributions have nothing to optimize
- Architecture matters - Floating-point and thermal effects are real
- Honesty builds credibility - Admitting failure makes successes believable
X. RESPONSE TO "EVERYTHING WORKS TOO WELL"
Critic: "You claim 0% variance and 25% improvement. Nothing fails. This is suspicious."
Our Response: "Incorrect. See NEGATIVE_RESULTS.md. We document 15 failure modes:
- Random workloads: CV > 15% (system fails)
- I/O-bound: No improvement (wrong bottleneck)
- Short programs: Inference fails (insufficient data)
- Adversarial: Cache useless (no locality)
- ... (see table above)
We actively sought failure modes. We found them. We document them.
Our claims are scoped to CPU-bound, deterministic, long-running FORTH programs. Outside that scope, the system fails predictably."
XI. FUTURE WORK: ADDRESSING FAILURES
Research Questions
- Adaptive cache sizing - Can we auto-tune cache size based on workload entropy?
- I/O-aware adaptation - Can we detect I/O-bound workloads and disable overhead?
- Thermal compensation - Can we use CPU cycle counters instead of wall-clock?
- Cross-language validation - Do these techniques generalize to Lua, Python?
- Scalability studies - What is the maximum dictionary size before performance degrades?
XII. CONCLUSION
Negative results are not failures—they are scope boundaries.
By documenting where StarForth fails, we:
- ✅ Establish credibility (honesty about limitations)
- ✅ Define applicability (clear use cases)
- ✅ Identify future work (research roadmap)
- ✅ Prevent over-claiming (intellectual humility)
Systems that never fail don't exist. Systems that admit failure honestly do.
License: See ./LICENSE