Files
LithosAnanake/docs/working/experiments/02-experiments/physics-optimization/proposals.md
T

17 KiB
Raw Blame History

Physics-Driven Optimization Proposals for StarForth

Executive Summary

The hot-words cache experiment proves that physics-driven optimization (metrics collection → real-time decisions → zero manual tuning) delivers measurable performance improvements while remaining verifiable and maintainable.

This document proposes 9 additional optimization opportunities in StarForth using similar methodologies. Each leverages the execution_heat tracking already in place and requires no external libraries or floating-point arithmetic.

Note on JIT Compilation: JIT is NOT recommended for StarForth. See section below.


Why Not JIT?

JIT Violates StarForth's Core Constraints

  1. Runtime Code Generation → Dynamic memory management → Heap complexity
  2. Formal Verification Nightmare → Generated code can't be proven correct
  3. No Floating-Point Allowed → JIT compilation itself requires FP overhead
  4. L4Re Incompatibility → Microkernel doesn't allow arbitrary code generation
  5. Security Risk → JIT compilation surfaces are attack vectors

Why Physics-Driven Beats JIT

Aspect JIT Physics-Driven
Verification Hard (generated code) Easy (static analysis)
L4Re Compatible No Yes
Floating-Point Required Yes No (Q48.16)
Memory Overhead 📈 High (code cache) 📉 Low (metadata only)
Incremental All-or-nothing Apply where ROI is highest
Profitability 📊 30-100× improvement 📊 1.78-3× improvement (sufficient!)

Verdict: Physics-driven approach is superior for StarForth's constraints. We gain 1.78× speedup with zero verification burden and perfect L4Re compatibility. That's a winning trade.


Opportunity #1: Return Stack Prediction & Colon Word Inlining

Problem

When a colon definition calls another word, we execute a full dictionary lookup + function call. For frequently-called small words, this overhead dominates.

Example:

: SQUARE ( n -- n^2 )  DUP * ;
: DISTANCE ( x y -- r )  SQUARE SWAP SQUARE + SQRT ;

Every call to DISTANCE triggers:

  • Dictionary lookup for SQUARE (even though it's predictable)
  • Return stack push/pop overhead
  • Small function body execution

Solution: Execution Heat-Based Inlining

Metrics to Track:

  • word->execution_heat (already tracked!)
  • definition->call_count (how many words call this colon definition)
  • definition->size_bytes (total body size)

Decision Rule:

// Inline a colon definition if:
// 1. It's called frequently (execution_heat > 100)
// 2. It's small enough (size_bytes < 256)
// 3. Total expanded size won't exceed L-cache (8KB)
if (word->execution_heat > 100 &&
    definition->size_bytes < 256 &&
    total_expanded < 8192) {
    // Inline the word body instead of calling it
    vm_compile_inlined(definition);
}

Expected Performance Impact:

  • Eliminate dictionary lookups for inlined words
  • Remove return stack overhead
  • Expected gain: 1.32.0× for call-heavy programs

Implementation Complexity: Medium (requires compile-time analysis)

Verification: Static (can prove inlining correctness via Isabelle)


Opportunity #2: Block I/O Prefetching Based on Access Patterns

Problem

Block storage I/O is expensive. When a program accesses BLOCK 10, it often follows with BLOCK 11, BLOCK 12, etc.

Current: Each BLOCK word triggers I/O independently (no lookahead).

Solution: Execution Heat-Tracked Access Patterns

Metrics to Track:

  • block->last_access_ns (monotonic timestamp)
  • block->access_sequence (last N block IDs accessed)
  • block->next_block_heat[MAX_BLOCKS] (how often block X follows block Y)

Decision Rule:

// When accessing block N, if we predict block M comes next:
// Prefetch block M into buffer cache
if (predicted_next_block != NULL &&
    predicted_next_block->access_heat > PREFETCH_THRESHOLD) {
    blk_prefetch_into_cache(predicted_next_block);
}

// Transition heat calculation (Markov-style):
transition_heat[from_block][to_block]++;

Expected Performance Impact:

  • Eliminate blocking I/O wait for sequential access patterns
  • Cache warm start for predictable block sequences
  • Expected gain: 1.53.0× for block-heavy programs

Implementation Complexity: Medium (requires Markov state tracking)

Verification: Dynamic (observable via block I/O metrics)


Opportunity #3: Stack Operation Fusion & Pattern Detection

Problem

Common patterns like DUP DROP, SWAP ROT, OVER SWAP are executed frequently but could be fused into single operations.

Current:

DUP DROP    \ 2 lookups, 2 executions
SWAP ROT    \ 2 lookups, 2 executions
OVER SWAP   \ 2 lookups, 2 executions

Expected: ~6-10 nanoseconds per operation pair

Solution: Pattern-Driven Fusion at Compile Time

Metrics to Track:

  • Track consecutive word execution pairs: word[i]word[i+1]
  • pair_heat[WORD_A][WORD_B] (how often does A precede B?)
  • Threshold: fusion candidate if pair_heat > 100

Decision Rule:

// During compilation, detect hot patterns:
if (pair_heat[DUP][DROP] > 100) {
    // Fuse into single primitive: DUP_DROP (no-op)
    register_word(vm, "DUP-DROP", forth_DUP_DROP);
}

if (pair_heat[SWAP][ROT] > 100) {
    // Fuse: SWAP followed by ROT = SWAP_ROT primitive
    register_word(vm, "SWAP-ROT", forth_SWAP_ROT);
}

Expected Performance Impact:

  • Eliminate dictionary lookups for fused patterns
  • Reduce call overhead for small operations
  • Expected gain: 1.21.5× for stack-heavy programs

Implementation Complexity: Low (pattern detection + primitive registration)

Verification: Static (proven by composition of verified primitives)


Opportunity #4: Memory Allocation Pattern Prediction

Problem

Dynamic memory allocation (ALLOCATE, RESIZE) has overhead. Patterns (e.g., "allocate 256 bytes, then 256 bytes again") repeat.

Current: Every allocation goes through full allocator path

Solution: Heat-Tracked Allocation Sizes

Metrics to Track:

  • allocation_heat[SIZE] (how many times is SIZE requested?)
  • allocation_pattern_sequence (last N allocation sizes)
  • allocation_time[SIZE] (execution time for allocations)

Decision Rule:

// Pre-allocate a pool for hot sizes:
if (allocation_heat[256] > 1000) {
    // Keep 10 pre-allocated 256-byte blocks warm
    allocator_warmpool_size(256, 10);
}

// For allocation-heavy patterns, reuse buffers:
if (SEQUENCE == [256, 256, 512, 256, 256, 512]) {
    // Detected repeating pattern - cache allocations
    allocator_pattern_cache(pattern);
}

Expected Performance Impact:

  • Eliminate allocator overhead for pre-warmed sizes
  • Reduce fragmentation via pattern-aware pooling
  • Expected gain: 1.31.8× for allocation-heavy programs

Implementation Complexity: Medium (requires allocator integration)

Verification: Dynamic (observable via allocation metrics)


Opportunity #5: Control Flow Branch Prediction

Problem

IF/THEN/ELSE branches have no prediction. In loops, the same branch is taken repeatedly (predictable).

Current:

: SEARCH ( n -- ? )
  BEGIN
    DUP 10 <
    IF  ... ELSE ... THEN
  UNTIL

Every iteration re-evaluates the condition and potentially mis-speculates.

Solution: Execution Heat-Tracked Branch Patterns

Metrics to Track:

  • Track outcomes of IF/THEN/ELSE at each location
  • branch_heat[location][taken] (how often does branch X go True vs. False?)
  • branch_sequence (what's the typical pattern?)

Decision Rule:

// For hot branches, set a prediction hint:
if (branch_heat[location][TRUE] > 90 &&
    branch_heat[location][FALSE] < 10) {
    // Mark this branch as "likely true"
    vm->current_executing_entry->flags |= BRANCH_LIKELY_TRUE;
}

// Inner interpreter can use hint for:
// - CPU branch prediction (x86 PREFETCHNTA hints)
// - Reordering of code paths
// - Speculative stack preparation

Expected Performance Impact:

  • Reduce branch mispredictions in CPU pipeline
  • Reorder code for better CPU cache behavior
  • Expected gain: 1.11.3× (CPU-dependent)

Implementation Complexity: Medium (requires CPU-specific assembly work)

Verification: Dynamic (observable via perf branch-misses)


Opportunity #6: Vocabulary Search Path Optimization

Problem

Multi-vocabulary systems (FORTH, EDITOR, SYSTEM) search vocabularies in order. If most lookups hit FORTH, other vocabularies waste cycles.

Current: Always search: EDITOR → SYSTEM → FORTH

Solution: Dynamic Search Order Based on Execution Heat

Metrics to Track:

  • vocab_hit_rate[EDITOR], vocab_hit_rate[SYSTEM], vocab_hit_rate[FORTH]
  • vocab_lookup_time[VOCAB] (latency per vocabulary)
  • Track which vocabulary actually had the hit

Decision Rule:

// Reorder search path based on hit rates:
// If FORTH has 85% hits, search it first
sorted_vocabs = sort_by_hit_rate([EDITOR, SYSTEM, FORTH]);

// Update search order dynamically:
if (hit_rate[FORTH] > hit_rate[EDITOR]) {
    swap_search_order(FORTH, EDITOR);
}

Expected Performance Impact:

  • Reduce dictionary lookup depth for hot vocabulary
  • Fewer misses on wrong vocabularies
  • Expected gain: 1.21.6× (vocab-dependent)

Implementation Complexity: Low (requires search path reordering)

Verification: Static (observable via lookup metrics)


Opportunity #7: String Operation Batching

Problem

String operations (TYPE, EMIT, PARSE) are called frequently on small strings. Each call has interpretation overhead.

Current:

." Hello " ." World " CR

Compiles to 3 separate operations with 3 lookups each.

Solution: Compile-Time String Fusion

Metrics to Track:

  • Detect consecutive string literals: ." A " ." B " pattern
  • string_op_heat[TYPE] (how often are consecutive string ops used?)
  • String size statistics

Decision Rule:

// During compilation, fuse consecutive string operations:
if (word_sequence == [EMIT, EMIT, EMIT] &&
    all_immediate) {
    // Fuse into single STRING_OUTPUT operation
    // Batch output: less interpreter overhead
    emit_fused_string_output(concatenated_string);
}

Expected Performance Impact:

  • Eliminate overhead for small string operations
  • Improve I/O efficiency via batching
  • Expected gain: 1.11.4× (I/O bound dependent)

Implementation Complexity: Low (compile-time pattern detection)

Verification: Static (observable via string operation metrics)


Opportunity #8: Arithmetic Operation Reordering & Associativity

Problem

Commutative operations (+ - * AND OR) are executed in fixed order. Some orders have better CPU cache behavior.

Example:

A B +  C +  D +    \ Left associative
A B C D + + +      \ Right associative (different cache behavior)

Solution: Heat-Tracked Operand Ordering

Metrics to Track:

  • operand_heat[RESULT] (how hot is the result of this operation?)
  • Track which operand sizes are common
  • CPU cache line utilization for different orderings

Decision Rule:

// For hot commutative operations, try both orderings:
if (add_heat[location] > 1000 &&
    operand_size_a != operand_size_b) {
    // Try reordering for better cache behavior
    // A=large, B=small: load small first (better prefetch)
    if (size_b < size_a) {
        emit_operation(B, A, ADD);  // Reordered
    }
}

Expected Performance Impact:

  • Improve CPU cache line reuse
  • Reduce memory latency for hot arithmetic
  • Expected gain: 1.051.15× (CPU-dependent)

Implementation Complexity: Medium (requires cache analysis)

Verification: Dynamic (observable via perf cache-misses)


Opportunity #9: Word Placement Optimization (Code Layout)

Problem

Dictionary layout is insertion-order (no optimization). Hot words might be far from each other, causing CPU cache misses.

Current: Words stored in linked-list order they were defined

Solution: Heat-Driven Memory Compaction

Metrics to Track:

  • word->execution_heat (already tracked!)
  • word->co_execution_heat[OTHER_WORD] (how often are these words called together?)
  • word->memory_address (track actual memory layout)

Decision Rule:

// During garbage collection / compaction:
// 1. Calculate which words are frequently called together
// 2. Place hot word clusters adjacent in memory
// 3. Minimize cache line misses for hot word sequences

hot_cluster = [EXIT, LIT, CR]  // Frequently co-executed
memory_compact_together(hot_cluster);  // Colocate in memory

// Result: prefetch-friendly dictionary layout

Expected Performance Impact:

  • Reduce CPU instruction cache misses
  • Better temporal locality for hot code paths
  • Expected gain: 1.051.2× (CPU architecture-dependent)

Implementation Complexity: High (requires memory layout changes, GC integration)

Verification: Dynamic (observable via perf icache-misses)


Implementation Roadmap

Phase 1: Foundation (Already Complete )

  • Execute heat tracking in VM
  • Hot-words cache (1.78× speedup proven)
  • Bayesian inference for confidence
  • Q48.16 fixed-point arithmetic
  1. Stack Operation Fusion (#3) - Highest ROI, lowest complexity
  2. Vocabulary Search Reordering (#6) - Minimal complexity
  3. Return Stack Prediction (#1) - Medium complexity, high impact

Phase 3: Medium Complexity (Future)

  1. Block I/O Prefetching (#2) - Medium complexity
  2. String Operation Batching (#7) - Low/medium complexity
  3. Memory Allocation Patterns (#4) - Moderate complexity

Phase 4: Advanced (Future)

  1. Control Flow Branch Prediction (#5) - High complexity, CPU-specific
  2. Arithmetic Operation Reordering (#8) - Medium complexity
  3. Word Placement Optimization (#9) - High complexity (requires GC)

Unified Metrics Infrastructure

All 9 opportunities require a common metrics collection framework:

typedef struct {
    // Execution frequency (already have this)
    uint64_t execution_heat;

    // Co-execution patterns (NEW)
    uint64_t co_execution_heat[DICT_SIZE];

    // Timing information (ENHANCED)
    uint64_t total_time_ns;
    uint64_t call_count;

    // Cache behavior (NEW)
    uint32_t cache_line_hits;
    uint32_t cache_line_misses;

    // Sequence patterns (NEW)
    circular_buffer_t recent_execution;

} PhysicsMetrics;

One-time investment in unified metrics → all 9 opportunities unlocked.


Expected Cumulative Performance Gain

Conservative Estimate

Baseline:                1.0×
+ Hot-words cache:       1.78×
+ Stack fusion:          2.13× (1.78 × 1.2)
+ Vocab reordering:      2.56× (2.13 × 1.2)
+ Return prediction:     3.30× (2.56 × 1.3)
──────────────────────────────
Total (Phase 2):         3.3× speedup

Ambitious Estimate (All 9)

Cumulative with all optimizations:  58× speedup possible

But Remember

  • Not all programs benefit from all optimizations
  • Each optimization has CPU/workload dependency
  • Diminishing returns apply (law of diminishing returns)
  • 1.78× from cache alone is already substantial!

Why Physics-Driven Beats Alternatives

Approach Speedup Verification L4Re Compatible Complexity
JIT Compilation 30100× Impossible No Very High
Static Optimization 1.051.2× Easy Yes Medium
Physics-Driven (current) 1.78× Easy Yes LowMedium
Physics-Driven (all 9) 58× Easy Yes MediumHigh

Verdict: Physics-driven is optimal for StarForth's constraints. We get measurable gains (1.78×+), complete verifiability, and L4Re compatibility all in one package.


References

  • Experiment Report: docs/PHYSICS_HOTWORDS_CACHE_EXPERIMENT.md
  • Reproduction Guide: docs/REPRODUCE_PHYSICS_EXPERIMENT.md
  • Physics Architecture: include/physics_metadata.h, include/physics_hotwords_cache.h
  • CLAUDE.md: Core design patterns

Summary: Strategic Direction

  1. Physics-Driven Optimization is the right approach for StarForth
  2. Hot-words cache (1.78×) is proven and ready for production
  3. 9 additional opportunities can stack to 58× total gain
  4. Low-hanging fruit (stack fusion, vocab reordering) should come next
  5. JIT is wrong direction (violates constraints, harder to verify)
  6. Incremental approach matches StarForth's design philosophy

The physics model is the foundation. Build on it, measure everything, verify rigorously.