17 KiB
Physics-Driven Optimization Proposals for StarForth
Executive Summary
The hot-words cache experiment proves that physics-driven optimization (metrics collection → real-time decisions → zero manual tuning) delivers measurable performance improvements while remaining verifiable and maintainable.
This document proposes 9 additional optimization opportunities in StarForth using similar methodologies. Each leverages the execution_heat tracking already in place and requires no external libraries or floating-point arithmetic.
Note on JIT Compilation: JIT is NOT recommended for StarForth. See section below.
Why Not JIT?
JIT Violates StarForth's Core Constraints
- Runtime Code Generation → Dynamic memory management → Heap complexity
- Formal Verification Nightmare → Generated code can't be proven correct
- No Floating-Point Allowed → JIT compilation itself requires FP overhead
- L4Re Incompatibility → Microkernel doesn't allow arbitrary code generation
- Security Risk → JIT compilation surfaces are attack vectors
Why Physics-Driven Beats JIT
| Aspect | JIT | Physics-Driven |
|---|---|---|
| Verification | ❌ Hard (generated code) | ✅ Easy (static analysis) |
| L4Re Compatible | ❌ No | ✅ Yes |
| Floating-Point Required | ✅ Yes | ❌ No (Q48.16) |
| Memory Overhead | 📈 High (code cache) | 📉 Low (metadata only) |
| Incremental | ❌ All-or-nothing | ✅ Apply where ROI is highest |
| Profitability | 📊 30-100× improvement | 📊 1.78-3× improvement (sufficient!) |
Verdict: Physics-driven approach is superior for StarForth's constraints. We gain 1.78× speedup with zero verification burden and perfect L4Re compatibility. That's a winning trade.
Opportunity #1: Return Stack Prediction & Colon Word Inlining
Problem
When a colon definition calls another word, we execute a full dictionary lookup + function call. For frequently-called small words, this overhead dominates.
Example:
: SQUARE ( n -- n^2 ) DUP * ;
: DISTANCE ( x y -- r ) SQUARE SWAP SQUARE + SQRT ;
Every call to DISTANCE triggers:
- Dictionary lookup for SQUARE (even though it's predictable)
- Return stack push/pop overhead
- Small function body execution
Solution: Execution Heat-Based Inlining
Metrics to Track:
word->execution_heat(already tracked!)definition->call_count(how many words call this colon definition)definition->size_bytes(total body size)
Decision Rule:
// Inline a colon definition if:
// 1. It's called frequently (execution_heat > 100)
// 2. It's small enough (size_bytes < 256)
// 3. Total expanded size won't exceed L-cache (8KB)
if (word->execution_heat > 100 &&
definition->size_bytes < 256 &&
total_expanded < 8192) {
// Inline the word body instead of calling it
vm_compile_inlined(definition);
}
Expected Performance Impact:
- Eliminate dictionary lookups for inlined words
- Remove return stack overhead
- Expected gain: 1.3–2.0× for call-heavy programs
Implementation Complexity: Medium (requires compile-time analysis)
Verification: Static (can prove inlining correctness via Isabelle)
Opportunity #2: Block I/O Prefetching Based on Access Patterns
Problem
Block storage I/O is expensive. When a program accesses BLOCK 10, it often follows with BLOCK 11, BLOCK 12, etc.
Current: Each BLOCK word triggers I/O independently (no lookahead).
Solution: Execution Heat-Tracked Access Patterns
Metrics to Track:
block->last_access_ns(monotonic timestamp)block->access_sequence(last N block IDs accessed)block->next_block_heat[MAX_BLOCKS](how often block X follows block Y)
Decision Rule:
// When accessing block N, if we predict block M comes next:
// Prefetch block M into buffer cache
if (predicted_next_block != NULL &&
predicted_next_block->access_heat > PREFETCH_THRESHOLD) {
blk_prefetch_into_cache(predicted_next_block);
}
// Transition heat calculation (Markov-style):
transition_heat[from_block][to_block]++;
Expected Performance Impact:
- Eliminate blocking I/O wait for sequential access patterns
- Cache warm start for predictable block sequences
- Expected gain: 1.5–3.0× for block-heavy programs
Implementation Complexity: Medium (requires Markov state tracking)
Verification: Dynamic (observable via block I/O metrics)
Opportunity #3: Stack Operation Fusion & Pattern Detection
Problem
Common patterns like DUP DROP, SWAP ROT, OVER SWAP are executed frequently but could be fused into single operations.
Current:
DUP DROP \ 2 lookups, 2 executions
SWAP ROT \ 2 lookups, 2 executions
OVER SWAP \ 2 lookups, 2 executions
Expected: ~6-10 nanoseconds per operation pair
Solution: Pattern-Driven Fusion at Compile Time
Metrics to Track:
- Track consecutive word execution pairs:
word[i]→word[i+1] pair_heat[WORD_A][WORD_B](how often does A precede B?)- Threshold: fusion candidate if
pair_heat > 100
Decision Rule:
// During compilation, detect hot patterns:
if (pair_heat[DUP][DROP] > 100) {
// Fuse into single primitive: DUP_DROP (no-op)
register_word(vm, "DUP-DROP", forth_DUP_DROP);
}
if (pair_heat[SWAP][ROT] > 100) {
// Fuse: SWAP followed by ROT = SWAP_ROT primitive
register_word(vm, "SWAP-ROT", forth_SWAP_ROT);
}
Expected Performance Impact:
- Eliminate dictionary lookups for fused patterns
- Reduce call overhead for small operations
- Expected gain: 1.2–1.5× for stack-heavy programs
Implementation Complexity: Low (pattern detection + primitive registration)
Verification: Static (proven by composition of verified primitives)
Opportunity #4: Memory Allocation Pattern Prediction
Problem
Dynamic memory allocation (ALLOCATE, RESIZE) has overhead. Patterns (e.g., "allocate 256 bytes, then 256 bytes again") repeat.
Current: Every allocation goes through full allocator path
Solution: Heat-Tracked Allocation Sizes
Metrics to Track:
allocation_heat[SIZE](how many times is SIZE requested?)allocation_pattern_sequence(last N allocation sizes)allocation_time[SIZE](execution time for allocations)
Decision Rule:
// Pre-allocate a pool for hot sizes:
if (allocation_heat[256] > 1000) {
// Keep 10 pre-allocated 256-byte blocks warm
allocator_warmpool_size(256, 10);
}
// For allocation-heavy patterns, reuse buffers:
if (SEQUENCE == [256, 256, 512, 256, 256, 512]) {
// Detected repeating pattern - cache allocations
allocator_pattern_cache(pattern);
}
Expected Performance Impact:
- Eliminate allocator overhead for pre-warmed sizes
- Reduce fragmentation via pattern-aware pooling
- Expected gain: 1.3–1.8× for allocation-heavy programs
Implementation Complexity: Medium (requires allocator integration)
Verification: Dynamic (observable via allocation metrics)
Opportunity #5: Control Flow Branch Prediction
Problem
IF/THEN/ELSE branches have no prediction. In loops, the same branch is taken repeatedly (predictable).
Current:
: SEARCH ( n -- ? )
BEGIN
DUP 10 <
IF ... ELSE ... THEN
UNTIL
Every iteration re-evaluates the condition and potentially mis-speculates.
Solution: Execution Heat-Tracked Branch Patterns
Metrics to Track:
- Track outcomes of
IF/THEN/ELSEat each location branch_heat[location][taken](how often does branch X go True vs. False?)branch_sequence(what's the typical pattern?)
Decision Rule:
// For hot branches, set a prediction hint:
if (branch_heat[location][TRUE] > 90 &&
branch_heat[location][FALSE] < 10) {
// Mark this branch as "likely true"
vm->current_executing_entry->flags |= BRANCH_LIKELY_TRUE;
}
// Inner interpreter can use hint for:
// - CPU branch prediction (x86 PREFETCHNTA hints)
// - Reordering of code paths
// - Speculative stack preparation
Expected Performance Impact:
- Reduce branch mispredictions in CPU pipeline
- Reorder code for better CPU cache behavior
- Expected gain: 1.1–1.3× (CPU-dependent)
Implementation Complexity: Medium (requires CPU-specific assembly work)
Verification: Dynamic (observable via perf branch-misses)
Opportunity #6: Vocabulary Search Path Optimization
Problem
Multi-vocabulary systems (FORTH, EDITOR, SYSTEM) search vocabularies in order. If most lookups hit FORTH, other vocabularies waste cycles.
Current: Always search: EDITOR → SYSTEM → FORTH
Solution: Dynamic Search Order Based on Execution Heat
Metrics to Track:
vocab_hit_rate[EDITOR],vocab_hit_rate[SYSTEM],vocab_hit_rate[FORTH]vocab_lookup_time[VOCAB](latency per vocabulary)- Track which vocabulary actually had the hit
Decision Rule:
// Reorder search path based on hit rates:
// If FORTH has 85% hits, search it first
sorted_vocabs = sort_by_hit_rate([EDITOR, SYSTEM, FORTH]);
// Update search order dynamically:
if (hit_rate[FORTH] > hit_rate[EDITOR]) {
swap_search_order(FORTH, EDITOR);
}
Expected Performance Impact:
- Reduce dictionary lookup depth for hot vocabulary
- Fewer misses on wrong vocabularies
- Expected gain: 1.2–1.6× (vocab-dependent)
Implementation Complexity: Low (requires search path reordering)
Verification: Static (observable via lookup metrics)
Opportunity #7: String Operation Batching
Problem
String operations (TYPE, EMIT, PARSE) are called frequently on small strings. Each call has interpretation overhead.
Current:
." Hello " ." World " CR
Compiles to 3 separate operations with 3 lookups each.
Solution: Compile-Time String Fusion
Metrics to Track:
- Detect consecutive string literals:
." A " ." B "pattern string_op_heat[TYPE](how often are consecutive string ops used?)- String size statistics
Decision Rule:
// During compilation, fuse consecutive string operations:
if (word_sequence == [EMIT, EMIT, EMIT] &&
all_immediate) {
// Fuse into single STRING_OUTPUT operation
// Batch output: less interpreter overhead
emit_fused_string_output(concatenated_string);
}
Expected Performance Impact:
- Eliminate overhead for small string operations
- Improve I/O efficiency via batching
- Expected gain: 1.1–1.4× (I/O bound dependent)
Implementation Complexity: Low (compile-time pattern detection)
Verification: Static (observable via string operation metrics)
Opportunity #8: Arithmetic Operation Reordering & Associativity
Problem
Commutative operations (+ - * AND OR) are executed in fixed order. Some orders have better CPU cache behavior.
Example:
A B + C + D + \ Left associative
A B C D + + + \ Right associative (different cache behavior)
Solution: Heat-Tracked Operand Ordering
Metrics to Track:
operand_heat[RESULT](how hot is the result of this operation?)- Track which operand sizes are common
- CPU cache line utilization for different orderings
Decision Rule:
// For hot commutative operations, try both orderings:
if (add_heat[location] > 1000 &&
operand_size_a != operand_size_b) {
// Try reordering for better cache behavior
// A=large, B=small: load small first (better prefetch)
if (size_b < size_a) {
emit_operation(B, A, ADD); // Reordered
}
}
Expected Performance Impact:
- Improve CPU cache line reuse
- Reduce memory latency for hot arithmetic
- Expected gain: 1.05–1.15× (CPU-dependent)
Implementation Complexity: Medium (requires cache analysis)
Verification: Dynamic (observable via perf cache-misses)
Opportunity #9: Word Placement Optimization (Code Layout)
Problem
Dictionary layout is insertion-order (no optimization). Hot words might be far from each other, causing CPU cache misses.
Current: Words stored in linked-list order they were defined
Solution: Heat-Driven Memory Compaction
Metrics to Track:
word->execution_heat(already tracked!)word->co_execution_heat[OTHER_WORD](how often are these words called together?)word->memory_address(track actual memory layout)
Decision Rule:
// During garbage collection / compaction:
// 1. Calculate which words are frequently called together
// 2. Place hot word clusters adjacent in memory
// 3. Minimize cache line misses for hot word sequences
hot_cluster = [EXIT, LIT, CR] // Frequently co-executed
memory_compact_together(hot_cluster); // Colocate in memory
// Result: prefetch-friendly dictionary layout
Expected Performance Impact:
- Reduce CPU instruction cache misses
- Better temporal locality for hot code paths
- Expected gain: 1.05–1.2× (CPU architecture-dependent)
Implementation Complexity: High (requires memory layout changes, GC integration)
Verification: Dynamic (observable via perf icache-misses)
Implementation Roadmap
Phase 1: Foundation (Already Complete ✅)
- ✅ Execute heat tracking in VM
- ✅ Hot-words cache (1.78× speedup proven)
- ✅ Bayesian inference for confidence
- ✅ Q48.16 fixed-point arithmetic
Phase 2: Low-Hanging Fruit (Recommended Next)
- Stack Operation Fusion (#3) - Highest ROI, lowest complexity
- Vocabulary Search Reordering (#6) - Minimal complexity
- Return Stack Prediction (#1) - Medium complexity, high impact
Phase 3: Medium Complexity (Future)
- Block I/O Prefetching (#2) - Medium complexity
- String Operation Batching (#7) - Low/medium complexity
- Memory Allocation Patterns (#4) - Moderate complexity
Phase 4: Advanced (Future)
- Control Flow Branch Prediction (#5) - High complexity, CPU-specific
- Arithmetic Operation Reordering (#8) - Medium complexity
- Word Placement Optimization (#9) - High complexity (requires GC)
Unified Metrics Infrastructure
All 9 opportunities require a common metrics collection framework:
typedef struct {
// Execution frequency (already have this)
uint64_t execution_heat;
// Co-execution patterns (NEW)
uint64_t co_execution_heat[DICT_SIZE];
// Timing information (ENHANCED)
uint64_t total_time_ns;
uint64_t call_count;
// Cache behavior (NEW)
uint32_t cache_line_hits;
uint32_t cache_line_misses;
// Sequence patterns (NEW)
circular_buffer_t recent_execution;
} PhysicsMetrics;
One-time investment in unified metrics → all 9 opportunities unlocked.
Expected Cumulative Performance Gain
Conservative Estimate
Baseline: 1.0×
+ Hot-words cache: 1.78×
+ Stack fusion: 2.13× (1.78 × 1.2)
+ Vocab reordering: 2.56× (2.13 × 1.2)
+ Return prediction: 3.30× (2.56 × 1.3)
──────────────────────────────
Total (Phase 2): 3.3× speedup
Ambitious Estimate (All 9)
Cumulative with all optimizations: 5–8× speedup possible
But Remember
- Not all programs benefit from all optimizations
- Each optimization has CPU/workload dependency
- Diminishing returns apply (law of diminishing returns)
- 1.78× from cache alone is already substantial!
Why Physics-Driven Beats Alternatives
| Approach | Speedup | Verification | L4Re Compatible | Complexity |
|---|---|---|---|---|
| JIT Compilation | 30–100× | ❌ Impossible | ❌ No | Very High |
| Static Optimization | 1.05–1.2× | ✅ Easy | ✅ Yes | Medium |
| Physics-Driven (current) | 1.78× | ✅ Easy | ✅ Yes | Low–Medium |
| Physics-Driven (all 9) | 5–8× | ✅ Easy | ✅ Yes | Medium–High |
Verdict: Physics-driven is optimal for StarForth's constraints. We get measurable gains (1.78×+), complete verifiability, and L4Re compatibility all in one package.
References
- Experiment Report:
docs/PHYSICS_HOTWORDS_CACHE_EXPERIMENT.md - Reproduction Guide:
docs/REPRODUCE_PHYSICS_EXPERIMENT.md - Physics Architecture:
include/physics_metadata.h,include/physics_hotwords_cache.h - CLAUDE.md: Core design patterns
Summary: Strategic Direction
- Physics-Driven Optimization is the right approach for StarForth
- Hot-words cache (1.78×) is proven and ready for production
- 9 additional opportunities can stack to 5–8× total gain
- Low-hanging fruit (stack fusion, vocab reordering) should come next
- JIT is wrong direction (violates constraints, harder to verify)
- Incremental approach matches StarForth's design philosophy
The physics model is the foundation. Build on it, measure everything, verify rigorously.