Files
LithosAnanake/docs/working/architecture/03-architecture/physics-engine/feedback-loops-analysis.md
T

14 KiB
Raw Blame History

StarForth Feedback Loops & Self-Optimization Analysis

Executive Summary

StarForth implements 6 feedback loops for automatic self-tuning during execution. Of these:

  • 3 fully wired AND operational (actively driving optimization decisions)
    1. Hot-words cache promotion (1.78× speedup measured)
    2. Rolling window adaptive shrinking (FIXED 2025-11-08, enforcement now works)
    3. Pipelining transition speculation (WIRED 2025-11-08, speculative prefetch active)
  • 1 wired AND utilized, but NOT validated ⚠️ (running, but unsupported design choices) 3. ⚠️ Heat decay of execution history (arbitrary linear function, no tuning knob, unvalidated)
  • 1 wired but NOT utilized (Phase 2 experimental, infrastructure ready) 5. Context-aware window tuning (binary chop algorithm stub, needs VM-level metrics aggregation)
  • 1 observability-only (metrics collected for formal verification) 6. Physics metadata temperature tracking

CURRENT STATUS:

  • A_BASELINE, A_B_CACHE, A_C_FULL, A_B_C_FULL configurations now test complete optimization stack:
    • Loop #1 (cache) enabled via A_B_CACHE flag
    • Loop #2 (window shrinking) always enabled, active in all configs
    • Loop #3 (decay) always enabled, active in all configs
    • Loop #4 (speculation) disabled by default (ENABLE_PIPELINING=0), ready for A_C_FULL testing
  • BUT: Loop #3 (heat decay) lacks principled design—slope is arbitrary, function choice unvalidated
  • Gap: No measurements showing decay actually helps; DECAY_RATE has NO DoE knob to tune it

FULLY WIRED & UTILIZED FEEDBACK LOOPS

1. Hot-Words Cache Promotion Feedback Loop DONE

What's Measured:

  • execution_heat counter on each dictionary entry (incremented per execution)

Trigger Condition:

  • When execution_heat > HOTWORDS_EXECUTION_HEAT_THRESHOLD (default: 50)

Automatic Action:

  • hotwords_cache_promote() moves word to 32-entry LRU cache
  • Subsequent lookups hit cache (1-2 ns vs 20-30 ns in bucket)

Code Locations:

  • Metric tracking: src/dictionary_management.c:170
  • Threshold check: src/physics_hotwords_cache.c:155
  • Promotion logic: src/physics_hotwords_cache.c:420-450

Observed Impact:

  • 1.78× speedup on cache-hit paths (measured in DoE phase 1)
  • Status: FULLY OPERATIONAL

2. Rolling Window Adaptive Shrinking Feedback Loop FIXED

What's Measured:

  • Pattern diversity: count of unique word transitions in execution history

Measurement Function:

// src/rolling_window_of_truth.c:546-577 (FIXED)
uint64_t rolling_window_measure_diversity(const RollingWindowOfTruth* window)
{
    // Count unique adjacent transitions (word_a → word_b)
    // NOW RESPECTS effective_window_size - only scans recent entries
}

Trigger Condition:

  • When pattern diversity growth rate < ADAPTIVE_GROWTH_THRESHOLD (default: 1%)
  • Checked every ADAPTIVE_CHECK_FREQUENCY executions (default: 256)

Automatic Action:

  • Shrink effective_window_size to ADAPTIVE_SHRINK_RATE percent of current (default: 75%)
  • Window shrinks: 4096 → 3072 → 2304 → 1728 → ... → ADAPTIVE_MIN_WINDOW_SIZE (256)

Code Locations:

  • Check function: src/rolling_window_of_truth.c:589-645
  • Shrink calculation: Line 626
  • Knobs: include/rolling_window_knobs.h

Integration:

  • Called from rolling_window_record_execution() on every word execution
  • Status: FULLY OPERATIONAL

Fix Applied (2025-11-08):

  • Line 559: Now limits scan to effective_window_size when window is warm
  • Lines 561-564: Circular indexing ensures only recent entries within effective window are scanned
  • Effect: New data arrives every tick, old data falls off after effective_window_size ticks (was: always 4096)

3. Linear Decay of Execution Heat Feedback Loop

What's Measured:

  • execution_heat value per word (time-sensitive metric)

Decay Formula (INTENTIONALLY BASELINE):

  • H(t) = max(0, H_0 - decay_rate × Δt)LINEAR decay
  • Configurable decay rate: DECAY_RATE_PER_US_Q16 (default: 1 = 1/65536 heat/µs)
  • Derived half-life: ~6-7 seconds for 100-heat word

Automatic Action:

  • Stale words automatically lose heat weight
  • Only frequently-executed words stay above promotion threshold

Code Locations:

  • Decay logic: src/physics_metadata.c:164-203 (physics_metadata_apply_linear_decay())
  • Integration: src/vm.c:524 (called before each execution in inner loop)

Design Rationale (From Phase 2 Specification):

  • Chose LINEAR over exponential for:
    1. Performance (integer arithmetic, no floating-point)
    2. Predictability (bounded convergence time)
    3. Simplicity (fewer lines, easier to verify formally)
    4. StarshipOS applicability (models "task context switch forces heat reset")
  • Explicitly noted as baseline: "Can always switch to halflife after validating linear approach"

STATUS: WIRED & UTILIZED, BUT NOT VALIDATED ⚠️

CRITICAL GAPS:

  1. No empirical validation: Does linear decay actually help? No measurements.
  2. Arbitrary initial slope: Why 1 µs⁻¹? No evidence this is optimal.
  3. Wrong function?: Could exponential, log, parabolic, or sinusoidal be better?
  4. Tuning blind: DECAY_RATE_PER_US_Q16 has NO KNOB in DoE configurations
  5. No feedback loop closure: Decay happens, but no measurement of whether it improved anything

IMPLEMENTATION FIXES APPLIED

Rolling Window Shrinking - FIXED

The Bug (Now Fixed): The adaptive shrinking mechanism was shrinking the reported metric without enforcing the limit.

Old Code (Line 557):

for (uint32_t i = 0; i < ROLLING_WINDOW_SIZE; i++)  // ← ALWAYS full size (4096)
{
    uint32_t current = window->execution_history[i];
    // Count unique transitions...
}

Fixed Code (Lines 559-574):

uint32_t scan_limit = (window->is_warm) ? window->effective_window_size : ROLLING_WINDOW_SIZE;

for (uint32_t i = 0; i < scan_limit; i++)
{
    /* Index into the recent window (window_pos is the next write position) */
    uint32_t idx = (window->window_pos + ROLLING_WINDOW_SIZE - scan_limit + i) % ROLLING_WINDOW_SIZE;
    uint32_t current = window->execution_history[idx];
    // Count unique transitions in recent entries only
}

What Changed:

  • Now limits scan to effective_window_size when window is warm
  • Uses circular indexing to ensure only RECENT entries are scanned
  • New data arrives every tick, old data falls off after effective_window_size ticks

User's Requirement Met:

"The window should advance with every tick so new data comes in and stale data falls off."

Now data falls off after effective_window_size ticks, not always 4096


4. Pipelining Transition Metrics NOW FUNCTIONAL

What's Measured:

  • Transition patterns: (word_A → word_B) frequencies via transition_metrics_record()
  • Probability calculations: P(B | A) in Q48.16 fixed-point format
  • Prefetch accuracy: Hits vs. misses from speculative cache promotions

Trigger Conditions:

  • SPECULATION_THRESHOLD_Q48 (50% minimum confidence)
  • MIN_SAMPLES_FOR_SPECULATION (10 minimum observations)

Automatic Action (NOW WIRED):

  • After recording each transition, update probability cache
  • Check: Does next word probability exceed threshold?
  • If yes: Find dictionary entry and pre-promote to hotwords cache
  • Effect: When next word is actually looked up, it's cache-warm
  • Record: Track prefetch_attempts for feedback measurement

Code Locations:

  • Metric collection: src/physics_pipelining_metrics.c:92-109 (transition_metrics_record())
  • Probability calculation: src/physics_pipelining_metrics.c:112-123
  • Decision logic: src/physics_pipelining_metrics.c:164-186 (transition_metrics_should_speculate())
  • Action wiring: src/vm.c:544-583 (execute_colon_word, speculative prefetch logic)

Status: FULLY WIRED & FUNCTIONAL

Wiring Details (2025-11-08):

  • Line 548: transition_metrics_record() - record the transition
  • Line 551: transition_metrics_update_cache() - find most likely next word
  • Line 554-555: transition_metrics_should_speculate() - check confidence threshold
  • Lines 560-581: If confident, look up predicted word and promote to hotwords cache
  • Line 580: Record the prefetch attempt for metrics feedback

Loop Closure:

  1. Metrics collected: transition_heat array counts (word_B | word_A)
  2. Probability calculated: Q48.16 format
  3. Decision made: Confidence check against threshold
  4. Action taken: Speculative promotion to hotwords cache
  5. Feedback recorded: prefetch_attempts counter

5. Context-Aware Window Size Tuning (Phase 2 Placeholder)

What's Measured:

  • Multi-word context patterns (sequences of 2-4 consecutive words)

Intended Action (Not Connected):

  • Use binary search to find optimal effective_window_size for given workload
  • Function: transition_metrics_binary_chop_suggest_window() (lines 340-357)

Why Not Utilized:

  • Phase 2 research feature
  • Experimental algorithm not yet validated
  • No integration point in main execution path

Status: PROTOTYPE, NOT INTEGRATED


OBSERVABILITY-ONLY (Not Decision-Driving)

6. Physics Metadata Temperature Tracking

What's Measured:

  • temperature_q8: Exponential moving average of execution_heat
  • Used for observability and formal verification model tracing

Current Use:

  • Profiling and debugging (not driving optimization decisions)
  • Forms foundation for Isabelle/HOL physics state machine proofs

Status: WORKING AS INTENDED (observability, not optimization)


Bootstrap Seeding Mechanisms

These run once at startup to initialize feedback loop state:

POST (Power-On Self-Test) Seeding

  • Function: rolling_window_seed_hotwords_cache() (lines 258-329)
  • Action: Records POST execution, promotes high-heat words to cache
  • Effect: Cache is "warm" when user workload begins

Pipelining Context Seeding

  • Function: rolling_window_seed_pipelining_context() (lines 341-415)
  • Action: Records POST sequences, establishes transition baseline
  • Effect: Pipelining has initial pattern knowledge

Configuration Knobs for Feedback Loops

All knobs are tunable via Makefile:

# Hot-words cache promotion
make HOTWORDS_EXECUTION_HEAT_THRESHOLD=75

# Adaptive window shrinking
make ADAPTIVE_SHRINK_RATE=75                    # Keep 75%, shrink by 25%
make ADAPTIVE_MIN_WINDOW_SIZE=256               # Never shrink below 256
make ADAPTIVE_CHECK_FREQUENCY=256               # Check every 256 executions
make ADAPTIVE_GROWTH_THRESHOLD=1                # Shrink when growth < 1%

# Heat decay
make DECAY_RATE_PER_US_Q16=16384               # Decay rate (units: heat/μs)
make DECAY_MIN_INTERVAL=1000000               # Min interval between decays (ns)

# Pipelining (Phase 2, disabled by default)
make ENABLE_PIPELINING=1
make SPECULATION_THRESHOLD_Q48=$((50 << 16))  # 50% confidence
make MIN_SAMPLES_FOR_SPECULATION=10

Summary Table: All Feedback Loops

Loop Measured Threshold Action Wired? Utilized? Status
1. Hot-words promotion execution_heat > 50 Add to cache VALIDATED
2. Adaptive shrinking pattern_diversity growth < 1% Shrink window FIXED
3. Heat decay Time since last execution ~6-7s half-life Reduce heat ⚠️ NOT VALIDATED
4. Pipelining transitions transition_heat[] > 50% confidence Speculative prefetch Phase 2
5. Context window tuning Multi-word sequences N/A Binary chop Phase 2
6. Temperature tracking temperature_q8 N/A Observability only OBSERVABILITY

Completed Actions

DONE: Fix Rolling Window Shrinking Implementation

Status: APPLIED (2025-11-08)

Modified src/rolling_window_of_truth.c:546-577 (function rolling_window_measure_diversity()):

  • Now respects effective_window_size when window is warm
  • Uses circular indexing to scan only recent entries
  • New data arrives every tick, old data falls off after effective_window_size ticks (not always 4096)
  • A_B_C_FULL configuration now tests both hot-words cache + adaptive window shrinking together

Next Steps for Phase 2 (If Proceeding)

OPTIONAL 1: Connect Pipelining Decision Logic

When ENABLE_PIPELINING=1:

  • Call transition_metrics_should_speculate() from vm inner loop
  • If returns true: Pre-load next word from context prediction
  • Measure cache miss reduction vs. A_BASELINE (Phase 2 experimental validation)

Status: Metrics collection ready, decision logic not wired

OPTIONAL 2: Validate Context Window Tuning (Phase 2)

For adaptive window size optimization:

  • Implement binary chop search algorithm in main execution path
  • Tune effective_window_size dynamically based on context complexity
  • Measure pattern capture efficiency vs. memory overhead

Status: Binary chop function exists, integration point not defined


Files & Line Numbers Reference

Purpose File Lines
Rolling window structure include/vm.h 58-70
Rolling window implementation src/rolling_window_of_truth.c -
Adaptive shrinking src/rolling_window_of_truth.c 583-639
Bug location src/rolling_window_of_truth.c 557
Diversity measurement src/rolling_window_of_truth.c 546-571
Shrinking knobs include/rolling_window_knobs.h -
Hot-words cache src/physics_hotwords_cache.c -
Promotion threshold src/physics_hotwords_cache.c 155
Heat decay src/physics_metadata.c 164-203
Pipelining metrics src/physics_pipelining_metrics.c -
Pipelining decision src/physics_pipelining_metrics.c 164-186
Main integration src/vm.c 524, 542-548

Analysis Complete: 2025-11-08 User Request: Investigate feedback loops, identify implementation problems Finding: 3 active + 2 dormant + 1 observability feedback loops; 1 critical bug in rolling window shrinking