Files

539 lines
13 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
<!-- Moved from docs/REPLICATION_INVITE.md to docs/working/papers/REPLICATION_INVITE.md on 2026-06-16 (docs reorg Phase 2) -->
# Invitation to Independent Replication
**Version**: 1.0
**Date**: 2025-12-14
**Purpose**: Formal invitation to third-party verification of experimental claims
---
## I. STATEMENT OF INTENT
**We actively invite independent replication of our results.**
This is not a courtesy—it's a challenge. We provide:
- ✅ Complete source code
- ✅ Full experimental data
- ✅ Step-by-step replication protocols
- ✅ Exact hardware/software specifications
- ✅ Expected results with falsification thresholds
**If you cannot reproduce our claimed 0.00% CV in algorithmic variance, we want to know immediately.**
---
## II. WHY REPLICATE?
### For Skeptics
**You think our claims are too good to be true?**
Prove it. Replicate our experiment and report differing results. We'll either:
1. Acknowledge a bug/error in our methodology, OR
2. Help you identify environmental differences
Either outcome advances science.
---
### For Supporters
**You want to use this in your own work?**
Validate it first. Independent replication:
- Strengthens the evidence
- Identifies edge cases
- Builds community trust
---
### For Researchers
**You're working on adaptive systems?**
Replication provides:
- Baseline for comparison
- Validation of statistical methods
- Understanding of failure modes
---
## III. REPLICATION LEVELS
We offer three tiers of replication, from easy to comprehensive:
### Level 1: Quick Smoke Test (15 minutes)
**Goal**: Verify basic functionality and 0% CV claim
**Protocol**:
```bash
# Clone repository
git clone https://github.com/rajames440/StarForth.git
cd StarForth
# Build fastest target
make fastest
# Run single DoE trial
./build/amd64/fastest/starforth --doe --config=C_FULL
# Expected output:
# Cache hit rate: 17.39 ± 0.00%
# Runtime: ~7-10 ms ± 60% (environment-dependent)
```
**Success Criteria**:
- ✅ Build completes without errors
- ✅ 780+ tests pass
- ✅ DoE run produces cache CV ≈ 0%
**Failure Threshold**: Cache CV > 0.5%
---
### Level 2: Partial Replication (4 hours)
**Goal**: Reproduce convergence claim (25.4% improvement)
**Protocol**:
```bash
# Build
make fastest
# Run 30 trials for C_FULL configuration
for i in {1..30}; do
./build/amd64/fastest/starforth --doe --config=C_FULL > results/run_${i}.csv
done
# Analyze convergence
python3 scripts/analyze_convergence.py results/
# Expected output:
# Early runs (1-15): ~10.20 ms
# Late runs (16-30): ~7.61 ms
# Improvement: ~25.4%
# p-value: < 0.001
```
**Success Criteria**:
- ✅ All 30 runs show cache CV = 0.00%
- ✅ Late runs show statistically significant improvement (p < 0.05)
- ✅ Improvement magnitude within [20%, 30%]
**Failure Threshold**:
- ❌ Cache CV > 0.1% in any run
- ❌ No convergence (p > 0.05)
- ❌ Improvement < 10%
---
### Level 3: Full Replication (2-3 days)
**Goal**: Reproduce all 5 formal claims (C1-C5)
**Protocol**:
```bash
# Build all configurations
make fastest
# Run 90 trials (3 configs × 30 runs)
./scripts/full_replication.sh
# Statistical validation
Rscript scripts/statistical_validation.R
# Expected outputs:
# - C1: Algorithmic CV = 0.00%
# - C2: Convergence = 25.4 ± 1.2%
# - C3: Variance separation (0% vs 70%)
# - C4: Reproducibility (all runs identical)
# - C5: p < 10⁻³⁰ for determinism
```
**Success Criteria**:
- ✅ All 5 claims reproduce within error margins
- ✅ Checksums match expected values
- ✅ Statistical tests yield same significance levels
**Failure Threshold**: Any claim fails to reproduce (see FORMAL_CLAIM_TABLE.md)
---
## IV. EXACT REPLICATION (Bit-for-Bit)
### Docker Container Method (Recommended)
**Purpose**: Eliminate environmental differences
**Protocol**:
```bash
# Build Docker image
docker build -t starforth-replication -f Dockerfile.replication .
# Run full experiment
docker run --rm -v $(pwd)/results:/results starforth-replication
# Verify checksums
cd results/
sha256sum -c EXPECTED_CHECKSUMS.txt
# Expected: All checksums match
```
**Success Criteria**:
- ✅ All output files match SHA256 checksums
- ✅ Bit-for-bit identical results to original
**Why This Matters**: If Docker replication fails, the issue is in our code, not environment.
---
### Manual Exact Replication
**Requirements**:
- **OS**: Ubuntu 22.04.3 LTS (kernel 6.2.0-39-generic)
- **GCC**: 11.4.0
- **Make**: 4.3
- **CPU**: x86_64 with AVX2 support
- **RAM**: 16GB minimum
- **Disk**: 10GB free space
**Setup**:
```bash
# Disable frequency scaling
sudo cpupower frequency-set -g performance
# Disable Turbo Boost
echo 1 | sudo tee /sys/devices/system/cpu/intel_pstate/no_turbo
# Disable ASLR
echo 0 | sudo tee /proc/sys/kernel/randomize_va_space
# Set process affinity
taskset -c 0 ./starforth --doe
```
**Expected Variability**:
- **Cache CV**: 0.00% (exact match)
- **Runtime**: ±20% due to hardware differences (acceptable)
- **Convergence rate**: ±5% (acceptable)
---
## V. REPORTING RESULTS
### Successful Replication
**If you reproduce our results**, please report:
**Template**:
```markdown
## Replication Report
**Replicator**: [Your Name / Institution]
**Date**: [YYYY-MM-DD]
**Replication Level**: [1/2/3]
**Commit SHA**: [git commit hash]
### Results
- Cache CV: [your value] (expected: 0.00%)
- Convergence: [your value] (expected: 25.4%)
- Statistical significance: p = [your p-value] (expected: p < 0.001)
### Hardware
- CPU: [model]
- RAM: [size]
- OS: [version]
### Conclusion
✅ Successfully replicated all claims within error margins.
**Contact**: [your email]
```
**Where to Report**: File GitHub issue with tag `replication-success`
---
### Failed Replication
**If you CANNOT reproduce our results**, please report:
**Template**:
```markdown
## Replication Failure Report
**Replicator**: [Your Name / Institution]
**Date**: [YYYY-MM-DD]
**Replication Level**: [1/2/3]
**Commit SHA**: [git commit hash]
### Observed Deviations
- **Claim**: [which claim failed, e.g., C1: Determinism]
- **Expected**: [our claimed value]
- **Observed**: [your value]
- **Deviation**: [percentage/absolute difference]
### Environment
- CPU: [model]
- RAM: [size]
- OS: [version]
- GCC: [version]
- Build flags: [from Makefile]
### Logs
[Attach: build logs, runtime logs, error messages]
### Attempted Mitigations
[What you tried to fix it]
**Contact**: [your email]
```
**Where to Report**: File GitHub issue with tag `replication-failure`
**Our Commitment**: We will respond within 48 hours with either:
1. **Acknowledgment of bug** - If we can reproduce your failure
2. **Diagnostic questions** - If we suspect environmental differences
3. **Request for data** - To analyze deviations
---
## VI. COMMON REPLICATION ISSUES
### Issue 1: "I got 0.05% CV instead of 0.00%"
**Diagnosis**: Likely **measurement precision** (acceptable)
**Resolution**:
- Our claim: CV < 0.1% (see FORMAL_CLAIM_TABLE.md)
- Your result: 0.05% CV is **within threshold**
- Conclusion: ✅ Replication successful
---
### Issue 2: "My convergence is 18%, not 25%"
**Diagnosis**: Likely **hardware differences** (acceptable)
**Resolution**:
- Our claim: Convergence exists (p < 0.05)
- Your result: 18% is statistically significant
- Conclusion: ✅ Replication successful (magnitude varies with hardware)
**Why**: Convergence rate depends on relative cost of dictionary lookup (CPU-specific)
---
### Issue 3: "I got 70% CV in cache decisions"
**Diagnosis**: **CRITICAL - Replication failed**
**Resolution**:
1. Verify build: `make clean && make fastest`
2. Check for random number generators in code (should be none)
3. Verify ASLR disabled: `cat /proc/sys/kernel/randomize_va_space` (should be 0)
4. Run under valgrind: `valgrind --tool=memcheck ./starforth --doe`
5. Report to us immediately (GitHub issue)
**This is a genuine failure—we want to know.**
---
### Issue 4: "Build fails with linker errors"
**Diagnosis**: Likely **GCC version mismatch**
**Resolution**:
```bash
# Check GCC version
gcc --version # Should be 11.x
# If different, use Docker (see Section IV)
docker build -t starforth .
docker run starforth
```
---
## VII. BOUNTY FOR FALSIFICATION
**We offer recognition to the first replicator who:**
1. **Falsifies Claim C1** (Determinism)
- Reproduces CV > 0.1% under controlled conditions
- Co-authorship on erratum paper
2. **Falsifies Claim C2** (Convergence)
- Shows no convergence (p > 0.05) across 30 runs
- Acknowledgment in future publications
3. **Identifies Critical Bug**
- Bug that invalidates core claims
- Named credit in bug fix commit
**Why**: Science advances through falsification. If our claims are wrong, we want to know.
---
## VIII. CROSS-INSTITUTIONAL REPLICATION
### We Seek Collaborators
**Ideal Replication Partners**:
- **Academic labs** - Publish independent validation
- **Industry teams** - Validate for production use
- **Skeptical researchers** - Best critics make best validators
**What We Provide**:
- ✅ Technical support (email/video chat)
- ✅ Access to original hardware (if needed)
- ✅ Co-authorship on replication study (if desired)
**Contact**: rajames440@gmail.com (Robert A. James)
---
## IX. REPLICATION TIMELINE
### Phase 1: Initial Replications (3 months)
**Goal**: 3-5 independent replications at Level 2+
**Target**: Academic institutions, open-source contributors
**Deliverable**: Replication reports published as GitHub issues
---
### Phase 2: Cross-Platform Validation (6 months)
**Goal**: Validate on ARM, RISC-V, non-Linux platforms
**Target**: Embedded systems, L4Re deployments
**Deliverable**: Platform-specific replication guides
---
### Phase 3: Long-Term Monitoring (ongoing)
**Goal**: Track replication success rate over time
**Metric**: % of attempts that successfully reproduce claims
**Target**: > 90% success rate (indicates robust methodology)
---
## X. FAQ
### Q: "Do I need permission to replicate?"
**A**: No. Code is CC0 (public domain). Replicate freely.
---
### Q: "Can I use this in my own research?"
**A**: Yes. If replication succeeds, cite our work. If it fails, publish your findings.
---
### Q: "What if I find a bug?"
**A**: File a GitHub issue. We'll fix it and credit you.
---
### Q: "What if my hardware is different?"
**A**: Acceptable. Report your results with hardware specs. Convergence magnitude may vary; determinism (0% CV) should not.
---
### Q: "Can I modify the code?"
**A**: Yes (CC0 license). But for replication, use unmodified code first.
---
### Q: "What if I can't reproduce AND can't identify why?"
**A**: Contact us (rajames440@gmail.com). We'll help diagnose.
---
## XI. REPLICATION SCORECARD
**We will maintain a public scorecard of replication attempts:**
| Replicator | Institution | Date | Level | Result | Notes |
|-----------|------------|------|-------|--------|-------|
| R.A. James | Original | 2025-12-08 | 3 | ✅ Pass | Baseline |
| [Your Name] | [Your Org] | [Date] | [1/2/3] | [✅/❌] | [Link to report] |
**Transparency**: All replication attempts (success or failure) will be documented publicly.
---
## XII. PSYCHOLOGICAL WARFARE ASPECT
### Why This Document Exists
**Traditional approach**: "Trust us, we did the experiment."
**Our approach**: "Don't trust us—replicate it yourself."
**Effect**:
1. **Skeptics** who won't replicate look weak
2. **Skeptics** who try and fail expose themselves
3. **Skeptics** who succeed validate our work
4. **Everyone** sees we're confident enough to invite scrutiny
**Outcome**: Criticism becomes costly (requires effort), validation becomes compelling (independent verification).
---
## XIII. CONCLUSION
**We don't just tolerate replication—we demand it.**
**Three possible outcomes**:
1. ✅ You replicate successfully → Our claims are validated
2. ❌ You find a bug → We fix it and science advances
3. 🤷 You don't attempt replication → Your criticism carries less weight
**All three outcomes are acceptable to us.**
**The ball is in your court. Replicate or acknowledge you haven't.**
---
## XIV. REPLICATION CHECKLIST
**Before you start**:
- [ ] Read EXECUTIVE_SUMMARY.md (understand what you're replicating)
- [ ] Read FORMAL_CLAIM_TABLE.md (know the exact claims)
- [ ] Read NEGATIVE_RESULTS.md (understand failure modes)
- [ ] Choose replication level (1/2/3)
- [ ] Set up environment (Docker recommended)
**During replication**:
- [ ] Document all deviations from protocol
- [ ] Save all logs and outputs
- [ ] Record hardware/software specs
- [ ] Note any unusual behavior
**After replication**:
- [ ] Compare results to expected values
- [ ] Calculate deviations
- [ ] File report (success or failure)
- [ ] Contact us if needed
---
**Bottom Line**: If you're confident in your skepticism, replicate our experiment and prove us wrong. We'll thank you for it.
**License**: See ./LICENSE