539 lines
13 KiB
Markdown
539 lines
13 KiB
Markdown
<!-- Moved from docs/REPLICATION_INVITE.md to docs/working/papers/REPLICATION_INVITE.md on 2026-06-16 (docs reorg Phase 2) -->
|
||
# Invitation to Independent Replication
|
||
|
||
**Version**: 1.0
|
||
**Date**: 2025-12-14
|
||
**Purpose**: Formal invitation to third-party verification of experimental claims
|
||
|
||
---
|
||
|
||
## I. STATEMENT OF INTENT
|
||
|
||
**We actively invite independent replication of our results.**
|
||
|
||
This is not a courtesy—it's a challenge. We provide:
|
||
- ✅ Complete source code
|
||
- ✅ Full experimental data
|
||
- ✅ Step-by-step replication protocols
|
||
- ✅ Exact hardware/software specifications
|
||
- ✅ Expected results with falsification thresholds
|
||
|
||
**If you cannot reproduce our claimed 0.00% CV in algorithmic variance, we want to know immediately.**
|
||
|
||
---
|
||
|
||
## II. WHY REPLICATE?
|
||
|
||
### For Skeptics
|
||
|
||
**You think our claims are too good to be true?**
|
||
|
||
Prove it. Replicate our experiment and report differing results. We'll either:
|
||
1. Acknowledge a bug/error in our methodology, OR
|
||
2. Help you identify environmental differences
|
||
|
||
Either outcome advances science.
|
||
|
||
---
|
||
|
||
### For Supporters
|
||
|
||
**You want to use this in your own work?**
|
||
|
||
Validate it first. Independent replication:
|
||
- Strengthens the evidence
|
||
- Identifies edge cases
|
||
- Builds community trust
|
||
|
||
---
|
||
|
||
### For Researchers
|
||
|
||
**You're working on adaptive systems?**
|
||
|
||
Replication provides:
|
||
- Baseline for comparison
|
||
- Validation of statistical methods
|
||
- Understanding of failure modes
|
||
|
||
---
|
||
|
||
## III. REPLICATION LEVELS
|
||
|
||
We offer three tiers of replication, from easy to comprehensive:
|
||
|
||
### Level 1: Quick Smoke Test (15 minutes)
|
||
|
||
**Goal**: Verify basic functionality and 0% CV claim
|
||
|
||
**Protocol**:
|
||
```bash
|
||
# Clone repository
|
||
git clone https://github.com/rajames440/StarForth.git
|
||
cd StarForth
|
||
|
||
# Build fastest target
|
||
make fastest
|
||
|
||
# Run single DoE trial
|
||
./build/amd64/fastest/starforth --doe --config=C_FULL
|
||
|
||
# Expected output:
|
||
# Cache hit rate: 17.39 ± 0.00%
|
||
# Runtime: ~7-10 ms ± 60% (environment-dependent)
|
||
```
|
||
|
||
**Success Criteria**:
|
||
- ✅ Build completes without errors
|
||
- ✅ 780+ tests pass
|
||
- ✅ DoE run produces cache CV ≈ 0%
|
||
|
||
**Failure Threshold**: Cache CV > 0.5%
|
||
|
||
---
|
||
|
||
### Level 2: Partial Replication (4 hours)
|
||
|
||
**Goal**: Reproduce convergence claim (25.4% improvement)
|
||
|
||
**Protocol**:
|
||
```bash
|
||
# Build
|
||
make fastest
|
||
|
||
# Run 30 trials for C_FULL configuration
|
||
for i in {1..30}; do
|
||
./build/amd64/fastest/starforth --doe --config=C_FULL > results/run_${i}.csv
|
||
done
|
||
|
||
# Analyze convergence
|
||
python3 scripts/analyze_convergence.py results/
|
||
|
||
# Expected output:
|
||
# Early runs (1-15): ~10.20 ms
|
||
# Late runs (16-30): ~7.61 ms
|
||
# Improvement: ~25.4%
|
||
# p-value: < 0.001
|
||
```
|
||
|
||
**Success Criteria**:
|
||
- ✅ All 30 runs show cache CV = 0.00%
|
||
- ✅ Late runs show statistically significant improvement (p < 0.05)
|
||
- ✅ Improvement magnitude within [20%, 30%]
|
||
|
||
**Failure Threshold**:
|
||
- ❌ Cache CV > 0.1% in any run
|
||
- ❌ No convergence (p > 0.05)
|
||
- ❌ Improvement < 10%
|
||
|
||
---
|
||
|
||
### Level 3: Full Replication (2-3 days)
|
||
|
||
**Goal**: Reproduce all 5 formal claims (C1-C5)
|
||
|
||
**Protocol**:
|
||
```bash
|
||
# Build all configurations
|
||
make fastest
|
||
|
||
# Run 90 trials (3 configs × 30 runs)
|
||
./scripts/full_replication.sh
|
||
|
||
# Statistical validation
|
||
Rscript scripts/statistical_validation.R
|
||
|
||
# Expected outputs:
|
||
# - C1: Algorithmic CV = 0.00%
|
||
# - C2: Convergence = 25.4 ± 1.2%
|
||
# - C3: Variance separation (0% vs 70%)
|
||
# - C4: Reproducibility (all runs identical)
|
||
# - C5: p < 10⁻³⁰ for determinism
|
||
```
|
||
|
||
**Success Criteria**:
|
||
- ✅ All 5 claims reproduce within error margins
|
||
- ✅ Checksums match expected values
|
||
- ✅ Statistical tests yield same significance levels
|
||
|
||
**Failure Threshold**: Any claim fails to reproduce (see FORMAL_CLAIM_TABLE.md)
|
||
|
||
---
|
||
|
||
## IV. EXACT REPLICATION (Bit-for-Bit)
|
||
|
||
### Docker Container Method (Recommended)
|
||
|
||
**Purpose**: Eliminate environmental differences
|
||
|
||
**Protocol**:
|
||
```bash
|
||
# Build Docker image
|
||
docker build -t starforth-replication -f Dockerfile.replication .
|
||
|
||
# Run full experiment
|
||
docker run --rm -v $(pwd)/results:/results starforth-replication
|
||
|
||
# Verify checksums
|
||
cd results/
|
||
sha256sum -c EXPECTED_CHECKSUMS.txt
|
||
|
||
# Expected: All checksums match
|
||
```
|
||
|
||
**Success Criteria**:
|
||
- ✅ All output files match SHA256 checksums
|
||
- ✅ Bit-for-bit identical results to original
|
||
|
||
**Why This Matters**: If Docker replication fails, the issue is in our code, not environment.
|
||
|
||
---
|
||
|
||
### Manual Exact Replication
|
||
|
||
**Requirements**:
|
||
- **OS**: Ubuntu 22.04.3 LTS (kernel 6.2.0-39-generic)
|
||
- **GCC**: 11.4.0
|
||
- **Make**: 4.3
|
||
- **CPU**: x86_64 with AVX2 support
|
||
- **RAM**: 16GB minimum
|
||
- **Disk**: 10GB free space
|
||
|
||
**Setup**:
|
||
```bash
|
||
# Disable frequency scaling
|
||
sudo cpupower frequency-set -g performance
|
||
|
||
# Disable Turbo Boost
|
||
echo 1 | sudo tee /sys/devices/system/cpu/intel_pstate/no_turbo
|
||
|
||
# Disable ASLR
|
||
echo 0 | sudo tee /proc/sys/kernel/randomize_va_space
|
||
|
||
# Set process affinity
|
||
taskset -c 0 ./starforth --doe
|
||
```
|
||
|
||
**Expected Variability**:
|
||
- **Cache CV**: 0.00% (exact match)
|
||
- **Runtime**: ±20% due to hardware differences (acceptable)
|
||
- **Convergence rate**: ±5% (acceptable)
|
||
|
||
---
|
||
|
||
## V. REPORTING RESULTS
|
||
|
||
### Successful Replication
|
||
|
||
**If you reproduce our results**, please report:
|
||
|
||
**Template**:
|
||
```markdown
|
||
## Replication Report
|
||
|
||
**Replicator**: [Your Name / Institution]
|
||
**Date**: [YYYY-MM-DD]
|
||
**Replication Level**: [1/2/3]
|
||
**Commit SHA**: [git commit hash]
|
||
|
||
### Results
|
||
- Cache CV: [your value] (expected: 0.00%)
|
||
- Convergence: [your value] (expected: 25.4%)
|
||
- Statistical significance: p = [your p-value] (expected: p < 0.001)
|
||
|
||
### Hardware
|
||
- CPU: [model]
|
||
- RAM: [size]
|
||
- OS: [version]
|
||
|
||
### Conclusion
|
||
✅ Successfully replicated all claims within error margins.
|
||
|
||
**Contact**: [your email]
|
||
```
|
||
|
||
**Where to Report**: File GitHub issue with tag `replication-success`
|
||
|
||
---
|
||
|
||
### Failed Replication
|
||
|
||
**If you CANNOT reproduce our results**, please report:
|
||
|
||
**Template**:
|
||
```markdown
|
||
## Replication Failure Report
|
||
|
||
**Replicator**: [Your Name / Institution]
|
||
**Date**: [YYYY-MM-DD]
|
||
**Replication Level**: [1/2/3]
|
||
**Commit SHA**: [git commit hash]
|
||
|
||
### Observed Deviations
|
||
- **Claim**: [which claim failed, e.g., C1: Determinism]
|
||
- **Expected**: [our claimed value]
|
||
- **Observed**: [your value]
|
||
- **Deviation**: [percentage/absolute difference]
|
||
|
||
### Environment
|
||
- CPU: [model]
|
||
- RAM: [size]
|
||
- OS: [version]
|
||
- GCC: [version]
|
||
- Build flags: [from Makefile]
|
||
|
||
### Logs
|
||
[Attach: build logs, runtime logs, error messages]
|
||
|
||
### Attempted Mitigations
|
||
[What you tried to fix it]
|
||
|
||
**Contact**: [your email]
|
||
```
|
||
|
||
**Where to Report**: File GitHub issue with tag `replication-failure`
|
||
|
||
**Our Commitment**: We will respond within 48 hours with either:
|
||
1. **Acknowledgment of bug** - If we can reproduce your failure
|
||
2. **Diagnostic questions** - If we suspect environmental differences
|
||
3. **Request for data** - To analyze deviations
|
||
|
||
---
|
||
|
||
## VI. COMMON REPLICATION ISSUES
|
||
|
||
### Issue 1: "I got 0.05% CV instead of 0.00%"
|
||
|
||
**Diagnosis**: Likely **measurement precision** (acceptable)
|
||
|
||
**Resolution**:
|
||
- Our claim: CV < 0.1% (see FORMAL_CLAIM_TABLE.md)
|
||
- Your result: 0.05% CV is **within threshold**
|
||
- Conclusion: ✅ Replication successful
|
||
|
||
---
|
||
|
||
### Issue 2: "My convergence is 18%, not 25%"
|
||
|
||
**Diagnosis**: Likely **hardware differences** (acceptable)
|
||
|
||
**Resolution**:
|
||
- Our claim: Convergence exists (p < 0.05)
|
||
- Your result: 18% is statistically significant
|
||
- Conclusion: ✅ Replication successful (magnitude varies with hardware)
|
||
|
||
**Why**: Convergence rate depends on relative cost of dictionary lookup (CPU-specific)
|
||
|
||
---
|
||
|
||
### Issue 3: "I got 70% CV in cache decisions"
|
||
|
||
**Diagnosis**: **CRITICAL - Replication failed**
|
||
|
||
**Resolution**:
|
||
1. Verify build: `make clean && make fastest`
|
||
2. Check for random number generators in code (should be none)
|
||
3. Verify ASLR disabled: `cat /proc/sys/kernel/randomize_va_space` (should be 0)
|
||
4. Run under valgrind: `valgrind --tool=memcheck ./starforth --doe`
|
||
5. Report to us immediately (GitHub issue)
|
||
|
||
**This is a genuine failure—we want to know.**
|
||
|
||
---
|
||
|
||
### Issue 4: "Build fails with linker errors"
|
||
|
||
**Diagnosis**: Likely **GCC version mismatch**
|
||
|
||
**Resolution**:
|
||
```bash
|
||
# Check GCC version
|
||
gcc --version # Should be 11.x
|
||
|
||
# If different, use Docker (see Section IV)
|
||
docker build -t starforth .
|
||
docker run starforth
|
||
```
|
||
|
||
---
|
||
|
||
## VII. BOUNTY FOR FALSIFICATION
|
||
|
||
**We offer recognition to the first replicator who:**
|
||
|
||
1. **Falsifies Claim C1** (Determinism)
|
||
- Reproduces CV > 0.1% under controlled conditions
|
||
- Co-authorship on erratum paper
|
||
|
||
2. **Falsifies Claim C2** (Convergence)
|
||
- Shows no convergence (p > 0.05) across 30 runs
|
||
- Acknowledgment in future publications
|
||
|
||
3. **Identifies Critical Bug**
|
||
- Bug that invalidates core claims
|
||
- Named credit in bug fix commit
|
||
|
||
**Why**: Science advances through falsification. If our claims are wrong, we want to know.
|
||
|
||
---
|
||
|
||
## VIII. CROSS-INSTITUTIONAL REPLICATION
|
||
|
||
### We Seek Collaborators
|
||
|
||
**Ideal Replication Partners**:
|
||
- **Academic labs** - Publish independent validation
|
||
- **Industry teams** - Validate for production use
|
||
- **Skeptical researchers** - Best critics make best validators
|
||
|
||
**What We Provide**:
|
||
- ✅ Technical support (email/video chat)
|
||
- ✅ Access to original hardware (if needed)
|
||
- ✅ Co-authorship on replication study (if desired)
|
||
|
||
**Contact**: rajames440@gmail.com (Robert A. James)
|
||
|
||
---
|
||
|
||
## IX. REPLICATION TIMELINE
|
||
|
||
### Phase 1: Initial Replications (3 months)
|
||
|
||
**Goal**: 3-5 independent replications at Level 2+
|
||
|
||
**Target**: Academic institutions, open-source contributors
|
||
|
||
**Deliverable**: Replication reports published as GitHub issues
|
||
|
||
---
|
||
|
||
### Phase 2: Cross-Platform Validation (6 months)
|
||
|
||
**Goal**: Validate on ARM, RISC-V, non-Linux platforms
|
||
|
||
**Target**: Embedded systems, L4Re deployments
|
||
|
||
**Deliverable**: Platform-specific replication guides
|
||
|
||
---
|
||
|
||
### Phase 3: Long-Term Monitoring (ongoing)
|
||
|
||
**Goal**: Track replication success rate over time
|
||
|
||
**Metric**: % of attempts that successfully reproduce claims
|
||
|
||
**Target**: > 90% success rate (indicates robust methodology)
|
||
|
||
---
|
||
|
||
## X. FAQ
|
||
|
||
### Q: "Do I need permission to replicate?"
|
||
|
||
**A**: No. Code is CC0 (public domain). Replicate freely.
|
||
|
||
---
|
||
|
||
### Q: "Can I use this in my own research?"
|
||
|
||
**A**: Yes. If replication succeeds, cite our work. If it fails, publish your findings.
|
||
|
||
---
|
||
|
||
### Q: "What if I find a bug?"
|
||
|
||
**A**: File a GitHub issue. We'll fix it and credit you.
|
||
|
||
---
|
||
|
||
### Q: "What if my hardware is different?"
|
||
|
||
**A**: Acceptable. Report your results with hardware specs. Convergence magnitude may vary; determinism (0% CV) should not.
|
||
|
||
---
|
||
|
||
### Q: "Can I modify the code?"
|
||
|
||
**A**: Yes (CC0 license). But for replication, use unmodified code first.
|
||
|
||
---
|
||
|
||
### Q: "What if I can't reproduce AND can't identify why?"
|
||
|
||
**A**: Contact us (rajames440@gmail.com). We'll help diagnose.
|
||
|
||
---
|
||
|
||
## XI. REPLICATION SCORECARD
|
||
|
||
**We will maintain a public scorecard of replication attempts:**
|
||
|
||
| Replicator | Institution | Date | Level | Result | Notes |
|
||
|-----------|------------|------|-------|--------|-------|
|
||
| R.A. James | Original | 2025-12-08 | 3 | ✅ Pass | Baseline |
|
||
| [Your Name] | [Your Org] | [Date] | [1/2/3] | [✅/❌] | [Link to report] |
|
||
|
||
**Transparency**: All replication attempts (success or failure) will be documented publicly.
|
||
|
||
---
|
||
|
||
## XII. PSYCHOLOGICAL WARFARE ASPECT
|
||
|
||
### Why This Document Exists
|
||
|
||
**Traditional approach**: "Trust us, we did the experiment."
|
||
|
||
**Our approach**: "Don't trust us—replicate it yourself."
|
||
|
||
**Effect**:
|
||
1. **Skeptics** who won't replicate look weak
|
||
2. **Skeptics** who try and fail expose themselves
|
||
3. **Skeptics** who succeed validate our work
|
||
4. **Everyone** sees we're confident enough to invite scrutiny
|
||
|
||
**Outcome**: Criticism becomes costly (requires effort), validation becomes compelling (independent verification).
|
||
|
||
---
|
||
|
||
## XIII. CONCLUSION
|
||
|
||
**We don't just tolerate replication—we demand it.**
|
||
|
||
**Three possible outcomes**:
|
||
1. ✅ You replicate successfully → Our claims are validated
|
||
2. ❌ You find a bug → We fix it and science advances
|
||
3. 🤷 You don't attempt replication → Your criticism carries less weight
|
||
|
||
**All three outcomes are acceptable to us.**
|
||
|
||
**The ball is in your court. Replicate or acknowledge you haven't.**
|
||
|
||
---
|
||
|
||
## XIV. REPLICATION CHECKLIST
|
||
|
||
**Before you start**:
|
||
- [ ] Read EXECUTIVE_SUMMARY.md (understand what you're replicating)
|
||
- [ ] Read FORMAL_CLAIM_TABLE.md (know the exact claims)
|
||
- [ ] Read NEGATIVE_RESULTS.md (understand failure modes)
|
||
- [ ] Choose replication level (1/2/3)
|
||
- [ ] Set up environment (Docker recommended)
|
||
|
||
**During replication**:
|
||
- [ ] Document all deviations from protocol
|
||
- [ ] Save all logs and outputs
|
||
- [ ] Record hardware/software specs
|
||
- [ ] Note any unusual behavior
|
||
|
||
**After replication**:
|
||
- [ ] Compare results to expected values
|
||
- [ ] Calculate deviations
|
||
- [ ] File report (success or failure)
|
||
- [ ] Contact us if needed
|
||
|
||
---
|
||
|
||
**Bottom Line**: If you're confident in your skepticism, replicate our experiment and prove us wrong. We'll thank you for it.
|
||
|
||
**License**: See ./LICENSE |