Learning skill measured by a proper scoring rule — not a forecast.
Mechanism validation — historical walk-forward is SUGGESTIVE, not proof. Skill measured by proper scoring rule (log-loss / Brier) vs climatology & persistence baselines. Not a forecast, no edge, no direction.
T1–T4 learning targets as-of
source: reproduce_leg/…/SEER_LEARNING_REPORT.json
source: reproduce_leg/…/LEARN_T2_REPORT.md
source: reproduce_leg/…/LEARN_T3_REPORT.md
source: reproduce_leg/…/LEARN_T4_REPORT.json
Convention: a proper scoring rule (log-loss / Brier), lower is better; skill = baseline − learned. T2 accuracy ties climatology because the regime is near-constant (baseline-easy) — the win is calibration, not accuracy. T3 is refuted-degenerate (its a-priori terciles make the target near-constant) — refutation is reported plainly. A calibration improvement over a baseline is not a directional call and is not a trading signal.
Calibration vs baselines (SEER) as-of
Calibration, not accuracy. Skill measured by Brier / log-loss vs climatology & persistence baselines. Not a forecast.
| Target | horizon | scored | learned Brier | climatology | persistence | skill vs clim. | skill vs pers. |
|---|---|---|---|---|---|---|---|
| T1_VOL_REGIME | 5 | 3258 | 0.2783 | 0.3206 | 0.6157 | +0.0423 | +0.3374 |
| T2_STRUCTURE_REGIME | 5 | 3258 | 0.0899 | 0.0885 | 0.1488 | -0.0014 | +0.0588 |
| T3_ENERGY_STRESS | 5 | 3258 | 0.0196 | 0.0205 | 0.0307 | +0.0009 | +0.0110 |
| T1_VOL_REGIME | 3 | 3260 | 0.3119 | 0.3688 | 0.6221 | +0.0569 | +0.3102 |
| T2_STRUCTURE_REGIME | 3 | 3260 | 0.0860 | 0.0880 | 0.1389 | +0.0020 | +0.0529 |
| T3_ENERGY_STRESS | 3 | 3260 | 0.0194 | 0.0201 | 0.0282 | +0.0006 | +0.0088 |
| T1_VOL_REGIME | 1 | 3262 | 0.4055 | 0.4692 | 0.6830 | +0.0638 | +0.2775 |
| T2_STRUCTURE_REGIME | 1 | 3262 | 0.0765 | 0.0874 | 0.1020 | +0.0109 | +0.0255 |
| T3_ENERGY_STRESS | 1 | 3262 | 0.0174 | 0.0196 | 0.0196 | +0.0021 | +0.0022 |
Historical walk-forward — mechanism validation, suggestive only, NOT proof of skill.
Sigma-sync episodes — edge verdict FRESH
- independent episodes
- 46
- beats / ties / worse
- 18 / 1 / 27
- beats_rate
- 0.4
- binom_p
- 0.9324
- diagnostic only
- YES
crypto assets ~0.57 correlated -> effective independent episodes < raw count; proxy X
Prospective pre-registration ledger FRESH
A predictive claim would require many pre-registered, out-of-sample confirmations accumulated over real elapsed time. None exist yet.
- sealed (future-target, append-only)
- 95
- outcomes recorded
- 23
- confirmed out-of-sample
- 22
- inconclusive / pending
- 0
outcome tally: SUPPORTED: 22 REFUTED: 1
A single SUPPORTED result is NOT a skill claim; pending seals are shown as INCONCLUSIVE_PENDING until their target time elapses.
source: reproduce_leg/…/prereg_ledger/prereg_ledger.jsonl · reproduce_leg/…/prereg_ledger/prereg_outcomes.jsonl