This page summarizes a walk-forward study of models that forecast next-five-trading-day annualized realized volatility for the S&P 500 ETF (SPY). Training uses data through 2018; out-of-sample evaluation runs from 2019 to the latest available session. Models are compared against naive persistence and GARCH(1,1) under a common protocol. Results are research aggregates only and are not trading recommendations.
The model guesses how choppy SPY will be over the next week. Below: that guess next to what really happened.
Recent forecast origins vs a prior reference window. Monthly auto-promote uses the same PSI bands (SAFE / CAUTION / CRITICAL). PSI detects distribution shift - it does not prove residual skill vs naive/GARCH (usefulness is labeled residual error on the track record). High PSI often tracks a vol-regime change (baselines move too). Architecture freeze stays locked.
| PSI | Tier | Pipeline action |
|---|---|---|
| < 0.10 | SAFE | Continue golden + shadow; auto-promote if those pass |
| 0.10 - 0.25 | CAUTION | WARN log; continue gates; promote only if later gates pass |
| ≥ 0.25 | CRITICAL | Block promote; keep current active; alert (no unfreeze) |
Out-of-sample MAE by year. Elevated points coincide with major volatility-regime breaks (2020, 2022).
Technical note (PDF) : protocol, freeze, QLIKE, embargo, Phase 4, gate, coverage
Each row is a completed walk-forward specification. Lower MAE, RMSE, and QLIKE are better. Primary ranking uses MAE on the raw volatility scale. QLIKE is a variance-scale diagnostic and is not the freeze criterion.
| # | Specification | MAE | RMSE | QLIKE | Retrain | Window | Model | Retrains |
|---|
Normal zero-drift bands from annualized σ̂ → next-5-session log-return intervals. Research diagnostic only - not a training signal and not a promote gate. Under-coverage means bands are too tight (fat tails and/or σ̂ too low). Architecture freeze stays locked.
Month-level MAE for the production freeze (monthly expanding baseline) versus the research MAE champion (lowest completed MAE; not live), with naive persistence on the same dates. Average MAE can hide crisis failure - Mar 2020 is the key check.
| Period | Production MAE | Research champ. | Naive | Prod vs champ | Prod vs naive |
|---|
QLIKE is a standard academic score for volatility forecasts. It looks at variance (the square of the vol number). Lower is better. It punishes a forecast that is too calm before a spike more than MAE does. Longer bars below are better. This is a diagnostic: it does not change the monthly production freeze.
Each row is one test on one loss: absolute error (MAE) or QLIKE, never mixed. d = freeze loss minus the other forecast's loss. Negative d means the freeze has smaller loss. Labels overlap for five days, so the test uses a serial-correlation correction (Harvey-Leybourne-Newbold). Monthly versus weekly is the same tree with a different retrain clock. A small p-value there does not change the live freeze.
| Pair | Loss | mean d | Who is lower | HLN p |
|---|
The 5-day realized-vol label at origin t uses returns t+1 through t+5. Walk-forward training through t therefore shares returns with the first test origins. Purging five labeled rows at the cut removes that overlap. The live freeze does not use an embargo.
| Slice | embargo=0 MAE | embargo=5 MAE | Ranking vs naive/GARCH |
|---|
| Specification | MAE | RMSE | Friday MAE | Midweek MAE | Target |
|---|
Martial Systems LLC. For investor or partnership inquiries, email martialsys@gmail.com.