Martial Systems LLC
SPY Research · realized volatility

Short-horizon realized volatility forecasts for SPY

This page summarizes a walk-forward study of models that forecast next-five-trading-day annualized realized volatility for the S&P 500 ETF (SPY). Training uses data through 2018; out-of-sample evaluation runs from 2019 to the latest available session. Models are compared against naive persistence and GARCH(1,1) under a common protocol. Results are research aggregates only and are not trading recommendations.

Loading…
Selected model MAE
-
-
Naive persistence
-
trailing 5-day RV
GARCH(1,1)
-
multi-step baseline
Monthly LGBM reference
-
expanding window

Predicted vs what actually happened

The model guesses how choppy SPY will be over the next week. Below: that guess next to what really happened.

What the model predicted What actually happened
As-of (origin close)
-
next 5 sessions after this date
Live epoch start
-
unattended not started
Freshness
-
age of last any prediction
Counts
-
sim / bootstrap / live

Distribution stability (PSI / KS)

Recent forecast origins vs a prior reference window. Monthly auto-promote uses the same PSI bands (SAFE / CAUTION / CRITICAL). PSI detects distribution shift - it does not prove residual skill vs naive/GARCH (usefulness is labeled residual error on the track record). High PSI often tracks a vol-regime change (baselines move too). Architecture freeze stays locked.

Tier
-
loading…
Model PSI
-
SAFE if < 0.10
Model KS
-
two-sample vs reference
Pipeline action
-
naive / GARCH context
PSI Tier Pipeline action
< 0.10 SAFE Continue golden + shadow; auto-promote if those pass
0.10 - 0.25 CAUTION WARN log; continue gates; promote only if later gates pass
≥ 0.25 CRITICAL Block promote; keep current active; alert (no unfreeze)

Error by calendar year (selected model)

Out-of-sample MAE by year. Elevated points coincide with major volatility-regime breaks (2020, 2022).

Method summary

  • Target: annualized realized volatility over the five trading sessions after each close.
  • Features: information available at or before that close only (trailing RV, GARCH-family volatility, returns, volume, technicals; optional VIX and calendar fields in secondary tests).
  • Validation: chronological walk-forward; no shuffled cross-validation.
  • Baselines: trailing five-day realized vol (persistence) and causal GARCH(1,1).
  • Selection: production freeze is monthly expanding baseline LightGBM, chosen on stress-period MAE and retrain cost, not lowest aggregate MAE.
  • Phase 4 (2026-09-18): feature ablations, named HAR-RV, and joint QQQ/IWM. Drop-volume and HAR-RV lags meet the research exit bar. Quantiles parked. Live freeze unchanged.
  • QLIKE: Patton (2011) variance-scale score, reported as a diagnostic. It does not choose the freeze.
  • Diebold-Mariano: DM-MAE and DM-QLIKE run separately with a Harvey-Leybourne-Newbold adjustment for overlapping 5-day labels. Not a freeze input.
  • Embargo: a 5-row label-overlap purge on the evaluate walk-forward keeps the MAE ranking versus naive and GARCH, including March 2020. Not a freeze input.
  • Live data quality: a forecast row is not appended if the raw cache is stale versus the last NYSE session, missing sessions, or has non-positive volume or a suspicious adj-close jump.

Specification comparison (MAE ascending)

Each row is a completed walk-forward specification. Lower MAE, RMSE, and QLIKE are better. Primary ranking uses MAE on the raw volatility scale. QLIKE is a variance-scale diagnostic and is not the freeze criterion.

# Specification MAE RMSE QLIKE Retrain Window Model Retrains

Limitations and diagnostics

Interval coverage (research)

Normal zero-drift bands from annualized σ̂ → next-5-session log-return intervals. Research diagnostic only - not a training signal and not a promote gate. Under-coverage means bands are too tight (fat tails and/or σ̂ too low). Architecture freeze stays locked.

80% band hit
-
nominal 80%
95% band hit
-
nominal 95%
Under-coverage?
-
gap vs nominal
Worst year (80%)
-
by-year coverage_80

Stress periods - production freeze vs research MAE champion

Month-level MAE for the production freeze (monthly expanding baseline) versus the research MAE champion (lowest completed MAE; not live), with naive persistence on the same dates. Average MAE can hide crisis failure - Mar 2020 is the key check.

Period Production MAE Research champ. Naive Prod vs champ Prod vs naive

Investors

Martial Systems LLC. For investor or partnership inquiries, email martialsys@gmail.com.