feat(research): XSR01 "Cross-Sectional Residual" — candidato nuovo in forward-monitor
Primo candidato da molte ondate a superare marginale-ADDS, deflated-Sharpe 0.95, tre null avversariali e il test di lag, con corr ~0 a TUTTI e 5 gli sleeve attivi. COS'E'. Il meccanismo congelato di STATARB-RESID (W=45, sgn=+1, residuo OLS causale su BTC, z-score, tanh, vol-target 20%) applicato ai 50 alt HL, con posizioni DEMEANATE cross-sezionalmente ogni giorno. Il demeaning annulla ALGEBRICAMENTE la gamba BTC comune — sum(p_i-p̄)(r_i-r_btc) = sum(p_i-p̄)r_i — quindi non e' un paniere di coppie ma una strategia cross-sectional dollar-neutral, e l'ampiezza effettiva passa da 4.5 a 37.4. NB: non e' nato da un'idea nuova ma dalla DIAGNOSI di un fallimento (l'ampiezza effettiva 4.5 indicava un fattore comune da togliere). NUMERI (2.6 anni): Sharpe netta 1.82 (lorda 2.70), maxDD -2.6%, vol 2.3%. GATE: marginale ADDS (robust_oos + beats_noise_null + non-hedge + has_insample_edge); deflated-Sharpe 0.985 PASS; corr TP01 0.060 / XS01 -0.019 / VRP01 -0.001 / SKH01 0.001 / GTAA01 0.071; book a w=15% HOLD 2.36 -> 2.51, DD invariato. SCETTICO (r0725_statarb_demean_skeptic.py): 3 null A FEE-NEUTRALE (permutazione cross-sezionale / casuale / temporale a blocchi) centrati a ~0.01, candidato lordo 2.70 sopra il MASSIMO di 300 estrazioni (p<0.004); lag 1.82/1.19/0.81/0.51 a +0/1/2/3g = decadimento dolce di segnale lento, non firma di look-ahead; non ridondante con XS-momentum semplice (corr 0.17-0.29). Nota di metodo: i null vanno confrontati A FEE ZERO — permutare un segnale ne fa esplodere il turnover, quindi a fee piene il null perderebbe per COSTO invece che per assenza di informazione (p-value trionfale e falso). NON NEL BOOK, 3 motivi dichiarati: (a) storia 2.6a monoregime e CRESCENTE (Sh 2024 1.03 / 2025 1.98 / 2026 3.11) -> edge recente; (b) weights_tilt_null FALLITO (gate_pass=False a 10% e 15%); (c) margine di costo sottile — 1.82 a 0.05%/gamba, 0.93 a 0.10%, NEGATIVA a 0.20%, turnover 38.5% del lordo/g su 50 alt anche illiquidi con slippage NON modellato (rischio #1). ESEGUIBILITA': ticket/gamba $1.73 a $600 (sotto min-order $5 -> STAT-MODE oggi), $5.76 a $2000, $14.41 a $5000 -> diventa reale a ~$5k, non ai ~$20k di XS01. CABLATO: scripts/live/paper_xsr.py (config CONGELATA, 3 libri MODELED $2000 / REAL $600 / REAL $5000 con min-order per gamba, stato append-only) in cron_daily.sh; gate PRE-REGISTRATO a forward-day 0 (r0725_xsr_deploy_gate.py): decisione 2026-10-23, deploy solo se Sharpe>=1.0 E haircut di eseguibilita' a $5000 <=40% (se sfonda -> RITIRO a prescindere dallo Sharpe). tests/test_paper_xsr.py: 7 casi che bloccano config, dollar-neutralita', causalita' dello step, min-order e la trappola dei timestamp tz-aware (astype int64 -> epoche 1970, gemella di quella dell'ondata 2026-07-01). Book e pesi INVARIATI. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This commit is contained in:
@@ -67,6 +67,7 @@ scripts/research/blind/leaderboard.json
|
||||
data/paper_prevday/
|
||||
data/paper_combo/
|
||||
data/paper_statarb/
|
||||
data/paper_xsr/
|
||||
|
||||
# log esecuzioni del book live (stato runtime, contiene fill/fee del conto reale)
|
||||
data/live/
|
||||
|
||||
@@ -11,6 +11,7 @@ mkdir -p logs
|
||||
uv run python scripts/live/paper_portfolio.py # avanza paper TP01+XS01
|
||||
uv run python scripts/live/paper_prevday.py # forward-monitor lead prevday-breakout (PAPER, non deploy)
|
||||
uv run python scripts/live/paper_statarb.py # forward-monitor lead STATARB-RESID ETH/BTC ortogonale (PAPER, non deploy)
|
||||
uv run python scripts/live/paper_xsr.py # forward-monitor XSR01 cross-sectional residual 50 alt HL (PAPER, gate 2026-10-23)
|
||||
uv run python scripts/live/cc01_regime_watch.py # trigger regime CC01 (read-only: WARN>=10%/ALERT>=15% funding 30g)
|
||||
uv run python scripts/research/r0724_stable_snapshot.py # snapshot point-in-time supply stablecoin (sblocca WATCH STABLE a 12 mesi)
|
||||
# NB: l'esecuzione Deribit e' passata al BOOK (TP01+SKH01 nettati) via scripts/cron_book.sh a
|
||||
|
||||
@@ -0,0 +1,222 @@
|
||||
"""FORWARD-MONITOR — XSR01 (Cross-Sectional Residual, 50 alt Hyperliquid), PAPER.
|
||||
|
||||
NON e' esecuzione reale. E' il monitoraggio forward-only del candidato trovato il 2026-07-25
|
||||
(diario `2026-07-25-mat01-statarb-generalization.md`, sezione XSR01). Stesso trattamento che il
|
||||
progetto da' a ogni lead: config CONGELATA, finestra out-of-sample vera che parte da ora, nessun
|
||||
edge creduto prima.
|
||||
|
||||
COS'E'. Il meccanismo di STATARB-RESID (residuo OLS rolling causale di log(alt) su log(BTC),
|
||||
z-score su W, tanh, vol-target sullo spread) applicato a TUTTI i 50 alt certificati di Hyperliquid,
|
||||
con le posizioni DEMEANATE cross-sezionalmente ogni giorno. Il demeaning annulla algebricamente la
|
||||
gamba BTC comune ( sum_i (p_i - p̄)(r_i - r_btc) = sum_i (p_i - p̄) r_i ) e con essa il fattore
|
||||
che teneva l'ampiezza effettiva del paniere a 4.5: dopo il demeaning e' **37.4**. Non e' quindi un
|
||||
paniere di coppie ma una strategia CROSS-SECTIONAL dollar-neutral sui 50 alt.
|
||||
|
||||
PERCHE' E' IN MONITOR E NON NEL BOOK (i tre motivi, tutti dichiarati):
|
||||
1. **Storia 2.6 anni, un regime solo** (bear alt 2024-2026), e il rendimento e' CRESCENTE nel
|
||||
tempo (Sharpe 2024 1.03 / 2025 1.98 / 2026 3.11): la maggior parte dell'edge e' recente.
|
||||
2. **`weights_tilt_null` NON passa** (gate_pass=False a w=10% e 15%): delta_insample negativo,
|
||||
effetto strutturale della storia corta — ma un gate fallito resta fallito.
|
||||
3. **Margine di costo sottile**: netta 1.82 a 0.05%/gamba, 0.93 a 0.10%, NEGATIVA a 0.20%. Il
|
||||
turnover e' 38.5% del lordo al giorno su 50 alt fra cui parecchi illiquidi (GALA, BLUR, JTO,
|
||||
ORDI, WIF): lo **slippage non e' modellato** ed e' il rischio numero uno. Questo monitor
|
||||
serve soprattutto a misurare quanto costa davvero.
|
||||
|
||||
COSA HA GIA' SUPERATO (r0725_statarb_basket_gate.py + r0725_statarb_demean_skeptic.py):
|
||||
marginale vs TP01 = ADDS (robust_oos, beats_noise_null, non-hedge, has_insample_edge);
|
||||
deflated-Sharpe 0.985 PASS; corr ~0 a TUTTI e 5 gli sleeve attivi (max |0.071|);
|
||||
tre null a fee-neutrale (permutazione cross-sezionale, casuale, permutazione temporale a blocchi)
|
||||
tutti centrati a ~0.01 con il candidato LORDO 2.70 sopra il massimo di 300 estrazioni (p<0.004);
|
||||
decadimento con lag dolce (1.82 / 1.19 / 0.81 / 0.51 a +0/1/2/3g) = firma di segnale lento, NON
|
||||
di look-ahead; non ridondante con un momentum cross-sectional semplice (corr 0.17-0.29).
|
||||
|
||||
CONFIG CONGELATA (ogni ritocco AZZERA la finestra forward):
|
||||
W=45, sgn=+1, vol-target 20%, cap 2x, demean cross-sezionale giornaliero,
|
||||
universo = i 50 alt certificati in data/raw/hl_*_1d.parquet (BTC escluso), fee 0.05%/gamba.
|
||||
Riusa il segnale ESATTO di scripts/research/r0725_statarb_multi.signal (nessuna reimplementazione).
|
||||
|
||||
TRE LIBRI IN PARALLELO (onesta' sull'eseguibilita'):
|
||||
MODELED $2000 : ribilanciamento continuo di ogni gamba (limite superiore teorico).
|
||||
REAL $600 : min-order $5 per gamba. Ticket medio atteso $1.73 -> **quasi nulla si esegue**:
|
||||
a questo capitale la strategia e' STAT-MODE, e il libro lo mostra invece di
|
||||
nasconderlo.
|
||||
REAL $5000 : ticket medio ~$14 -> soglia realistica di eseguibilita'. E' il numero che dice
|
||||
a che capitale questa strategia diventa vera.
|
||||
|
||||
Stato: data/paper_xsr/{state.json, returns.jsonl} (append-only).
|
||||
|
||||
uv run python scripts/live/paper_xsr.py # avanza col dato disponibile
|
||||
uv run python scripts/live/paper_xsr.py --status # solo stato, non avanza
|
||||
uv run python scripts/live/paper_xsr.py --reset # azzera (riparte da ora)
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
|
||||
PROJECT_ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(PROJECT_ROOT))
|
||||
sys.path.insert(0, str(PROJECT_ROOT / "scripts" / "research"))
|
||||
|
||||
# Segnale ESATTO dello studio (nessuna reimplementazione -> niente drift col backtest).
|
||||
from r0725_statarb_multi import BASE, MIN_BARS, load_hl, signal, universe # noqa: E402
|
||||
|
||||
STATE_DIR = PROJECT_ROOT / "data" / "paper_xsr"
|
||||
STATE_FILE = STATE_DIR / "state.json"
|
||||
RETURNS_FILE = STATE_DIR / "returns.jsonl"
|
||||
|
||||
# --- CONFIG CONGELATA (frozen) -----------------------------------------------------------------
|
||||
FEE_LEG = 0.0005 # 0.05% per gamba sul nozionale mosso
|
||||
MODELED_CAPITAL = 2000.0
|
||||
REAL_BOOKS = ((600.0, "REAL-$600"), (5000.0, "REAL-$5000"))
|
||||
MIN_ORDER = 5.0
|
||||
ANN = np.sqrt(365.0)
|
||||
|
||||
|
||||
def build_panel():
|
||||
"""(ts, dt, W target [n x k], R_alt [n x k], r_btc [n], simboli). Tutto causale."""
|
||||
base_px = load_hl(BASE)
|
||||
pos_cols, ret_cols = {}, {}
|
||||
for sym in universe():
|
||||
try:
|
||||
tgt = load_hl(sym)
|
||||
except FileNotFoundError:
|
||||
continue
|
||||
ix = base_px.index.intersection(tgt.index)
|
||||
if len(ix) < MIN_BARS:
|
||||
continue
|
||||
p, _ = signal(base_px[ix], tgt[ix])
|
||||
pos_cols[sym] = pd.Series(p, index=ix)
|
||||
ret_cols[sym] = tgt[ix].pct_change().fillna(0.0)
|
||||
P = pd.concat(pos_cols, axis=1, sort=True).sort_index().fillna(0.0)
|
||||
R = pd.concat(ret_cols, axis=1, sort=True).sort_index().reindex(P.index).fillna(0.0)[P.columns]
|
||||
Q = P.sub(P.mean(axis=1), axis=0) # demean cross-sezionale
|
||||
W = Q / max(P.shape[1], 1) # peso di portafoglio per gamba
|
||||
rb = base_px.pct_change().fillna(0.0).reindex(P.index).fillna(0.0)
|
||||
# epoca ms ESPLICITA: `.astype("int64")`/`.view("int64")` su DatetimeIndex tz-aware non-ns
|
||||
# danno la scala sbagliata in pandas 2.x — trappola gia' pagata dal progetto (CLAUDE.md,
|
||||
# ondata 2026-07-01: merge_asof broadcastato = look-ahead invisibile a causality_ok).
|
||||
ts = np.array([int(t.timestamp() * 1000) for t in P.index], dtype="int64")
|
||||
return ts, P.index, W.to_numpy(float), R.to_numpy(float), rb.to_numpy(float), list(P.columns)
|
||||
|
||||
|
||||
def _state_io(write: dict | None = None):
|
||||
STATE_DIR.mkdir(parents=True, exist_ok=True)
|
||||
if write is not None:
|
||||
STATE_FILE.write_text(json.dumps(write, indent=2, default=str))
|
||||
return write
|
||||
return json.loads(STATE_FILE.read_text()) if STATE_FILE.exists() else None
|
||||
|
||||
|
||||
def _append(path: Path, rec: dict):
|
||||
STATE_DIR.mkdir(parents=True, exist_ok=True)
|
||||
with path.open("a") as f:
|
||||
f.write(json.dumps(rec) + "\n")
|
||||
|
||||
|
||||
def _book(cap: float, k: int) -> dict:
|
||||
return dict(cap=cap, cap0=cap, peak=cap, dd=0.0, w=[0.0] * k, n_skip=0, n_fill=0)
|
||||
|
||||
|
||||
def init_state() -> dict:
|
||||
ts, _, W, _, _, syms = build_panel()
|
||||
return dict(start_ts=int(ts[-1]), last_ts=int(ts[-1]), n_bars=0, syms=syms,
|
||||
modeled=_book(MODELED_CAPITAL, len(syms)),
|
||||
reals={name: _book(c, len(syms)) for c, name in REAL_BOOKS})
|
||||
|
||||
|
||||
def _step(bk: dict, w_target: np.ndarray, r_alt: np.ndarray, r_btc: float,
|
||||
min_order: float | None) -> float:
|
||||
"""Un passo di un libro: rende il netto della barra e aggiorna capitale/peak/DD in place."""
|
||||
w_held = np.asarray(bk["w"], float)
|
||||
ret = float(np.dot(w_held, r_alt) - w_held.sum() * r_btc) # gamba BTC = -somma dei pesi alt
|
||||
if min_order is None:
|
||||
w_new = w_target.copy()
|
||||
else:
|
||||
move = np.abs(w_target - w_held) * bk["cap"] >= min_order
|
||||
w_new = np.where(move, w_target, w_held)
|
||||
bk["n_fill"] += int(move.sum())
|
||||
bk["n_skip"] += int((~move).sum())
|
||||
d_alt = np.abs(w_new - w_held)
|
||||
d_btc = abs(w_new.sum() - w_held.sum())
|
||||
net = ret - FEE_LEG * (float(d_alt.sum()) + d_btc)
|
||||
bk["cap"] *= (1.0 + max(net, -0.99))
|
||||
bk["peak"] = max(bk["peak"], bk["cap"])
|
||||
bk["dd"] = max(bk["dd"], (bk["peak"] - bk["cap"]) / bk["peak"] if bk["peak"] > 0 else 0.0)
|
||||
bk["w"] = w_new.tolist()
|
||||
return net
|
||||
|
||||
|
||||
def advance(st: dict) -> dict:
|
||||
ts, dt, W, R, rb, syms = build_panel()
|
||||
if syms != st["syms"]: # universo cambiato -> non si avanza
|
||||
print(f" [XSR01] universo cambiato ({len(st['syms'])} -> {len(syms)}): "
|
||||
"la finestra forward richiede universo costante. Usa --reset per ripartire.")
|
||||
return st
|
||||
new = [i for i in range(len(ts)) if ts[i] > st["last_ts"]]
|
||||
if not new:
|
||||
return st
|
||||
for i in new:
|
||||
nm = _step(st["modeled"], W[i], R[i], float(rb[i]), None)
|
||||
rec = dict(ts=int(ts[i]), dt=str(pd.Timestamp(dt[i])), net_modeled=round(nm, 6))
|
||||
for _, name in REAL_BOOKS:
|
||||
rec[f"net_{name}"] = round(_step(st["reals"][name], W[i], R[i], float(rb[i]),
|
||||
MIN_ORDER), 6)
|
||||
rec["gross"] = round(float(np.abs(W[i]).sum()), 4)
|
||||
_append(RETURNS_FILE, rec)
|
||||
st["last_ts"] = int(ts[new[-1]])
|
||||
st["n_bars"] = st.get("n_bars", 0) + len(new)
|
||||
return st
|
||||
|
||||
|
||||
def print_status(st: dict) -> None:
|
||||
ts, _, W, _, _, _ = build_panel()
|
||||
days = (int(ts[-1]) - st["start_ts"]) / 86_400_000
|
||||
print("\n XSR01 forward-monitor (PAPER — cross-sectional residual 50 alt HL, NON deploy)")
|
||||
print(" config CONGELATA: W=45 sgn=+1 demean cross-sezionale, vol-target 20%, fee 0.05%/gamba")
|
||||
print(f" forward da {pd.Timestamp(st['start_ts'], unit='ms', tz='UTC').date()} "
|
||||
f"({st['n_bars']} barre 1d ~{days:.0f}g) gambe: {len(st['syms'])}")
|
||||
print(f" lordo corrente {np.abs(W[-1]).sum()*100:.1f}% del capitale "
|
||||
f"(dollar-neutral: netto {W[-1].sum():+.1e})")
|
||||
m = st["modeled"]
|
||||
print(f" MODELED ($2000, ribil. continuo): {(m['cap']/m['cap0']-1)*100:+6.2f}% "
|
||||
f"eq ${m['cap']:.2f} maxDD {m['dd']*100:.1f}%")
|
||||
for _, name in REAL_BOOKS:
|
||||
b = st["reals"][name]
|
||||
tot = b["n_fill"] + b["n_skip"]
|
||||
print(f" {name:<12} (min-order $5) : {(b['cap']/b['cap0']-1)*100:+6.2f}% "
|
||||
f"eq ${b['cap']:.2f} maxDD {b['dd']*100:.1f}% "
|
||||
f"gambe eseguite {100*b['n_fill']/tot if tot else 0:.0f}%")
|
||||
print(" -> il divario MODELED-REAL e' l'haircut di eseguibilita'; a $600 e' atteso GRANDE")
|
||||
print(f" log: {RETURNS_FILE}\n")
|
||||
|
||||
|
||||
def main() -> None:
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--status", action="store_true")
|
||||
ap.add_argument("--reset", action="store_true")
|
||||
args = ap.parse_args()
|
||||
if args.reset:
|
||||
for p in (STATE_FILE, RETURNS_FILE):
|
||||
if p.exists():
|
||||
p.unlink()
|
||||
st = _state_io(init_state())
|
||||
print("forward-monitor XSR01 inizializzato (forward-only da ora).")
|
||||
print_status(st); return
|
||||
st = _state_io()
|
||||
if st is None:
|
||||
st = _state_io(init_state())
|
||||
print("forward-monitor XSR01 inizializzato (forward-only da ora).")
|
||||
print_status(st); return
|
||||
if not args.status:
|
||||
st = advance(st); _state_io(st)
|
||||
print_status(st)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,219 @@
|
||||
"""STATARB-BASKET — il paniere multi-coppia e' uno SLEEVE? Giudizio coi gate veri del progetto.
|
||||
|
||||
DA DOVE VIENE. `r0725_statarb_multi.py` ha applicato il meccanismo CONGELATO (W=45, sgn=+1, preso
|
||||
da un altro studio: qui non e' stato cercato nulla) alle 50 coppie alt/BTC di Hyperliquid e ha
|
||||
trovato un paniere EW con Sharpe 0.82, maxDD -6.0%, corr a XS01 solo 0.207, 0/50 coppie degeneri.
|
||||
Quello script pero' si fermava alla statistica descrittiva. Questo lo porta davanti ai gate che il
|
||||
progetto usa per ammettere uno sleeve:
|
||||
|
||||
study/marginal_vs_tp01 -> ADDS / HEDGE / NOISE / REDUNDANT / DILUTES / NEUTRAL
|
||||
(include multi-cut, noise-null a corr-zero, hedge-vs-alpha)
|
||||
deflated_sharpe -> PASS >= 0.95, coi trial DAVVERO fatti
|
||||
weights_tilt_null -> ogni proposta di peso vs il null dei tilt casuali
|
||||
corr per-sleeve -> ridondanza contro i 5 sleeve attivi
|
||||
|
||||
PRECEDENTE CHE RENDE LA DOMANDA LEGITTIMA: XS01 e' nel book al 15% pur essendo STAT-MODE (19 gambe,
|
||||
non eseguibili a $600). Un paniere di coppie a 51 gambe sarebbe ammissibile alle STESSE condizioni,
|
||||
se supera i gate. Non e' quindi l'eseguibilita' a decidere qui, e' l'edge.
|
||||
|
||||
VARIANTI PRE-REGISTRATE (3, contate nel deflated-Sharpe; nessuna scelta guardando i risultati):
|
||||
V1 ALL50 — tutte le coppie alt/BTC valutabili. Zero liberta' di selezione.
|
||||
V2 MAJ19 — solo i 19 major liquidi, cioe' l'universo XS_UNIVERSE definito il 2026-06-19 per
|
||||
ragioni di LIQUIDITA' in un altro studio: sottoinsieme a priori, non scelto oggi.
|
||||
V3 DEMEAN — ALL50 con le posizioni demeanate cross-sezionalmente ogni giorno. Motivo strutturale
|
||||
dichiarato prima di guardare: le 50 coppie condividono la gamba BTC, per questo
|
||||
l'ampiezza effettiva era 4.5 invece di 50; togliere la componente comune dovrebbe
|
||||
alzarla. E' una costruzione standard (neutralizzare il fattore comune), non un fit.
|
||||
|
||||
uv run python scripts/research/r0725_statarb_basket_gate.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT))
|
||||
sys.path.insert(0, str(ROOT / "scripts" / "research"))
|
||||
sys.path.insert(0, str(ROOT / "scripts" / "research" / "alt"))
|
||||
|
||||
from r0725_statarb_multi import (BASE, BLOCK, MIN_BARS, SEED, load_hl, pnl, signal, universe,
|
||||
_dd, _sh)
|
||||
|
||||
ANN = np.sqrt(365.0)
|
||||
|
||||
|
||||
def pair_frames() -> tuple[pd.DataFrame, pd.DataFrame]:
|
||||
"""Matrici [data x coppia] di posizione e ritorno-spread, per tutte le coppie valutabili."""
|
||||
base_px = load_hl(BASE)
|
||||
pos_cols, spr_cols = {}, {}
|
||||
for sym in universe():
|
||||
try:
|
||||
tgt = load_hl(sym)
|
||||
except FileNotFoundError:
|
||||
continue
|
||||
ix = base_px.index.intersection(tgt.index)
|
||||
if len(ix) < MIN_BARS:
|
||||
continue
|
||||
p, s = signal(base_px[ix], tgt[ix])
|
||||
pos_cols[sym] = pd.Series(p, index=ix)
|
||||
spr_cols[sym] = pd.Series(s, index=ix)
|
||||
P = pd.concat(pos_cols, axis=1, sort=True).sort_index()
|
||||
S = pd.concat(spr_cols, axis=1, sort=True).sort_index()
|
||||
return P, S
|
||||
|
||||
|
||||
def basket_from_positions(P: pd.DataFrame, S: pd.DataFrame, cols=None,
|
||||
demean: bool = False) -> pd.Series:
|
||||
"""Rendimento EW del paniere, ricalcolando fee su ogni gamba (2 per coppia)."""
|
||||
if cols is not None:
|
||||
cols = [c for c in cols if c in P.columns]
|
||||
P, S = P[cols], S[cols]
|
||||
if demean:
|
||||
P = P.sub(P.mean(axis=1), axis=0)
|
||||
out = {}
|
||||
for c in P.columns:
|
||||
p = P[c].to_numpy(float)
|
||||
s = S[c].to_numpy(float)
|
||||
out[c] = pd.Series(pnl(p, s), index=P.index)
|
||||
return pd.concat(out, axis=1, sort=True).mean(axis=1, skipna=True).dropna()
|
||||
|
||||
|
||||
def eff_breadth(P: pd.DataFrame, S: pd.DataFrame, cols=None, demean=False) -> float:
|
||||
if cols is not None:
|
||||
cols = [c for c in cols if c in P.columns]
|
||||
P, S = P[cols], S[cols]
|
||||
if demean:
|
||||
P = P.sub(P.mean(axis=1), axis=0)
|
||||
R = {}
|
||||
for c in P.columns:
|
||||
R[c] = pd.Series(pnl(P[c].to_numpy(float), S[c].to_numpy(float)), index=P.index)
|
||||
M = pd.concat(R, axis=1, sort=True)
|
||||
C = M.corr().values
|
||||
off = C[~np.eye(len(C), dtype=bool)]
|
||||
rbar = float(np.nanmean(off))
|
||||
return len(C) / (1.0 + (len(C) - 1) * rbar) if rbar > -1 / (len(C) - 1) else float(len(C))
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 100)
|
||||
print(" STATARB-BASKET — il paniere multi-coppia supera i gate di ammissione a sleeve?")
|
||||
print("=" * 100)
|
||||
|
||||
from src.portfolio.sleeves import XS_UNIVERSE
|
||||
|
||||
P, S = pair_frames()
|
||||
print(f"\n coppie valutabili: {P.shape[1]} barre: {len(P)} "
|
||||
f"da {P.index[0].date()} a {P.index[-1].date()}")
|
||||
|
||||
maj = [s for s in XS_UNIVERSE if s != BASE]
|
||||
variants = {
|
||||
"V1 ALL50": dict(cols=None, demean=False),
|
||||
"V2 MAJ19": dict(cols=maj, demean=False),
|
||||
"V3 DEMEAN": dict(cols=None, demean=True),
|
||||
}
|
||||
series, srs = {}, []
|
||||
print("\n" + "-" * 100)
|
||||
print(" (1) LE TRE VARIANTI PRE-REGISTRATE")
|
||||
print("-" * 100)
|
||||
print(f" {'variante':<12}{'gambe':>7}{'Sharpe':>9}{'maxDD':>9}{'ret tot':>10}{'ampiezza eff':>15}")
|
||||
for nm, kw in variants.items():
|
||||
r = basket_from_positions(P, S, **kw)
|
||||
series[nm] = r
|
||||
srs.append(_sh(r.values))
|
||||
nlegs = (P.shape[1] if kw["cols"] is None else len([c for c in kw["cols"] if c in P.columns]))
|
||||
print(f" {nm:<12}{nlegs:>7}{_sh(r.values):>9.2f}{_dd(r.values)*100:>8.1f}%"
|
||||
f"{(np.prod(1+r.values)-1)*100:>9.1f}%{eff_breadth(P, S, **kw):>15.1f}")
|
||||
|
||||
# ---------------- gate del progetto su ciascuna variante
|
||||
from altlib import deflated_sharpe, marginal_vs_tp01
|
||||
|
||||
print("\n" + "-" * 100)
|
||||
print(" (2) GATE MARGINALE vs TP01 (il gate che decide, non lo Sharpe assoluto)")
|
||||
print("-" * 100)
|
||||
reports = {}
|
||||
for nm, r in series.items():
|
||||
rr = r.copy()
|
||||
rr.index = pd.to_datetime(rr.index, utc=True)
|
||||
m = marginal_vs_tp01(rr)
|
||||
reports[nm] = m
|
||||
print(f"\n --- {nm} ---")
|
||||
print(f" verdetto : {m.get('marginal_verdict')}")
|
||||
for k in ("corr", "robust_oos", "beats_noise_null", "is_hedge", "has_insample_edge",
|
||||
"tp01_beta", "alpha_ann"):
|
||||
if k in m:
|
||||
v = m[k]
|
||||
print(f" {k:<15}: {v if not isinstance(v, float) else round(v, 4)}")
|
||||
bl = m.get("blends", {})
|
||||
for w, d in (bl.items() if isinstance(bl, dict) else []):
|
||||
if isinstance(d, dict):
|
||||
print(f" blend w={w}: " + " ".join(
|
||||
f"{k}={round(v,3) if isinstance(v,(int,float)) else v}" for k, v in d.items()))
|
||||
|
||||
# ---------------- deflated Sharpe coi trial veri
|
||||
print("\n" + "-" * 100)
|
||||
print(" (3) DEFLATED SHARPE — trial reali = 3 varianti (la config W=45/sgn=+1 viene da un")
|
||||
print(" altro studio: su QUESTI dati non e' stata cercata)")
|
||||
print("-" * 100)
|
||||
for nm, r in series.items():
|
||||
dsr, null_max = deflated_sharpe(_sh(r.values), srs, r.values, dpy=365.0)
|
||||
print(f" {nm:<12} Sharpe {_sh(r.values):>5.2f} DSR {dsr:>6.3f} "
|
||||
f"(max atteso sotto il null {null_max:>5.2f}) {'PASS' if dsr >= 0.95 else 'sotto 0.95'}")
|
||||
|
||||
# ---------------- correlazione ai 5 sleeve attivi + tilt-null
|
||||
print("\n" + "-" * 100)
|
||||
print(" (4) RIDONDANZA vs I 5 SLEEVE ATTIVI + weights_tilt_null sul peso proposto")
|
||||
print("-" * 100)
|
||||
from src.portfolio.portfolio import Sleeve, StrategyPortfolio, weights_tilt_null
|
||||
from src.portfolio.sleeves import active_sleeves
|
||||
|
||||
sl = active_sleeves()
|
||||
best = max(series.items(), key=lambda kv: _sh(kv[1].values))
|
||||
nm_best, r_best = best
|
||||
rb = r_best.copy()
|
||||
rb.index = pd.to_datetime(rb.index, utc=True)
|
||||
print(f" variante col miglior Sharpe: {nm_best}")
|
||||
for s in sl:
|
||||
j = pd.concat({"a": s.daily(), "b": rb}, axis=1, sort=True).dropna()
|
||||
if len(j) > 60:
|
||||
print(f" corr -> {s.name:<20} {j['a'].corr(j['b']):>6.3f} ({len(j)} giorni comuni)")
|
||||
|
||||
daily_cols = {s.name: s.daily() for s in sl}
|
||||
daily_cols["STATARB_BASKET"] = rb
|
||||
w_cur = {s.name: s.weight for s in sl}
|
||||
w_cur["STATARB_BASKET"] = 0.0
|
||||
for w_new in (0.10, 0.15):
|
||||
k = 1.0 - w_new
|
||||
w_prop = {s.name: s.weight * k for s in sl}
|
||||
w_prop["STATARB_BASKET"] = w_new
|
||||
try:
|
||||
res = weights_tilt_null(daily_cols, w_cur, w_prop)
|
||||
print(f"\n peso {w_new:.0%} (i 5 scalati x{k:.2f}):")
|
||||
for kk, vv in res.items():
|
||||
if isinstance(vv, (int, float, bool, str)):
|
||||
print(f" {kk:<22} {round(vv,4) if isinstance(vv,float) else vv}")
|
||||
except Exception as e:
|
||||
print(f" [tilt-null non calcolabile a w={w_new}: {e.__class__.__name__}: {e}]")
|
||||
|
||||
# ---------------- book con/senza
|
||||
print("\n" + "-" * 100)
|
||||
print(" (5) BOOK CON E SENZA — Sharpe/DD FULL e HOLD-OUT")
|
||||
print("-" * 100)
|
||||
b0 = StrategyPortfolio(sl).backtest()
|
||||
for w_new in (0.10, 0.15):
|
||||
k = 1.0 - w_new
|
||||
sl2 = [Sleeve(s.name, s.weight * k, s.daily_fn, s.pos_fn) for s in sl]
|
||||
sl2.append(Sleeve("STATARB_BASKET", w_new, lambda _r=rb: _r))
|
||||
b1 = StrategyPortfolio(sl2).backtest()
|
||||
print(f" w={w_new:.0%} FULL {b0['full']['sharpe']:.2f} -> {b1['full']['sharpe']:.2f} "
|
||||
f"HOLD {b0['holdout']['sharpe']:.2f} -> {b1['holdout']['sharpe']:.2f} "
|
||||
f"DD {b0['full']['maxdd']*100:.1f}% -> {b1['full']['maxdd']*100:.1f}%")
|
||||
|
||||
print("\n" + "=" * 100)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,227 @@
|
||||
"""SCETTICO su STATARB-DEMEAN (V3) — prima di crederci.
|
||||
|
||||
IL CANDIDATO. `r0725_statarb_basket_gate.py` ha prodotto la prima cosa che somiglia a uno sleeve
|
||||
nuovo da molte ondate: posizioni del meccanismo congelato (W=45, sgn=+1) sulle 50 coppie alt/BTC,
|
||||
DEMEANATE cross-sezionalmente ogni giorno. Sharpe 1.79, maxDD -2.5%, ampiezza effettiva 37.4,
|
||||
marginale ADDS, deflated-Sharpe 0.985 (PASS), corr ~0 a tutti e 5 gli sleeve attivi.
|
||||
|
||||
Numeri cosi' puliti sono, nella storia di questo progetto, il momento in cui si e' quasi sempre
|
||||
scoperto un artefatto (feed testnet, look-ahead ffill, CC01 Sharpe 11, celle 0-perdite). Quindi
|
||||
prima di scrivere "trovata", si prova a UCCIDERLA.
|
||||
|
||||
⚠ OSSERVAZIONE STRUTTURALE CHE MOTIVA IL TEST PRINCIPALE. Demeanare annulla algebricamente la
|
||||
gamba BTC: sum_i (p_i - p̄)(r_i - r_BTC) = sum_i (p_i - p̄) r_i , perche' sum_i (p_i - p̄) = 0.
|
||||
Quindi V3 NON e' piu' un paniere di coppie: e' una strategia CROSS-SECTIONAL sui 50 alt con pesi
|
||||
(p_i - p̄). Va quindi testata come tale, e in particolare contro il sospetto che il merito sia
|
||||
della COSTRUZIONE (demean + vol-target + 50 asset) e non del SEGNALE.
|
||||
|
||||
I TEST (un candidato deve sopravvivere a tutti):
|
||||
T1 NULL DI PERMUTAZIONE CROSS-SEZIONALE — ogni giorno si permutano le etichette-asset del vettore
|
||||
di posizione. Distribuzione di posizioni, demeaning, vol-target, costi: tutto identico; sparisce
|
||||
SOLO l'abbinamento segnale<->asset. Se il candidato non batte questo null, l'edge e' costruzione.
|
||||
T2 NULL CASUALE — posizioni casuali tanh(N(0,1)) con la stessa scalatura e lo stesso demeaning.
|
||||
T3 NULL DI PERMUTAZIONE TEMPORALE a blocchi (20g) — rompe l'allineamento nel tempo.
|
||||
T4 LAG — ritardare la posizione di 1 e 2 giorni. Un segnale lento degrada dolcemente; un
|
||||
look-ahead crolla a zero immediatamente.
|
||||
T5 SOTTOPERIODI anno per anno — 2.6 anni sono un regime solo (bear alt): dove vive il rendimento?
|
||||
T6 COSTI — la fee attuale addebita 2 gambe per coppia, ma dopo il demeaning la gamba BTC si
|
||||
annulla: e' quindi SOVRASTIMATA. Si verifica comunque a 0.05 / 0.10 / 0.20% per gamba.
|
||||
T7 RIDONDANZA con un momentum cross-sectional semplice sugli stessi 50 (e' XS01 travestito?).
|
||||
T8 STRUTTURA — esposizione lorda/netta, bilanciamento long/short, turnover, gambe che si muovono.
|
||||
|
||||
uv run python scripts/research/r0725_statarb_demean_skeptic.py
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import pandas as pd
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
sys.path.insert(0, str(ROOT))
|
||||
sys.path.insert(0, str(ROOT / "scripts" / "research"))
|
||||
|
||||
from r0725_statarb_multi import FEE_LEG, SEED, _dd, _sh
|
||||
from r0725_statarb_basket_gate import pair_frames
|
||||
|
||||
ANN = np.sqrt(365.0)
|
||||
N_NULL = 300
|
||||
BLOCK = 20
|
||||
|
||||
|
||||
def ret_from_pos(P: pd.DataFrame, S: pd.DataFrame, fee_leg: float = FEE_LEG,
|
||||
lag: int = 0, demean: bool = True) -> pd.Series:
|
||||
"""Rendimento EW del paniere da una matrice di posizioni [data x asset]. `lag` giorni in piu'."""
|
||||
Q = P.sub(P.mean(axis=1), axis=0) if demean else P
|
||||
q = Q.to_numpy(float)
|
||||
s = S.to_numpy(float)
|
||||
held = np.zeros_like(q)
|
||||
k = 1 + lag
|
||||
if k < len(q):
|
||||
held[k:] = q[:-k]
|
||||
turn = np.abs(np.diff(held, axis=0, prepend=np.zeros((1, q.shape[1]))))
|
||||
net = held * s - 2.0 * fee_leg * turn
|
||||
return pd.Series(np.nanmean(net, axis=1), index=P.index).dropna()
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 100)
|
||||
print(" SCETTICO — STATARB-DEMEAN (V3): si riesce a ucciderla?")
|
||||
print("=" * 100)
|
||||
|
||||
P, S = pair_frames()
|
||||
P = P.fillna(0.0)
|
||||
S = S.reindex(P.index)[P.columns]
|
||||
base = ret_from_pos(P, S)
|
||||
sh0 = _sh(base.values)
|
||||
print(f"\n candidato: {P.shape[1]} asset, {len(P)} barre "
|
||||
f"({P.index[0].date()} -> {P.index[-1].date()})")
|
||||
print(f" Sharpe {sh0:.2f} maxDD {_dd(base.values)*100:.1f}% "
|
||||
f"vol ann {base.std()*ANN*100:.2f}% ret ann {base.mean()*365*100:.2f}%")
|
||||
|
||||
rng = np.random.default_rng(SEED)
|
||||
q = P.to_numpy(float)
|
||||
|
||||
# ⚠ I NULL SI CONFRONTANO A FEE ZERO. Permutare le posizioni ne fa esplodere il turnover
|
||||
# (un asset riceve ogni giorno una posizione scorrelata dalla precedente), quindi un null a
|
||||
# fee piene sarebbe battuto per motivi di COSTO e non di segnale: darebbe un p-value
|
||||
# artificialmente trionfale. A fee zero il confronto isola l'informazione.
|
||||
sh0_gross = _sh(ret_from_pos(P, S, fee_leg=0.0).values)
|
||||
print(f" Sharpe LORDA (fee 0) {sh0_gross:.2f} <- e' questa che i null devono battere")
|
||||
|
||||
# ---------------- T1 permutazione cross-sezionale
|
||||
print("\n" + "-" * 100)
|
||||
print(" T1 NULL DI PERMUTAZIONE CROSS-SEZIONALE — il test decisivo")
|
||||
print(" (stessa distribuzione di posizioni, stesso demeaning, stesso vol-target e costi;")
|
||||
print(" si rompe SOLO l'abbinamento segnale<->asset)")
|
||||
print("-" * 100)
|
||||
nulls = []
|
||||
for _ in range(N_NULL):
|
||||
qq = np.take_along_axis(q, rng.permuted(np.tile(np.arange(q.shape[1]), (q.shape[0], 1)),
|
||||
axis=1), axis=1)
|
||||
nulls.append(_sh(ret_from_pos(pd.DataFrame(qq, index=P.index, columns=P.columns), S, fee_leg=0.0).values))
|
||||
nulls = np.array(nulls)
|
||||
p1 = float((nulls >= sh0_gross).mean())
|
||||
print(f" null: media {nulls.mean():>6.2f} mediana {np.median(nulls):>6.2f} "
|
||||
f"p95 {np.percentile(nulls,95):>6.2f} max {nulls.max():>6.2f}")
|
||||
print(f" candidato LORDO {sh0_gross:.2f} -> p = {p1:.4f} "
|
||||
f"{'SOPRAVVIVE' if p1 < 0.05 else '*** UCCISO: l edge e la COSTRUZIONE, non il segnale ***'}")
|
||||
|
||||
# ---------------- T2 posizioni casuali
|
||||
print("\n" + "-" * 100)
|
||||
print(" T2 NULL CASUALE — posizioni tanh(N(0,1)) con stessa scala, demeaning e costi")
|
||||
print("-" * 100)
|
||||
scale = np.abs(q).mean()
|
||||
r2 = []
|
||||
for _ in range(N_NULL):
|
||||
qq = np.tanh(rng.normal(0, 1, size=q.shape)) * scale / max(np.abs(np.tanh(1.0)), 1e-9)
|
||||
qq[q == 0.0] = 0.0 # stessa maschera di "non ancora attivo"
|
||||
r2.append(_sh(ret_from_pos(pd.DataFrame(qq, index=P.index, columns=P.columns), S, fee_leg=0.0).values))
|
||||
r2 = np.array(r2)
|
||||
p2 = float((r2 >= sh0_gross).mean())
|
||||
print(f" null: media {r2.mean():>6.2f} p95 {np.percentile(r2,95):>6.2f} max {r2.max():>6.2f}"
|
||||
f" -> p = {p2:.4f} {'SOPRAVVIVE' if p2 < 0.05 else 'UCCISO'}")
|
||||
|
||||
# ---------------- T3 permutazione temporale a blocchi
|
||||
print("\n" + "-" * 100)
|
||||
print(" T3 NULL DI PERMUTAZIONE TEMPORALE (blocchi 20g)")
|
||||
print("-" * 100)
|
||||
nb = int(np.ceil(len(q) / BLOCK))
|
||||
r3 = []
|
||||
for _ in range(N_NULL):
|
||||
order = rng.permutation(nb)
|
||||
qq = np.concatenate([q[i * BLOCK:(i + 1) * BLOCK] for i in order])[:len(q)]
|
||||
r3.append(_sh(ret_from_pos(pd.DataFrame(qq, index=P.index, columns=P.columns), S, fee_leg=0.0).values))
|
||||
r3 = np.array(r3)
|
||||
p3 = float((r3 >= sh0_gross).mean())
|
||||
print(f" null: media {r3.mean():>6.2f} p95 {np.percentile(r3,95):>6.2f} -> p = {p3:.4f}"
|
||||
f" {'SOPRAVVIVE' if p3 < 0.05 else 'UCCISO'}")
|
||||
|
||||
# ---------------- T4 lag
|
||||
print("\n" + "-" * 100)
|
||||
print(" T4 LAG — un segnale lento degrada dolcemente, un look-ahead crolla")
|
||||
print("-" * 100)
|
||||
for lag in (0, 1, 2, 3):
|
||||
r = ret_from_pos(P, S, lag=lag)
|
||||
print(f" lag +{lag}g: Sharpe {_sh(r.values):>5.2f} maxDD {_dd(r.values)*100:>5.1f}%")
|
||||
|
||||
# ---------------- T5 sottoperiodi
|
||||
print("\n" + "-" * 100)
|
||||
print(" T5 SOTTOPERIODI — 2.6 anni sono UN regime; dove vive il rendimento?")
|
||||
print("-" * 100)
|
||||
for y in sorted(set(base.index.year)):
|
||||
w = base[base.index.year == y]
|
||||
if len(w) < 40:
|
||||
continue
|
||||
print(f" {y}: n={len(w):>4} Sharpe {_sh(w.values):>5.2f} "
|
||||
f"ret {(np.prod(1+w.values)-1)*100:>6.2f}% maxDD {_dd(w.values)*100:>5.1f}%")
|
||||
half = len(base) // 2
|
||||
print(f" prima meta' Sharpe {_sh(base.values[:half]):>5.2f} "
|
||||
f"seconda meta' {_sh(base.values[half:]):>5.2f}")
|
||||
|
||||
# ---------------- T6 costi
|
||||
print("\n" + "-" * 100)
|
||||
print(" T6 COSTI — nota: dopo il demeaning la gamba BTC si annulla, quindi addebitare")
|
||||
print(" 2 gambe per coppia SOVRASTIMA i costi. Si verifica comunque al rialzo.")
|
||||
print("-" * 100)
|
||||
for f in (0.0, 0.0005, 0.0010, 0.0020, 0.0040):
|
||||
r = ret_from_pos(P, S, fee_leg=f)
|
||||
print(f" fee {f*100:>5.2f}%/gamba: Sharpe {_sh(r.values):>5.2f} "
|
||||
f"ret ann {r.mean()*365*100:>6.2f}%")
|
||||
|
||||
# ---------------- T7 ridondanza con XS momentum semplice
|
||||
print("\n" + "-" * 100)
|
||||
print(" T7 RIDONDANZA — e' un momentum cross-sectional travestito?")
|
||||
print("-" * 100)
|
||||
from r0725_statarb_multi import load_hl, universe
|
||||
px = {}
|
||||
for s_ in universe():
|
||||
try:
|
||||
px[s_] = load_hl(s_)
|
||||
except FileNotFoundError:
|
||||
pass
|
||||
PX = pd.concat(px, axis=1, sort=True).reindex(P.index).ffill()
|
||||
for L in (30, 45, 90):
|
||||
mom = np.log(PX / PX.shift(L))
|
||||
z = mom.sub(mom.mean(axis=1), axis=0).div(mom.std(axis=1).replace(0, np.nan), axis=0)
|
||||
w = np.tanh(z.fillna(0.0))
|
||||
w = w.sub(w.mean(axis=1), axis=0)
|
||||
rx = ret_from_pos(w[P.columns].fillna(0.0), S, demean=False)
|
||||
j = pd.concat({"a": base, "b": rx}, axis=1, sort=True).dropna()
|
||||
print(f" XS-mom L={L:>3}: Sharpe {_sh(rx.values):>5.2f} "
|
||||
f"corr al candidato {j['a'].corr(j['b']):>6.3f}")
|
||||
|
||||
# ---------------- T8 struttura
|
||||
print("\n" + "-" * 100)
|
||||
print(" T8 STRUTTURA — esposizione, bilanciamento, turnover, gambe operative")
|
||||
print("-" * 100)
|
||||
Q = P.sub(P.mean(axis=1), axis=0)
|
||||
gross = Q.abs().sum(axis=1)
|
||||
net = Q.sum(axis=1)
|
||||
nlong = (Q > 0).sum(axis=1)
|
||||
turn = Q.diff().abs().sum(axis=1)
|
||||
print(f" esposizione LORDA media {gross.mean():>6.2f} (mediana {gross.median():.2f})")
|
||||
print(f" esposizione NETTA media {net.mean():>8.2e} (deve essere ~0 per costruzione)")
|
||||
print(f" gambe long/giorno media {nlong.mean():>5.1f} su {P.shape[1]} "
|
||||
f"(bilanciamento atteso ~50%)")
|
||||
print(f" turnover lordo/giorno {turn.mean():>6.2f} = {turn.mean()/max(gross.mean(),1e-9)*100:>5.1f}% del lordo")
|
||||
n_ass = max(P.shape[1], 1)
|
||||
# il rendimento e' la MEDIA sugli asset -> il peso di portafoglio della gamba i e' held_i/N.
|
||||
# Quindi il lordo del paniere e' gross/N, e il nozionale per gamba e' cap*|held_i|/N.
|
||||
gross_frac = gross.mean() / n_ass
|
||||
print(f" lordo del paniere {gross_frac*100:>5.1f}% del capitale (leva {gross_frac:.2f}x)")
|
||||
for cap in (600, 2000, 5000, 20000):
|
||||
tk = cap * gross_frac / n_ass
|
||||
print(f" cap ${cap:>6}: ticket medio per gamba ${tk:>7.2f} "
|
||||
f"{'-> sotto il min-order $5: NON eseguibile' if tk < 5 else '-> sopra il min-order'}")
|
||||
|
||||
print("\n" + "=" * 100)
|
||||
verdict = (p1 < 0.05 and p2 < 0.05 and p3 < 0.05)
|
||||
print(f" ESITO SCETTICO: {'i tre null NON la uccidono' if verdict else 'UCCISA da almeno un null'}")
|
||||
print("=" * 100)
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,100 @@
|
||||
"""r0725_xsr_deploy_gate — gate di decisione PRE-REGISTRATO per XSR01 (fissato 2026-07-25).
|
||||
|
||||
PERCHE' PRE-REGISTRATO. La regola si fissa oggi, forward-day 0, PRIMA di vedere qualunque barra
|
||||
out-of-sample. Decidere "a occhio" dopo aver visto i numeri e' selection-on-forward — il bias che
|
||||
questo progetto ha gia' pagato due volte (gate SELECTION-ON-HOLDOUT 2026-06-29; EW-STR refutato
|
||||
come selezione di 2° ordine 2026-07-01). Stesso trattamento gia' dato a STATARB-RESID.
|
||||
|
||||
CONTESTO. XSR01 e' il candidato del 2026-07-25: cross-sectional residual dollar-neutral sui 50 alt
|
||||
HL. In-sample (2.6 anni) fa Sharpe netta 1.82 / lorda 2.70, maxDD -2.6%, marginale ADDS,
|
||||
deflated-Sharpe 0.985, corr ~0 a tutti e 5 gli sleeve. Sopravvive a tre null fee-neutrali e al
|
||||
test di lag. NON e' nel book per tre motivi dichiarati: storia corta e monotona crescente,
|
||||
`weights_tilt_null` fallito, e margine di costo sottile con slippage NON modellato.
|
||||
|
||||
REGOLA (immutabile — ogni modifica va motivata nel diario come violazione):
|
||||
- Data decisione: 2026-10-23 (90 giorni forward dal 2026-07-25). Non prima.
|
||||
- Metrica primaria: Sharpe annualizzato di `net_modeled` su TUTTA la finestra forward.
|
||||
- GUARDIA DI COSTO, la piu' importante per questa strategia (turnover 38.5% del lordo/giorno su
|
||||
50 alt anche illiquidi): il rapporto Sharpe_forward / Sharpe_in_sample_netta(1.82) va letto
|
||||
insieme allo scarto MODELED-vs-REAL$5000. Se l'haircut di eseguibilita' a $5000 supera il 40%
|
||||
del rendimento modellato, la strategia e' un artefatto di costo e va RITIRATA a prescindere
|
||||
dallo Sharpe.
|
||||
- Esiti:
|
||||
Sharpe >= 1.0 AND haircut$5000 <= 40% -> CANDIDATO SLEEVE: proporre peso {5,10}% e
|
||||
giudicarlo con `weights_tilt_null` (che oggi NON passa) + maxDD combinato.
|
||||
Deploy solo se il tilt-null passa E il capitale e' >= $5000.
|
||||
0.3 <= Sh < 1.0 -> ESTENDI una sola volta di 90g (2027-01-21).
|
||||
Sharpe < 0.3 OR haircut > 40% -> RITIRO dal forward-monitor.
|
||||
- Guardie accessorie: config invariata (il test `tests/test_paper_xsr.py` la blocca);
|
||||
universo costante a 50 gambe (un cambio impone --reset e azzera la finestra).
|
||||
|
||||
Uso: `uv run python scripts/research/r0725_xsr_deploy_gate.py`
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import json
|
||||
from datetime import date
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[2]
|
||||
RETURNS = ROOT / "data" / "paper_xsr" / "returns.jsonl"
|
||||
|
||||
START = date(2026, 7, 25)
|
||||
DECISION = date(2026, 10, 23)
|
||||
DECISION_EXT = date(2027, 1, 21)
|
||||
SH_DEPLOY = 1.0
|
||||
SH_RETIRE = 0.3
|
||||
HAIRCUT_MAX = 0.40
|
||||
IS_SHARPE_NET = 1.82 # riferimento in-sample, per contesto
|
||||
|
||||
|
||||
def main() -> None:
|
||||
print("=" * 92)
|
||||
print(" XSR01 — gate di decisione PRE-REGISTRATO (fissato 2026-07-25, forward-day 0)")
|
||||
print("=" * 92)
|
||||
if not RETURNS.exists():
|
||||
print(f" nessun dato forward ancora: {RETURNS} assente.")
|
||||
print(f" data decisione: {DECISION} (mancano {(DECISION - date.today()).days} giorni)")
|
||||
return
|
||||
rows = [json.loads(x) for x in RETURNS.read_text().splitlines() if x.strip()]
|
||||
if len(rows) < 5:
|
||||
print(f" finestra forward troppo corta ({len(rows)} barre).")
|
||||
print(f" data decisione: {DECISION} (mancano {(DECISION - date.today()).days} giorni)")
|
||||
return
|
||||
rm = np.array([r["net_modeled"] for r in rows])
|
||||
key5 = "net_REAL-$5000"
|
||||
r5 = np.array([r.get(key5, np.nan) for r in rows])
|
||||
sh = float(rm.mean() / rm.std() * np.sqrt(365)) if rm.std() > 0 else 0.0
|
||||
eq = np.cumprod(1 + rm)
|
||||
dd = float(np.max((np.maximum.accumulate(eq) - eq) / np.maximum.accumulate(eq)))
|
||||
tot_m = float(np.prod(1 + rm) - 1)
|
||||
tot_5 = float(np.prod(1 + r5[np.isfinite(r5)]) - 1) if np.isfinite(r5).any() else float("nan")
|
||||
haircut = (tot_m - tot_5) / abs(tot_m) if tot_m not in (0.0,) and np.isfinite(tot_5) else float("nan")
|
||||
today = date.today()
|
||||
|
||||
print(f" finestra forward : {rows[0]['dt'][:10]} -> {rows[-1]['dt'][:10]} ({len(rows)} barre)")
|
||||
print(f" Sharpe fwd (mod) : {sh:+.2f} (in-sample netta era {IS_SHARPE_NET:.2f})")
|
||||
print(f" maxDD fwd : {dd:.1%}")
|
||||
print(f" ritorno MODELED : {tot_m*100:+.2f}% REAL-$5000: {tot_5*100:+.2f}%")
|
||||
print(f" haircut eseguib. : {haircut*100:.1f}% (guardia <= {HAIRCUT_MAX*100:.0f}%)")
|
||||
print(f" data decisione : {DECISION} (proroga unica: {DECISION_EXT})")
|
||||
|
||||
if today < DECISION:
|
||||
print(f"\n -> NESSUNA DECISIONE OGGI ({today}): mancano {(DECISION - today).days} giorni.")
|
||||
print(" Il numero corrente NON autorizza deploy anticipato (regola pre-registrata).")
|
||||
return
|
||||
if np.isfinite(haircut) and haircut > HAIRCUT_MAX:
|
||||
print("\n -> RITIRO: l'haircut di eseguibilita' supera la guardia. E' costo, non edge.")
|
||||
elif sh >= SH_DEPLOY:
|
||||
print("\n -> CANDIDATO SLEEVE: proporre peso {5,10}% e passarlo a weights_tilt_null "
|
||||
"(oggi NON passa) + maxDD combinato. Deploy solo con capitale >= $5000.")
|
||||
elif sh >= SH_RETIRE:
|
||||
print(f"\n -> SOTTO SOGLIA ma >= {SH_RETIRE}: estensione unica fino al {DECISION_EXT}.")
|
||||
else:
|
||||
print("\n -> RITIRO dal forward-monitor.")
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
main()
|
||||
@@ -0,0 +1,101 @@
|
||||
"""Lock del forward-monitor XSR01 (candidato trovato 2026-07-25).
|
||||
|
||||
Pin degli invarianti che rendono la finestra forward interpretabile. Se uno di questi salta, i
|
||||
numeri accumulati non valgono piu' e la finestra va azzerata consapevolmente (--reset), non
|
||||
lasciata scorrere con una config diversa.
|
||||
"""
|
||||
from __future__ import annotations
|
||||
|
||||
import sys
|
||||
from pathlib import Path
|
||||
|
||||
import numpy as np
|
||||
import pytest
|
||||
|
||||
ROOT = Path(__file__).resolve().parents[1]
|
||||
sys.path.insert(0, str(ROOT))
|
||||
sys.path.insert(0, str(ROOT / "scripts" / "live"))
|
||||
sys.path.insert(0, str(ROOT / "scripts" / "research"))
|
||||
|
||||
pytest.importorskip("pandas")
|
||||
import pandas as pd # noqa: E402
|
||||
|
||||
paper_xsr = pytest.importorskip("paper_xsr")
|
||||
|
||||
|
||||
def _panel():
|
||||
try:
|
||||
return paper_xsr.build_panel()
|
||||
except (FileNotFoundError, ValueError):
|
||||
pytest.skip("parquet Hyperliquid assenti")
|
||||
|
||||
|
||||
def test_config_congelata():
|
||||
"""La config e' il contratto della finestra forward: cambiarla la invalida."""
|
||||
from r0725_statarb_multi import CAP, SGN, TARGET_VOL, VOL_WIN, W
|
||||
assert (W, SGN, TARGET_VOL, VOL_WIN, CAP) == (45, +1, 0.20, 30, 2.0)
|
||||
assert paper_xsr.FEE_LEG == 0.0005
|
||||
assert paper_xsr.MIN_ORDER == 5.0
|
||||
assert paper_xsr.MODELED_CAPITAL == 2000.0
|
||||
|
||||
|
||||
def test_dollar_neutral_per_costruzione():
|
||||
"""Il demeaning DEVE dare somma dei pesi ~0: e' cio' che annulla la gamba BTC comune."""
|
||||
_, _, W, _, _, _ = _panel()
|
||||
net = np.abs(W.sum(axis=1))
|
||||
assert np.nanmax(net) < 1e-12, f"non dollar-neutral: max |somma pesi| = {np.nanmax(net)}"
|
||||
|
||||
|
||||
def test_timestamp_in_millisecondi_non_1970():
|
||||
"""Trappola pandas gia' pagata dal progetto: astype/view int64 su indice tz-aware.
|
||||
I ts devono essere ms epoch plausibili, non nanosecondi ne' secondi."""
|
||||
ts, idx, _, _, _, _ = _panel()
|
||||
assert ts.dtype == np.int64
|
||||
assert pd.Timestamp(int(ts[-1]), unit="ms", tz="UTC").year >= 2024
|
||||
assert pd.Timestamp(int(ts[0]), unit="ms", tz="UTC").year >= 2023
|
||||
# coerenza con l'indice da cui derivano
|
||||
assert pd.Timestamp(int(ts[-1]), unit="ms", tz="UTC").date() == idx[-1].date()
|
||||
|
||||
|
||||
def test_universo_50_alt_senza_btc():
|
||||
_, _, _, _, _, syms = _panel()
|
||||
assert len(syms) >= 45, f"attesi ~50 alt certificati, trovati {len(syms)}"
|
||||
assert "BTC" not in syms, "BTC e' il riferimento, non una gamba"
|
||||
|
||||
|
||||
def test_step_causale_la_posizione_nuova_non_guadagna_la_barra_corrente():
|
||||
"""Il rendimento della barra si calcola sui pesi TENUTI (decisi prima), mai sul target."""
|
||||
k = 3
|
||||
bk = paper_xsr._book(1000.0, k)
|
||||
bk["w"] = [0.10, -0.10, 0.0]
|
||||
w_target = np.array([-0.50, 0.50, 0.0]) # ribaltamento totale
|
||||
r_alt = np.array([0.10, -0.10, 0.0]) # premia i pesi VECCHI
|
||||
net = paper_xsr._step(bk, w_target, r_alt, 0.0, None)
|
||||
assert net > 0, "il netto deve riflettere i pesi tenuti, non il nuovo target"
|
||||
assert bk["w"] == pytest.approx(list(w_target)) # e dopo si e' ribilanciato
|
||||
|
||||
|
||||
def test_min_order_blocca_le_gambe_piccole():
|
||||
"""A capitale piccolo le gambe sub-min-order NON si muovono: e' il punto dei libri REAL."""
|
||||
k = 2
|
||||
bk = paper_xsr._book(600.0, k)
|
||||
w_target = np.array([0.001, -0.001]) # 600*0.001 = $0.60 << $5
|
||||
paper_xsr._step(bk, w_target, np.zeros(k), 0.0, paper_xsr.MIN_ORDER)
|
||||
assert bk["w"] == [0.0, 0.0], "gamba sub-min-order eseguita: il libro REAL sarebbe fittizio"
|
||||
assert bk["n_skip"] == 2
|
||||
|
||||
big = paper_xsr._book(600.0, k)
|
||||
paper_xsr._step(big, np.array([0.20, -0.20]), np.zeros(k), 0.0, paper_xsr.MIN_ORDER)
|
||||
assert big["w"] == pytest.approx([0.20, -0.20]) # $120 per gamba: si esegue
|
||||
assert big["n_fill"] == 2
|
||||
|
||||
|
||||
def test_fee_addebitata_su_alt_piu_gamba_btc():
|
||||
"""La fee copre le gambe alt E la gamba BTC netta: a ribilanciamento nullo, netto = 0."""
|
||||
k = 2
|
||||
bk = paper_xsr._book(1000.0, k)
|
||||
net = paper_xsr._step(bk, np.zeros(k), np.zeros(k), 0.0, None)
|
||||
assert net == pytest.approx(0.0)
|
||||
bk2 = paper_xsr._book(1000.0, k)
|
||||
net2 = paper_xsr._step(bk2, np.array([0.5, -0.5]), np.zeros(k), 0.0, None)
|
||||
assert net2 == pytest.approx(-paper_xsr.FEE_LEG * 1.0) # |0.5|+|−0.5| alt, BTC netto 0
|
||||
Reference in New Issue
Block a user