Final, Clean, Bug-Fixed S1-vs-S2-Comparison Context: आपके established-Point-6-Sequence (\"पहले-IS-पर-पूरा-risk-adjusted-profile, फिर-OOS-Validation\") को-genuinely-पूरा-किया — इस-दौरान 2-genuine-critical-bugs मिले-और-fix-हुए: Near-Zero-ATR-Bug: 2-events
Evidence from your latest S1 vs S2 comparison is now strong enough to treat S2 as the only serious core candidate and to move to the “role-clarity” + OOS step. The two bugs you caught were genuinely critical; with those fixed, the current picture is methodologically coherent.
---
1. Status of the S1 vs S2 decision (Point‑6 sequence)
Given:
- n = 13,761 events (good sample size)
- Sharpe: S1 = 0.008 vs S2 = 0.393
- Sortino: S1 = 0.017 vs S2 = 0.849
- Max DD (chronological): S1 = -213.51 vs S2 = -47.53
- Gain-to-Pain: S1 = 1.023 vs S2 = 3.047
- Skew: S1 ≈ 1.51, S2 ≈ 1.02 (both trend-following style, S1 more convex)
- Wilcoxon p-value for S2>S1: < 1e-6
- Bootstrap stability: 100% of resamples show S2>S1
Methodologically, for the relative question “S2 vs S1”:
- S2 dominates S1 on every risk-adjusted and distributional metric you’ve reported.
- The effect is large (not marginal) and statistically very robust (tiny p-value + 100% bootstrap superiority).
So within your earlier “Point‑6” structure:
- The in-sample risk-adjusted profile stage is complete for the strategy family S1/S2.
- You now have enough objective evidence to say:
- S2 is the core candidate.
- S1 does not justify a meaningful allocation as a core or co-core strategy on risk-adjusted grounds.
Any remaining “gaps” now are not about S1 vs S2, but about validating S2 itself for deployment.
---
2. Role clarity: S2 core vs S1 satellite/tail
Based on your metrics, a clean, objective framing would be:
2.1 S2 – Core strategy candidate
- Positive, non-trivial Sharpe/Sortino (for a long-short / event-based equity strategy, 0.39 / 0.85 is tilted to the “viable but not spectacular” side, which is realistic).
- Drawdown profile and Gain-to-Pain (≈3) are consistent with something that can be risk-managed at portfolio level.
- Skew still positive (trend-following flavour), so some convexity is already embedded; you’re not giving that up by choosing S2.
Conclusion: Treat S2 as the only serious core engine from this family.
2.2 S1 – Only if you deliberately want a lottery-style overlay
S1’s profile:
- Sharpe and Sortino are effectively zero.
- Max DD is extreme, suggesting very poor capital efficiency.
- Higher positive skew (1.51) hints at a “rare big winners, a lot of noise/pain” pattern.
This profile is only justifiable if:
- You can demonstrate, via conditional analysis, that S1 delivers very large wins in specific regimes you explicitly need (e.g., persistent multi-month directional trends) and
- You are willing to allocate a tiny, explicitly capped risk budget to buy that convex tail (e.g., a small “lottery ticket” overlay whose loss you can fully write off at the portfolio level).
If that regime-linked, tail-benefit does not show up clearly in a targeted analysis, the most objective conclusion is:
> S1 should be retired, not “kept as a satellite” out of sentiment.
So in pure role-clarity terms:
S2 = core. S1 = at most a micro-size convexity overlay; more realistically, zero allocation unless tail-regime benefits are empirically proven.
---
3. Remaining in-sample checks before OOS (if not already done)
Assuming you haven’t already done them exhaustively, these are the only IS-stage items I’d still treat as “critical” before you lock the spec and go to OOS:
1. Cost & frictions robustness
- Re-test S2 under pessimistic assumptions for:
- Brokerage, STT, GST, exchange fees
- Realistic entry/exit slippage (e.g., 1–2 ticks worse each side)
- Confirm Sharpe / DD / GtP degrade gradually, not collapse.
2. Sub-period / regime stability
- Break the IS history into 3–4 large, contiguous time blocks (e.g., pre-2010, 2010–2015, 2015–2020, 2020+ if the data allows).
- S2 should remain profitable and qualitatively similar across blocks (levels may differ, but no “one block carries everything” pathology).
3. Parameter / rule perturbation
- Small, local tweaks (e.g., ATR multipliers, lookback lengths) should not flip S2 from “good” to “terrible”.
- You want a plateau, not a spike, around your chosen setting.
If these three are already addressed and look reasonable, then your IS work is complete enough and the next legitimate step is OOS / walk-forward.
---
4. OOS / Walk-forward design: concrete suggestion
Two key principles:
1. Time-based, contiguous splits only.
- No random or cross-sectional shuffling; you already saw how order affects DD.
2. Freeze S2 now.
- From this point, no further rule-tweaking using any data that will later be called “OOS”.
You effectively want two layers:
- A strict final holdout (never touched for design decisions).
- A walk-forward CV on the earlier data to test robustness.
4.1 Hard OOS / final holdout
If your full history spans a long horizon (say, ≥12–15 years):
- Training (IS): first 70–75% of calendar time
- Final OOS: last 25–30% of calendar time
- Alternatively, fix a rule like: “Last 4–5 years as final OOS”, if that’s ≈25–30% of your sample.
Properties:
- The holdout is one single, contiguous block of recent data.
- No tuning, debugging, or parameter selection is allowed using that block.
This block becomes your “exam” after walk-forward.
If you have already “seen” the entire history while developing the bugs and metrics, then:
- Treat this held-out block as “pseudo-OOS”: still useful for robustness, but not fully independent.
- Your only truly virgin OOS will be data going forward in calendar time (paper/live).
4.2 Walk-forward on the training region
Within the earlier 70–75% (the training region), implement a rolling time-series CV such as:
Example template A (long history: ~12–20 years)
- IS window length: 8–10 years
- OOS window length (per fold): 1–2 years
- Step size: roll forward by 1 year.
Schematic (conceptual dates):
- Fold 1:
- Train: 2005–2012
- Test (OOS-1): 2013
- Fold 2:
- Train: 2006–2013
- Test (OOS-2): 2014
- Fold 3:
- Train: 2007–2014
- Test (OOS-3): 2015
- …continue until you hit the start of your final holdout block.
You then aggregate all these fold-level OOS segments (OOS-1,2,3,…) and compute:
- Sharpe, Sortino, DD, GtP, skew, etc. on the concatenated walk-forward OOS.
- Compare to IS metrics:
- You want a reasonable degradation, not a phase transition.
Example template B (shorter history: ~8–10 years)
If history is shorter, use:
- IS window: 5 years
- OOS per fold: 1 year
- Step: 6–12 months.
The logic remains the same: each fold simulates “design in the past, trade in the future”.
4.3 What to allow to change between folds
If S2 has very few true hyperparameters (mostly rule-based, not heavily estimated), then:
- Keep the core decision rules fixed across all folds.
- Allow only:
- Position-sizing scalers (e.g., risk per trade)
- Possibly universe / liquidity filters (if they are inherently time-varying, e.g., “top X by ADV”).
Do not re-optimise signal thresholds or ATR lookbacks each fold; that turns walk-forward into a hidden optimiser.
4.4 Evaluation criteria for OOS / walk-forward
You want the following to be true on:
- (a) concatenated walk-forward OOS, and
- (b) final holdout block:
1. Directionally same profile as IS:
- Positive Sharpe, positive GtP, positive skew.
2. Controlled deterioration:
- OOS Sharpe / GtP can be lower, but not “collapse to zero”.
3. Stable DD behaviour:
- No catastrophic new pattern (e.g., a block where the strategy is systematically short when the market keeps grinding up, or vice versa).
4. No single fold carrying everything:
- Profits should be distributed across OOS folds, not just in one lucky 2-year window.
If these hold, S2 passes a proper Point‑6 style validation and can be considered for real capital allocation (subject to broader portfolio context, capacity, and operational constraints).
---
Bottom line:
- For the relative S1 vs S2 question, your present evidence is already more than sufficient: S2 is the clear core candidate; S1 is at best a small, explicitly-defined convex overlay, and more likely a candidate for retirement.
- The next legitimate step in your sequence is indeed a time-based walk-forward + final holdout OOS, along the lines above, with the S2 spec frozen from this point forward.
If you have any further queries, please connect with us on 022-6290-10141 (Timings : 09.00 AM to 05.00 PM) or you can email us on info@cniinfoxchange.com