+687.5%
Total Return
89.0%
Annual CAGR
5.9%
Max Drawdown
15.0×
Calmar Ratio
71.3%
Win Rate
0.602R
Expectancy
1.25:1
Reward:Risk
5.75
T-Statistic
This system demonstrates a statistically confirmed positive expectancy across 3.24 years of backtest data encompassing 188 closed positions on XAUUSD M15. The strategy achieves 1.25:1 reward-to-risk, operating 26.8 percentage points above its mathematical breakeven threshold of 44.5%. Annualised CAGR of 89.0% relative to 5.9% maximum drawdown yields a Calmar ratio of 15.0×, significantly exceeding the professional benchmark range of 3–5×. Monte Carlo validation across 2,000 block-bootstrap simulations confirms structural consistency under adverse trade sequencing. 17 of 19 validation tests pass. 2 tests fall below the 70-point threshold and warrant review.
Section I
Analytical Findings & Observations
F.0
Statistical Plausibility Strength
The Statistical Plausibility gate scores this system 100/100. The combination of win rate, reward-to-risk and trade frequency is consistent with a genuine, durable edge; no plausibility concern was raised.
F.1
Statistical Significance Strength
The strongest dimension is Stat Significance (94/100). T-statistic of 5.75 exceeds the 99% two-tailed significance threshold of 2.576. (p = 0) Probability of results arising by chance is below 0.1%. The edge is statistically real given this 188-trade sample. This test applies a Welch t-test on the profit distribution and requires the mean return to be significantly different from zero.
F.2
Tail Risk Elevation Finding
CVaR (95%) measures 1.94× the average loss — within acceptable range. The worst 10 trades (5% of sample) average $1180.33 against a $607.7 mean loss. CVaR 99%: 2.15× average loss. Tail risk level: LOW. This elevation is partially structural: with a 1.25× RR ratio, the absolute average loss is modest, making tail events appear proportionally larger in ratio terms. Active monitoring of worst-case trade magnitude under live conditions is advisable.
F.3
MC Drawdown Envelope Observation
Block-bootstrap Monte Carlo (2,000 simulations, block size 10, AC lag-1: 0.164) produces a 95th-percentile maximum drawdown of 22.7% — approximately 3.8× the historical 5.9%. P50: 12.0%, P99: 30.3%. The historical sequence sits at the 4th percentile of the simulated distribution, confirming results were not predicated on an unusually favourable trade ordering. Risk management sizing against the MC P95 envelope rather than historical DD is advisable for live deployment.
F.4
Execution Sensitivity Observation
Under 10% execution degradation (wider spreads, adverse fills), expectancy retains 0.8× of its backtest level. At 0.602R base expectancy, the strategy remains profitable under this stress test. Forward testing under broker-accurate spread conditions is standard practice before capital deployment.
Development Considerations
Areas for Further Development
Edge Consistency
Edge Consistency scored 63/100. Day-of-week variance of 25.98 and profit factor log-variance of 0.306 indicate some session-dependent variation in edge quality. No structurally unprofitable weekday detected. Review whether the strategy has higher expectancy on specific days and whether disabling trading on the weakest session day would improve overall edge quality without meaningfully reducing trade count.
Monte Carlo Robustness
MC Robustness scored 66/100. Block-bootstrap CV of 0.207 indicates moderate sequence dependency. MC P95 DD of 22.7% vs historical 5.9% (3.8× expansion). Position sizing should be calibrated against the MC P95 envelope, not the historical DD. At 1% risk per trade, the P95 scenario implies up to 23.0% account drawdown — ensure capital allocation accounts for this rather than the 5.9% historical figure.
Drawdown Endurance
DD Endurance scored 72/100. The strategy spends 38% of calendar time below its high-water mark — equity is in drawdown for nearly three quarters of the backtest. Longest single episode: 82d 7h 57m, longest recovery: 38d 5h 10m. Recovery speed is strong (median penance 3.3× vs the theoretical 3.0× IID expectation), so the issue is frequency of drawdown entry, not recovery speed. Consider whether the entry filter can be tightened to reduce the number of small losing episodes — these accumulate the underwater time, not the major DD events.
Section II
Validation Test Results
94
Temporal
94
Statistical
88
Drawdown
93
Capital
88
Edge
63
Edge
80
Concentration
94
Ulcer
81
Sample
73
Return
74
MC
90
Consecutive
93
Cliff
66
MC
72
DD
94
Execution
90
Holding
79
Edge
81
Expected
Statistical Plausibility GATE100
Outcome combination consistent with a genuine edge
Applied as a score gate, not averaged into the Edge Score. See the technical explanation below.
Temporal Stability
94
EXCELLENT — All 6 periods profitable
All 6 of 6 equal calendar periods generated positive returns across the backtest horizon. No losing period detected. Return consistency CV of 0.43 confirms profitability is spread evenly, not concentrated in a single regime window. This score measures temporal robustness — a strategy that only profits in one or two periods may be regime-dependent rather than exhibiting a repeatable edge.
Statistical Significance
94
Highly significant edge (t=5.74, 99% confidence)
T-statistic of 5.75 exceeds the 99% two-tailed significance threshold of 2.576. (p = 0) Probability of results arising by chance is below 0.1%. The edge is statistically real given this 188-trade sample. This test applies a Welch t-test on the profit distribution and requires the mean return to be significantly different from zero.
Drawdown Analysis
88
MINIMAL drawdown (5.9% max, 2.8% avg episode)
Maximum drawdown of 5.9% with an average episode depth of 2.8%. The median recovery speed is 3.4 days per 1% of drawdown. 30 drawdown episodes were detected. No single episode dominates the overall drawdown profile, indicating consistent rather than event-driven risk. This test scores three components: max DD depth (50%), average episode depth (30%), and recovery quality in days per 1% of DD (20%).
Capital Efficiency
93
EXCELLENT — 88.9% annual, Calmar 15.0
Compound annual growth rate of 88.9% against 5.9% maximum drawdown. Calmar ratio of 15.0× significantly exceeds the professional benchmark of 3–5×. CAGR is computed using true compound growth (end equity / start equity)^(1/3.24 years), not simple annualisation. Capital efficiency rewards strategies that generate high risk-adjusted returns relative to their worst historical loss.
Edge Temporal Decay
88
STABLE — Edge is consistent with no meaningful decay
Rolling expectancy regression slope is positive (normalised +1.66), indicating the edge has strengthened over the backtest horizon. Second-half expectancy exceeds first-half by 142% (ratio 2.42). Profit factor across four quartiles (3.188, 3.637, 1.837, 3.818) is approximately stable (normalised slope 0.01). Win rate also trends downward across quartiles (normalised slope -0.09). This test detects whether a strategy's edge is eroding over time — a critical check for curve-fitted systems that perform well historically but deteriorate as market conditions evolve.
Edge Consistency
63
FAIR — Edge shows condition-dependent patterns
Win rate variance across weekdays falls within acceptable bounds. No structurally unprofitable weekday detected. Profit factor log-variance of 0.3056 and day-of-week variance of 25.9762 indicate edge quality does not fluctuate meaningfully by session day. This test checks whether the strategy's edge is consistent across all trading sessions or is heavily dependent on specific days or conditions.
Concentration Risk
80
GOOD — Acceptable distribution
Top 10% of winning trades account for 30.4% of total profit — above the 30% ideal-diversification threshold but below the 50% concentration-risk threshold. The largest single winner represents 4.2% of total profit, confirming no individual trade disproportionately sustains the overall result. Profit distribution is scored on two components: top-decile share (80%) and single largest winner share (20%). A well-distributed profit profile indicates genuine repeatable edge rather than lottery-dependent returns.
Ulcer Index
94
Excellent drawdown profile (UI: 1.9%)
Ulcer Index of 1.9% represents minimal cumulative drawdown pain. Max DD: 5.9%, avg DD: 1.15%, time underwater: 42.0%. Unlike maximum drawdown which captures a single worst point, the Ulcer Index integrates both depth and duration of all underwater periods — a UI below 5% indicates drawdowns are shallow, brief, and recover quickly.
Sample Adequacy
81
GOOD — 188 trades over 3.2y - solid validation
188 trades over 3.2 years is below the academic minimum of 250 trades (75% coverage) — confidence factor of 0.938 applied. MinTRL (minimum track record length) statistic: 30. Confidence factor applied to all other tests: 0.938. Sample adequacy is the foundational test — a backtest with insufficient trades cannot produce statistically valid conclusions regardless of how impressive the individual metrics appear.
Return Autocorrelation
73
FAIR — Minor positive correlation (AC: 0.162)
Lag-1 autocorrelation of 0.162 (lag-2: 0.137) — meaningful serial dependence detected. Returns are effectively independent. No martingale signature or hidden clustering pattern detected. Significant autocorrelation can indicate position-sizing escalation or regime-dependent behaviour that inflates backtest results.
MC DD Stability
74
FAIR — Minor sequence sensitivity detected
Under 1,000 permutation shuffles of the exact trade sequence, the 95th-percentile maximum drawdown reaches 12.1% — a 2.0× expansion from the 5.9% historical figure. 99th percentile: 14.6%. A ratio below 2.0× confirms the strategy does not rely on a particularly favourable trade ordering. This test measures whether the backtest drawdown is structurally representative or a statistical artefact of a lucky sequence of trades.
Consecutive Loss
90
MINIMAL — Statistically normal streak behavior
Maximum consecutive losing streak of 3 trades against a statistically expected maximum of 4.2 (ratio 0.71×). Loss clustering ratio of 0.92 — losses are not grouping more frequently than random distribution predicts. Worst streak required approximately 6 average wins to fully recover (damage ratio 4.1×). This test checks four dimensions: observed vs expected streak length (30%), loss clustering (25%), worst streak damage (25%), and recovery speed (20%).
Cliff Ratio
93
EXCELLENT — Healthy risk profile
95th-percentile loss of $1254.38 is 1.65× the average win of $757.94 — a healthy ratio indicating tail losses are not catastrophically larger than typical wins. Average loss: $607.7. Single largest loss ($1314.56) is 1.05× above the P95 level — no structural outlier. This test uses the 95th-percentile loss rather than the single largest loss as the primary metric, making the score more robust to one-off broker anomalies while still flagging structural outliers separately.
MC Robustness
66
FAIR — MODERATE
Block-bootstrap Monte Carlo (2,000 simulations, block size 10 preserving serial structure, AC lag-1: 0.164) produces a survival rate of 100.0% across all simulations. Coefficient of variation: 0.207. MC DD envelope — P50: 12.0%, P95: 22.7%. No position-scaling pattern detected — the strategy applies approximately uniform lot sizing regardless of recent outcomes. Block bootstrap preserves the serial correlation structure of returns (unlike naive IID resampling), producing more realistic stress scenarios.
DD Endurance
72
FAIR — MANAGEABLE (3.3x penance, 38% underwater)
Median penance ratio of 3.3× is near the theoretical IID expectation of 3.0× (Bailey & López de Prado, 2014). A ratio below 1.0 means recovery consistently takes less time than the drawdown formation period — a strong signal of genuine edge. Time spent underwater: 38.2%. Longest DD episode: 82d 7h 57m (6.9% of backtest). Longest recovery: 38d 5h 10m. 30 episodes detected. Scored on four components: penance ratio (35%), longest DD as % of backtest (25%), % time underwater (25%), and recovery consistency CV (15%).
Execution Cost Sensitivity
94
EXCELLENT — Edge survives execution degradation
Under a 10% uniform execution degradation scenario (wins reduced 10%, losses increased 10%), per-trade expectancy retains 0.8× of its backtest level. Original expectancy: $365.68 → degraded: $294.21 (19.5% impact). Strategy remains profitable under this stress test. At 0.602R base expectancy, the strategy retains meaningful cushion against real-world execution costs.
Holding Time
90
EXCELLENT — Winners held 1.6x longer than losers
Winners are held 1.56× longer than losers on average (winners: 3.6h, losers: 2.3h). This is a positive pattern — the strategy allows profitable trades to run while cutting losses relatively quickly. Discipline tier: EXCELLENT. Median ratio: 0.95×. At the current 0.602R expectancy, this does not materially impact performance.
Edge Quality
79
GOOD — Strong statistical edge
Expectancy of 0.602R per trade reflects a genuine and strong edge. Win rate of 71.3% operates 26.8 percentage points above the mathematical breakeven of 44.5%. Largest win is 5.69× the average win — some concentration in large outlier wins. Edge quality is scored on four dimensions: expectancy (35%), repeatability (30%), win rate margin (15%), and execution decay (20%).
Expected Shortfall
81
GOOD — Well-controlled tail risk (ES ratio: 1.9x)
CVaR (95%) measures 1.94× the average loss — within acceptable range. The worst 10 trades (5% of sample) average $1180.33 against a $607.7 mean loss. CVaR 99%: 2.15× average loss. Tail risk level: LOW. This elevation is partially structural: with a 1.25× RR ratio, the absolute average loss is modest, making tail events appear proportionally larger in ratio terms. Active monitoring of worst-case trade magnitude under live conditions is advisable.
Section II-b
Statistical Plausibility — Technical Explanation
Win rate, reward-to-risk and trade frequency jointly determine a strategy's Sharpe ratio (Lopez de Prado, Advances in Financial Machine Learning, Ch.15). Inverting that identity against the highest verified Sharpe on record gives the maximum win rate a genuine, durable edge could sustain for a given payoff. This test measures how far the observed result sits beyond that bound, combines it with a moment-corrected edge-strength signal, and applies the result as a score gate.
θ = √n · [ (π₊−π₋)·p + π₋ ] / [ (π₊−π₋)·√(p(1−p)) ]
Observed win rate (p)71.3%
Reward : risk1.25 : 1
Implied annualised Sharpe4.5
Ceiling Sharpe (θmax, verified)2.0
Max plausible win rate (p*)57.5%
Effective sample size188
Edge-smoothness signal (psr_z)6.4
Win-rate / payoff signal (z)3.8
The two signals are combined by taking the stronger of the two (a union of independent implausibility modes; a strategy need only be implausible on one axis). The dominant signal here is neither axis (within plausible range). The combined value is mapped through a continuous severity function on the statistic's own scale and applied as a score gate, not averaged into the Edge Score: a gated strategy's final score is compressed into a capped band whose position preserves its ranking against other gated strategies, while remaining below a genuine pass. The floor (0.25) bounds how far the score is moved, reflecting that statistical implausibility is evidence, not proof.
Basis: Lopez de Prado, Strategy Risk (AFML Ch.15); Probabilistic Sharpe Ratio (Bailey & Lopez de Prado, 2012); Sharpe ceiling anchored to the highest verified fund track record (Medallion-level, Cornell 2020). The ceiling θmax and the gate parameters are ErgodicLabs choices, stated as priors that extreme in-sample results tend not to persist out of sample. This measures statistical plausibility; it is not a determination of intent.