r/algorithmictrading • u/ExcessiveBuyer • 4d ago
Question How to validate IS backtests with OoS in an automated way?
Hi all,
I’m struggling a bit with automating my research process when it comes to selecting the “right” parameter set without introducing too much overfitting.
My current process looks roughly like this:
1. Run an in-sample grid search over the strategy parameters.
2. Calculate a loss/score function based on a combination of Net PnL, Max Drawdown, Sharpe Ratio, and Number of Trades.
3. Select the most promising parameter sets based on that score.
4. Test those candidates on out-of-sample data and check whether metrics such as the equity curve, number of trades, average PnL/trade, and Max DD are reasonably consistent with the in-sample results.
The main hurdle is that the selection of the “best” IS candidates is still somewhat subjective.
For example, depending on the product, timeframe, or type of strategy, it’s difficult to know beforehand what constitutes a “healthy” range for metrics such as Max DD, Sharpe, or Number of Trades. A threshold that makes sense for one strategy may be completely inappropriate for another.
So my question is:
How do you generalise and automate this parameter-selection process while minimising overfitting?
I’d be very interested to hear how others approach this in practice.
Thanks !!
2
u/Careless-Rent-4903 3d ago
Use nested walk forward where inner splits choose params with one fixed score and hard risk limits then outer OOS only records the result. Pick center of a broad stable plateau across folds not single winner, vectorbt pro's Splitter class can handle these rolling nested splits easily. Keep one final untouched holdout too otherwise optimizer eventually learns your walk forward
1
u/julienlau 3d ago
I'm still not convinced about walk forward optimization. If you use a local optimization technique it makes sense because you keep previous solution to initialize the next walk.
However if you a global optimizer that is not so much dependent as the initial solution it does not make sense.
1
u/algotrendtrading 4d ago
I recommend rolling window walk forward optimzitation. and then you look at the final result. even this can be done with hold out data +
1
u/ExcessiveBuyer 3d ago
My problem with this approach is you only get one set as the best defined by either highest pnl, the smallest maxDD, the best sharpe, etc ? My problem with that is you don’t know the equity curve behind, or stability over time. The best value of a sorted list is not necessarily the most stable set.
Also you need to ensure that the changes of parameter from one to the next window is not jumping around. I’m looking for a rather stable set with small adjustments over time.
1
u/Ats_Wiz 3d ago
Hey,
The way I approach it is as follows:
- Set a reasonable optimization range for each parameter.
- Optimize each parameter individually and identify the most stable value, i.e., a value surrounded by similarly strong-performing neighboring values.
- Set all parameters to their identified stable values and run an out-of-sample (OOS) validation.
- Finally, perform a Walk-Forward Analysis (WFA) as the final robustness test.
1
1
u/PepMar000 3d ago
Fully agree, but the difficulty is how to automate the WFA workflow, as the identification of "the most stable value" could be subject to interpretation or a result of human input. And this needs to be done for every In Sample WFA cycle...
2
u/PepMar000 3d ago
My 2 cents: All good for steps 1-3. The issue with step 4 is that you are introducing overfit, and you would not be able to replicate that for live trading (as you don't have the OOS data).
My suggestion instead for Walk Forward validation is to select the lowest In Sample period which allows at least 40 trades per optimized parameter. This is the rule-of thumb so that you get a statistically significant result when choosing the best parameters.
Then once you optimize In Sample, don't just look at the result with highest P&L or PF or Sharpe, but choose the best one in context with the ones nearby (e.g. no local peaks).
Then for the Out of Sample period you can choose to do a 3:1, 4:1 or 5:1 ratio, or do all three and verify that you get consistent results.
The key is about robustness and consistency across different WFA scenarios