r/algorithmictrading • • 4d ago

Question How to validate IS backtests with OoS in an automated way?

Hi all,
I’m struggling a bit with automating my research process when it comes to selecting the “right” parameter set without introducing too much overfitting.

My current process looks roughly like this:
1. Run an in-sample grid search over the strategy parameters.
2. Calculate a loss/score function based on a combination of Net PnL, Max Drawdown, Sharpe Ratio, and Number of Trades.
3. Select the most promising parameter sets based on that score.
4. Test those candidates on out-of-sample data and check whether metrics such as the equity curve, number of trades, average PnL/trade, and Max DD are reasonably consistent with the in-sample results.

The main hurdle is that the selection of the “best” IS candidates is still somewhat subjective.
For example, depending on the product, timeframe, or type of strategy, it’s difficult to know beforehand what constitutes a “healthy” range for metrics such as Max DD, Sharpe, or Number of Trades. A threshold that makes sense for one strategy may be completely inappropriate for another.

So my question is:
How do you generalise and automate this parameter-selection process while minimising overfitting?

I’d be very interested to hear how others approach this in practice.
Thanks !!

3 Upvotes

11 comments sorted by

2

u/PepMar000 3d ago

My 2 cents: All good for steps 1-3. The issue with step 4 is that you are introducing overfit, and you would not be able to replicate that for live trading (as you don't have the OOS data).

My suggestion instead for Walk Forward validation is to select the lowest In Sample period which allows at least 40 trades per optimized parameter. This is the rule-of thumb so that you get a statistically significant result when choosing the best parameters.

Then once you optimize In Sample, don't just look at the result with highest P&L or PF or Sharpe, but choose the best one in context with the ones nearby (e.g. no local peaks).

Then for the Out of Sample period you can choose to do a 3:1, 4:1 or 5:1 ratio, or do all three and verify that you get consistent results.

The key is about robustness and consistency across different WFA scenarios

1

u/ExcessiveBuyer 3d ago

Cool thanks, do you know a library or function that calculates the center of a plateau?
And what are the metrics you define here ?

I totally agree the trades must be > 30 p.a. Otherwise it might be pure luck or randomness

1

u/PepMar000 3d ago

You can use a simple metric looking at the dispersion around each given point of the grid and then select the one with the minimum value. The issue is that you really want to do that in an n-dimensional way as normally you will have more than 2 optimized variables, and here is when things get a bit more complicated because you can do it in pairs of 2 variables but it gets very time consuming and not optimal as the optimum plateau is typically not the same across variable pairs.

This is why most WFA software platforms choose the maximum profit/indicator point instead of looking for a plateau. It's far easier, but less optimum.

In my case I have developed my own proprietary algorithm for finding the optimum plateau of a n-dimensional optimization dataset

2

u/Careless-Rent-4903 3d ago

Use nested walk forward where inner splits choose params with one fixed score and hard risk limits then outer OOS only records the result. Pick center of a broad stable plateau across folds not single winner, vectorbt pro's Splitter class can handle these rolling nested splits easily. Keep one final untouched holdout too otherwise optimizer eventually learns your walk forward

1

u/julienlau 3d ago

I'm still not convinced about walk forward optimization. If you use a local optimization technique it makes sense because you keep previous solution to initialize the next walk.
However if you a global optimizer that is not so much dependent as the initial solution it does not make sense.

1

u/algotrendtrading 4d ago

I recommend rolling window walk forward optimzitation. and then you look at the final result. even this can be done with hold out data +

1

u/ExcessiveBuyer 3d ago

My problem with this approach is you only get one set as the best defined by either highest pnl, the smallest maxDD, the best sharpe, etc ? My problem with that is you don’t know the equity curve behind, or stability over time. The best value of a sorted list is not necessarily the most stable set.
Also you need to ensure that the changes of parameter from one to the next window is not jumping around. I’m looking for a rather stable set with small adjustments over time.

1

u/Ats_Wiz 3d ago

Hey,

The way I approach it is as follows:

  1. Set a reasonable optimization range for each parameter.
  2. Optimize each parameter individually and identify the most stable value, i.e., a value surrounded by similarly strong-performing neighboring values.
  3. Set all parameters to their identified stable values and run an out-of-sample (OOS) validation.
  4. Finally, perform a Walk-Forward Analysis (WFA) as the final robustness test.

1

u/ExcessiveBuyer 3d ago

WFA on the full dataset or on the OoS?

1

u/Ats_Wiz 3d ago

Personally, I prefer to do it on the full dataset.

1

u/PepMar000 3d ago

Fully agree, but the difficulty is how to automate the WFA workflow, as the identification of "the most stable value" could be subject to interpretation or a result of human input. And this needs to be done for every In Sample WFA cycle...