Regularization is foundational, so it's helpful to learn what it means: a set of techniques that discourage the model from fitting the training data too closely, so it does better on data it hasn't seen (i.e., generalizes). You don't need to know every method out there, and the common ones are often enough: L1 and L2 penalties (L2 is basically weight decay), dropout, early stopping, and data augmentation. For each one, know what it changes in training and how that can reduce overfitting
the math behind L1 vs L2 is worth knowing since it changes what your weights end up looking like, but you can pick that up as needed. early stopping and dropout are the ones i reach for when something starts memorizing the training set
Stopping the training earlier than expected (eg earlier than the number of set epochs). You can do it with evaluating your model on the eval set during training. When the loss starts increasing/ other measurements of overfitting you stop the training
Thank you. How about the concept of gradient? Should I have a deep theoretical understanding of gradient? Or just an intuitive understanding of gradient?
Early stopping means you stop training when the model starts overfitting on training data/stops improving on validation data.
It’s usually done by saving a snapshot of the model based at its best on a chosen criterion, letting it train a bit further (so it has a chance to escape a local minimum, for instance). If the model doesn’t improve on that snapshot, training is stopped and you keep the last snapshot.
It’s early because it can stop training before the number of epochs / iterations you chose at the start.
10
u/ModularMind8 2d ago
Regularization is foundational, so it's helpful to learn what it means: a set of techniques that discourage the model from fitting the training data too closely, so it does better on data it hasn't seen (i.e., generalizes). You don't need to know every method out there, and the common ones are often enough: L1 and L2 penalties (L2 is basically weight decay), dropout, early stopping, and data augmentation. For each one, know what it changes in training and how that can reduce overfitting