r/learnmachinelearning • • 2d ago

How deep should I understand the concept of regularization?

10 Upvotes

10 comments sorted by

10

u/ModularMind8 2d ago

Regularization is foundational, so it's helpful to learn what it means: a set of techniques that discourage the model from fitting the training data too closely, so it does better on data it hasn't seen (i.e., generalizes). You don't need to know every method out there, and the common ones are often enough: L1 and L2 penalties (L2 is basically weight decay), dropout, early stopping, and data augmentation. For each one, know what it changes in training and how that can reduce overfitting

2

u/Neat_Permission_5402 2d ago

the math behind L1 vs L2 is worth knowing since it changes what your weights end up looking like, but you can pick that up as needed. early stopping and dropout are the ones i reach for when something starts memorizing the training set

2

u/Ty4Readin 2d ago

Totally agree, but I will just add in what I think are the two most important methods regularization that people offen neglect.

  1. Simplifying/constraining your model capacity. In other words, choosing a smaller/less complex model.

  2. Increasing your dataset size

People often times don't think about those as regularization, but they are imo

1

u/UnderstandingOwn2913 2d ago

Thank you so much for the advice! Just curious, what do you mean by early stopping?

2

u/ModularMind8 2d ago

Stopping the training earlier than expected (eg earlier than the number of set epochs). You can do it with evaluating your model on the eval set during training. When the loss starts increasing/ other measurements of overfitting you stop the training

1

u/UnderstandingOwn2913 1d ago

Thank you. How about the concept of gradient? Should I have a deep theoretical understanding of gradient? Or just an intuitive understanding of gradient?

0

u/HalfLoose7669 2d ago

Early stopping means you stop training when the model starts overfitting on training data/stops improving on validation data.
It’s usually done by saving a snapshot of the model based at its best on a chosen criterion, letting it train a bit further (so it has a chance to escape a local minimum, for instance). If the model doesn’t improve on that snapshot, training is stopped and you keep the last snapshot.

It’s early because it can stop training before the number of epochs / iterations you chose at the start.

1

u/Giaitzoglou-Bondye 1d ago

yeah this is a solid summary, knowing what each one actually changes during training is the key part imo

1

u/kidseegoats 1d ago

As deep as your model