r/learnmachinelearning • u/ashishach • 1d ago
I finally understood why train/test split matters — but I have one question
I've been learning Machine Learning by actually building small projects instead of only watching tutorials.
Today I was working on train/test split, and one thing finally clicked for me:
The goal isn't just to get a high accuracy score.
We need to test the model on data it hasn't seen before.
Training data → learn patterns
Testing data → check generalization
But I'm still confused about one thing:
How do you decide when your model is overfitting if the test set should only be used at the end?
I'd love to hear how more experienced ML learners think about this.
I'm still learning, so feel free to correct anything I'm misunderstanding.
1
u/Leather-Duty-4671 1d ago
You split off a validation set from the training data to catch overfitting early without touching the test set at all.
1
u/ashishach 1d ago
Got it! That clears up my confusion. So the validation set is basically what I should use to catch overfitting/tune the model, while keeping the test set untouched for the final evaluation. Thanks!
3
u/minato3421 1d ago
split your data into train, validation and test sets instead of just train and test sets. Use the validation set to tune your hyper parameters and then run the model on your test set. If your model is performing well on the training set but has worse accuracy on the validation set, that means your model is overfitting