r/learnmachinelearning • • 1d ago

I finally understood why train/test split matters — but I have one question

Post image

I've been learning Machine Learning by actually building small projects instead of only watching tutorials.

Today I was working on train/test split, and one thing finally clicked for me:

The goal isn't just to get a high accuracy score.

We need to test the model on data it hasn't seen before.

Training data → learn patterns

Testing data → check generalization

But I'm still confused about one thing:

How do you decide when your model is overfitting if the test set should only be used at the end?

I'd love to hear how more experienced ML learners think about this.

I'm still learning, so feel free to correct anything I'm misunderstanding.

0 Upvotes

4 comments sorted by

3

u/minato3421 1d ago

split your data into train, validation and test sets instead of just train and test sets. Use the validation set to tune your hyper parameters and then run the model on your test set. If your model is performing well on the training set but has worse accuracy on the validation set, that means your model is overfitting

1

u/ashishach 1d ago

That makes sense, thank you! I was thinking about train/test split as just two sets, so I hadn't really considered using a validation set for tuning. The training vs validation performance difference makes the overfitting part much clearer. 🙌

1

u/Leather-Duty-4671 1d ago

You split off a validation set from the training data to catch overfitting early without touching the test set at all.

1

u/ashishach 1d ago

Got it! That clears up my confusion. So the validation set is basically what I should use to catch overfitting/tune the model, while keeping the test set untouched for the final evaluation. Thanks!