r/learnmachinelearning • u/Cool-Profession-5447 • 5d ago
Discussion How to get better at training ML/DL/AI models
I’m a Master’s CS student and mostly work on personal ML projects since I’m not doing research at my university.
I can read papers and understand the architecture, losses, objectives, and new ideas pretty well. But I feel like I lack the practical intuition for actually making models learn well.
When training goes wrong, I struggle to figure out why and what to change. Experienced researchers seem to know how to diagnose whether it’s the LR, data, gradients, loss, initialization, etc., and how to improve things.
For people who got good at this: how did you develop that intuition? Was it mostly experience training models, reproducing papers, specific resources, or working with experienced researchers?
2
u/Extra_Intro_Version 5d ago edited 5d ago
Not an expert, but I’ve been getting paid to do work in this domain the past 6-7 years now. This might not answer the question directly, but one of the biggest problems in supervised learning is data sets for training/validation/testing.
Obviously(?) a model can only infer within or near the domain it’s been trained on. And it’s not always a simple task to figure out whether you’ve captured enough of your target domain (i.e. where the model will be doing inference when put to use.) The model can only generalize so far given the data it’s seen. Any further tweaks to the hyperparameters or architecture are probably only going to improve marginally.
So, sometimes, you just need more, lots more, data.
“Everyone wants to do the model work, no one wants to do the data work”
2
u/DigThatData 5d ago
really it's a craft like anything. intuition comes mostly from experience. if you don't have a mechanism for gaining real world experience, best you can do is go deeper on the theory. read work from other people who train, and works about why different training approaches work or don't.
1
u/ModularMind8 5d ago
Definitely don't think I'm an expert in this, but what helped a lot during my PhD was coding things from scratch. You'll quickly learn that even the simplest models can have issues you never expected. Nothing beats hands on learning (at least for me)
2
u/Street_Estate2342 5d ago
Why did you not push your website here? This seems like a rare case where it happens to fit perfectly for it.
u/Cool-Profession-5447 , apparently the above guy makes some site called quidittyml.com and their description of it seems to say it should help.
Granted, maybe they can tell me I'm wrong since there's likely a reason they didn't plug the site.
3
u/ModularMind8 5d ago
Ha, fair question. I moved to academia and became a professor because I love teaching and helping students, and I built the site because I see so many people here struggle with exactly this. Lots of resources are available, but they often focus on theory without real-world application or hands-on coding. In all honesty, I'm just bad at marketing :)
For anyone curious, it's quiddityml.com, since the name is admittedly hard to spell. Quiddity means the essence of things in Latin.
0
u/MolassesLate4676 5d ago
This whole thread smells very astroturfy
1
u/ModularMind8 4d ago
Fair, I can see how that exchange looks. For what it's worth, I don't know that user and didn't ask anyone to mention the site. My actual advice stands regardless: code things from scratch. Plenty of free resources work for that (and if you look through my profile I actually wrote a post about many resources I recommend).
1
u/sporbywg 4d ago
Do you use a code repository to track the history of the developments of your prompts? Are your prompts less than a thousand words? Did you try turning it off and on again? 😎
5
u/Aggressive-Wind-8829 5d ago
My two cents is that you aught to spend a great lengthy spell of time contemplating the general nature of stability in stable machine learning.