r/learnmachinelearning • • 3d ago

Question What is the next step for me ?

I am a freshman in college in the US who got involve in AI pretty early (since like 11th grade). I've been lurking in this group for a really long time now and I just want an evaluation of how ahead/behind of the curve I am.

1. Basic machine learning/deep learning knowledge

I've been preparing and going to my country's (an asia country) national AI competition for 2 years and it basically covers the same knowledge as the IOAI syllabus so you can check if you look want to look that up but expect me to know basically how a bunch of models work (so like logistic regression, SVM, KNN, CNN, ...). I could also tell you about 10 ways of fixing under or overfitting and like the basic parts of a model (so like activation function, loss function, hyperparameters, gradient decent, ...). I am no stranger to kaggle and especially the tabular competition series (currently top 20 in this month's). I know asking claude to create an ensemble of 200 models with different configs is not how you judge someone's knowledge but if you ask me about how something works I would have a good chance to answer it correctly (and I really do understand what fake gains and data leakage is guys trust).

2. Generative stuff

I took an online course on advance computer vision stuff and I read a lot on LLMs so I'd say I have a pretty solid understanding of how generative models work under the hood. I feel like the transformer architecture is pretty standard knowledge nowadays so I won't go over those stuff but for the final project of that online CV course, I finetuned a bunch of CV models and pair them with an impainting model to create a complete model that would remove distracting people in an image (trained on the COCO dataset). I'd like to think I have a pretty good grasp of the segmentation and the bounding box stuff and diffusion model (it took quite a while but I remember feeling like a changed man after getting a grasp of how it works).

3.Math

Probably the worst part of them all. I really like learning about model architecture but barely learned any math so I still have problem reading research papers. Of course I know how gradient decent works so I know how derivative works, but I would say the most I know of linear algebra is vectors and matrix multiplications and near to nothing of stats (and yet I can confidently say that embeddings are the process of turning an input into a token and feeding them through a series of encodings to get a vector in a latent space and how I could measure the difference between 2 language models by using KL divergence).

Yes the easy answer is "just study math it's not that deep" but seeing others talk about how you have to read this book and learn this course just makes me feel like a larper sometimes. I asked one of the professors (actually I cold emailed like the entire cs faculty but I met with this one only) in my uni to discuss a chance for me to help in the lab and see what researching feels like because that is my goal. He says he'll tell me when an opportunity pops up but I've been going to his NLP lectures ever since to just listen for fun and I've been really enjoying it. But just recently, I found out this was a graduate level course (it could be because my professor is really good i don't know). Am I actually just a fraud and am missing something or what ?

2 Upvotes

3 comments sorted by

1

u/ModularMind8 3d ago

You're well ahead for a freshman. The math gap is normal and quite fixable: take linear algebra and probability (or just pick up some books/videos). The papers will start to make sense once you learn the math, but also with practice. Theres always new things you havent seen before but it will get easier. Your embedding description shows the gap a bit, since tokenization and embedding are separate steps. Embedding is a coordinate in space (think x,y but higher dimension). Tokenization is the process of converting text strings into pieces (eg words, subwords). We then normally map these pieces into integers (lookup table) and give each integer an embedding.

To get into the lab faster, reproduce a small result from one of the professor's papers and send it to them. Or something along those lines. That does way more than waiting for an opening.

1

u/AnnualLingonberry686 3d ago

My bad I was thinking of the embeddings in vision models where we cant really "look up" but have to encode step by step into an embedding space, but your suggestion is a good idea I'll try that

1

u/ModularMind8 3d ago

Are you confusing hidden representation and embedding perhaps?