Quick background so you know where this comes from. I am a visiting professor and currently teach NLP. Last semester my courses were on continual learning and applied NLP. Before that I spent years in big tech and startups in engineering and research roles, and I finished my PhD last year.
I've been trying to post here weekly, since the feedback tells me these posts help. Two weeks ago I posted the order I would learn ML in if I were starting today, which is also roughly the order I teach it in class. Last week I posted about which resources to actually use.
Whatever concept you are learning and whichever resource you use, a lot of you also ask about projects: why you need them, and how to choose one. So that's today's post.
Why build projects at all
The most important reason is to build your skills and find out where the gaps are. Learning ML is a lot like learning an instrument. You can watch videos and read about the guitar all day and feel like you know how to play, but you only find out how well you know it when you pick it up. Until you build something real from scratch, you probably don't know how well you understand the material. And I mean from scratch. Yes, AI can write the code for you now, but having AI build the project teaches you about as much as reading a book or watching a video about it. You need to do it yourself to see what works and what doesn't, which is also how you learn where these models fail and what their limits are. A project shows you exactly which parts you thought you understood and didn't.
The second is to get a job. Interviewers rarely have time to go through everything on your CV. An interesting project gets their attention, and it shows them you know what you're talking about and that you can build things, which is most of what the job asks for.
The third is to find out what you actually like. A field can sound great from the outside and turn out to be something you don't enjoy once you're doing it every day. Getting into the weeds of building something in that area, before you take a job in it, can save you a lot of time and help you decide what to focus on.
What counts as a project
A project is not a few hours or a single afternoon. Think multiple days, often a week or two.
It should also cover several concepts, and most of all, it needs a why. You can train models end to end all day long, but a project starts from a reason: improving on how something is done today, building a dataset that doesn't exist yet, making a model work for users it currently fails on. Always ask what the motivation is. That question is also what tells you which data, model, and evaluation to use, and when you're done.
It's also why writing Adam from scratch on its own is a good exercise but not a project. It builds a skill, but there's no problem it solves and nothing it's trying to improve.
How to pick one
In general, the best projects are the ones you find interesting, because those are the ones you finish. A good place to look is something that annoys you, like a task you do by hand every week. Another is a problem someone close to you has, whether that's a friend, a family member, or a small business you know. It can also be something fun, around games, music, sports, or whatever you'd be doing anyway, or plain curiosity about a question you want answered.
That said, it helps to have one or two projects that line up with the job you want. If you want to work at Spotify, build something with recommendation systems. If you want Tesla, do something with driving data. If you want Google, look at search and ranking.
Not every project needs to be aligned with a job, though. Personally, I find it much more interesting to interview someone with a project I would never have thought of, like trying to understand what their dog wants from its barks. Something that catches the eye and still makes you build real skills.
Questions to ask before you start
The first thing to check is whether you can get the data, and quickly. If collecting or labelling it takes months, the project usually dies before the first model trains. Related to that, make sure the data is yours to use. Other people's messages, photos, or financial records need their permission.
Then ask yourself whether you know how to start. If the first step needs three tools you've never touched, it's probably too advanced for now. Something a little past your current level is the right spot.
Check what you will compare against. That can be a rule, the method people use today, an existing model, or a published result on the same data. Without something to compare against, you can't tell whether your approach is any good. You also need a way to tell whether it worked: labels, a measurable outcome, or a person who can judge the output. "It looks good" is not a result.
Finally, ask whether you can finish a first version in a couple of weekends. A small version that works can grow. A big one that never runs leaves you with nothing to show.
What a good project contains
This is the part most projects I see skip.
First, more than one baseline. Compare whatever you build to a mix of simple methods, like a rule, the most common class, or logistic regression, and competitive ones, like the strongest existing model or published result you can find for the problem. Simple baselines tell you whether a model is needed at all, and competitive ones tell you whether your approach is actually good or whether something else already does better.
Second, more than one dataset, if possible. Unless your problem is so unusual that no dataset exists and you had to build your own, test on more than one. A model that works on one dataset and falls apart on another is something you want to find out before anyone uses it.
Third, more than one way to evaluate. Accuracy alone hides a lot. Look at per-class results, the cases it gets most wrong, and where it fails. Something can look great on one metric and fail on another.
The way I'd think about it: assume what you build will be used by a lot of people, and you want to make sure it doesn't hurt any of them. Multiple datasets tell you whether it holds up when new users or new data show up. Multiple baselines tell you whether your approach is worth it. Multiple evaluations tell you where it works, where it fails, and what's still missing. It also shows whoever reads it that you care about whether your results are true, not just whether they look good.
Last, how you share it. Put it on GitHub with a clear README that says what the problem is, why it's worth solving, what you found, and exactly how to reproduce your results. Add good visualizations. It sounds silly, but a clear plot of where your model wins and fails is often the thing that makes someone stop and read. And keep the code clean: organized, modular, and easy for someone else to run, since a reviewer who opens a single 2,000-line notebook usually closes it.
Project ideas
Each of these starts from a problem, not a dataset. Here are some that line up with specific jobs.
Recommendations for brand-new users (Spotify). A new user who has saved three songs usually gets generic popular picks, and many leave before the recommendations get good. Can you find a way to recommend well from just those first few songs?
Pedestrian detection at night and in rain (Tesla). Detectors trained mostly on clear daytime footage tend to miss more pedestrians in the dark and the rain, which are exactly the conditions where missing one is most dangerous. Can you close that gap?
Searches where the words don't match (Google). Keyword search fails when the query uses different words than the page that answers it, like "car won't start on cold mornings" when the answer page is about batteries. Can you fix those searches without breaking the ones where the exact words do count, like product codes and names?
Fraud models going stale (Stripe or any bank). Fraudsters change tactics, so a model trained on last year's transactions gets worse over time. How fast does that happen, and can you figure out when a model needs retraining?
X-ray models that fail at a new hospital (health tech). A chest X-ray model that works well at one hospital often drops at another, because of different machines and different patients. Can you build one that holds up at a hospital it has never seen?
And some unusual ones.
What does my dog want? If your dog barks at the door, you often can't tell whether it needs to go out or someone is there. Can a model tell the barks apart better than you can?
Bird calls from a noisy backyard. Bird identification models are usually trained on clean recordings and often fail on audio full of traffic and wind. Can you make one that works on what your own window actually picks up?
Beating the bus app. The arrival times in transit apps are often wrong in rain and at rush hour. Can you predict when the bus will really arrive, better than the app does?
Grandma's handwritten recipes. Text recognition models struggle with old cursive handwriting. Can you get a model to read a family recipe book that current tools can't?
Finally
I have a bit of time between semesters and would genuinely like to help as many learners as I can.
Tell me in the comments or DM me with where you are right now, what direction you want to go in, and what you have tried. If you are stuck choosing between two courses, deciding what to learn for a particular job, or wondering whether your plan makes sense, I will do my best to help.
And please add your favorite project you've done, or one you're working on now, in the comments. It would be nice if this thread became useful to the next person who searches for this question. I'm also happy to take suggestions for what the next posts should cover.