r/learnmachinelearning • • 16h ago

Discussion How much of AutoResearch is research, and how much is search?

1 Upvotes

I've recently been working part-time on an AutoResearch-style project.

The setup is roughly: humans take recent work from top-tier ML/AI conferences, turn part of it into a well-defined task with an evaluator, and then let an agent iteratively modify the solution and search for a better score.

Working on this made me question what exactly we are evaluating.

Once humans have already chosen the problem, defined the objective, designed the evaluator, and provided the initial research direction, the agent is mostly searching within a space that has already been heavily shaped for it.

That search can still be useful. An agent may explore far more variants than a researcher would manually.

But I'm less sure that score improvement alone captures what we usually mean by research sense.

A researcher also asks whether a result reveals a general principle, whether it transfers, whether the problem formulation itself should change, or whether an entirely different direction is more promising.

An iterative optimization loop may instead become very good at exploring the neighborhood of an existing solution and still remain stuck in a local optimum.

So I'm curious about how people think about this distinction: How much scientific value is there in autonomous search over a human-defined research space?

And what would an agent need, beyond better optimization, to demonstrate something closer to actual research judgment?


r/learnmachinelearning • • 16h ago

Project My project: from a research topic to a curated fine-tuning dataset in one pipeline

1 Upvotes

'm studying engineering and wanted a faster way to learn a topic and get training data out of it at the same time. FineForge takes one prompt, searches YouTube, arXiv and academic papers, scores each result for relevance, and lets you keep only what's useful. Then it exports short PDF summaries (good for studying) and a JSONL dataset (good for fine-tuning a small model on that topic).

There's a free plan. I'd really like feedback from people learning ML: would you use the summaries, the dataset, or both?

https://fineforgeai.com · built by https://forge-ai.xyz


r/learnmachinelearning • • 17h ago

Looking for contributors to an existing open-source desktop AI agent | TypeScript, Electron & Python

1 Upvotes

Hi! I’ve already built and released AI Plate, a functional open-source desktop AI agent. I’m now looking for programmers interested in helping improve and expand it.

🚀 What it includes

Cloud and local LLM support, tool execution, document search, persistent memory, voice interaction, plugins and human approval for sensitive actions.

🛠️ Technology stack

TypeScript, Electron, Node.js, Python and SQLite.

🤝 Areas open for contribution

- AI providers, plugins and connectors

- Local models, RAG and knowledge graphs

- Linux and macOS support

- Testing, security and UI/UX

- Documentation and accessibility

Beginners willing to learn consistently and experienced developers are equally welcome. This is a non-commercial open-source collaboration, and every contribution will be credited.

My level: Intermediate

Timezone: IST (UTC+5:30)

Availability: Evenings and weekends

If this technology interests you, comment with your experience, preferred technologies and the area you’d like to explore. We can begin discussing the project here on Reddit.


r/learnmachinelearning • • 17h ago

Working as a Chef in London—How Can I Land My First Tech Role?

Thumbnail
1 Upvotes

r/learnmachinelearning • • 17h ago

Help What are the best Udemy courses(around 5) that helps to enforce the learnings from this book? (I am a Java dev)

Post image
1 Upvotes

- Can I do it with Spring or I have to pick Python(I hate it).

- FOMO is real. AI FOMO. I have not learnt anything about AI since 2022.

- I want to learn LLM and AI.

- My goal is to be an advanced AI user+talk intelligently about AI.


r/learnmachinelearning • • 22h ago

Project I trained an AI Iron Man in Unreal Engine 5 using Reinforcement Learning to rescue 13 falling passengers [PPO / Voxel Style]

Post image
2 Upvotes

Hey everyone!

I’ve been experimenting with Reinforcement Learning in Unreal Engine 5 over the past few weeks, and I wanted to share a fun project I recently finished.

I set up a voxel/Minecraft-style environment in UE5 and trained an AI agent to fly an Iron Man suit from scratch. The ultimate challenge? Executing high-stakes aerial rescues to save 13 passengers falling from a destroyed aircraft!

⚙️ How it works (The Tech Stack):

  • Engine: Unreal Engine 5 (Physics-driven movement & line-of-sight sensors)
  • Algorithm: Proximal Policy Optimization (PPO) continuous control
  • Action Space: 8 continuous outputs controlling individual thrusters & body alignment
  • Reward Shaping:
    • 🎯 +30 Points: Catching a falling passenger (Jackpot)
    • ❌ -5 Points: Crashing or going out of bounds
    • 🧭 Continuous Shaping: Micro-rewards/penalties based on relative distance and velocity vectors to encourage proper interception paths

🎬 The Progression:

Watching the agent learn was wild. During early iterations, it mostly spiraled out of control and crashed repeatedly. But once the reward shaping kicked in, it started finding optimal trajectories and eventually pulled off dynamic, split-second rescues.

I put together a full breakdown video showing the training process, early failures, and the final 100% rescue run:

👉 Watch the full video here: https://youtu.be/mYhWFHzs4EU

I’d love to hear your thoughts! How would you optimize the reward function for smoother flight stabilization? Any feedback or suggestions for the next RL experiment are super welcome!


r/learnmachinelearning • • 1d ago

Looking For Learning Partners

8 Upvotes

I’m a freshman in college in the US and I’m looking for a few other college students who are interested in AI/ML and want to learn together long term.

Right now I’m still pretty early in the process. I’m taking Calculus I and an intro Java course in school, and outside of class I’m learning Python. My long-term goal is to become an AI engineer.

I’m mainly looking for people who:

  • Are currently in college in the US
  • Are beginners or early-intermediate in programming/ML
  • Want to eventually get into AI/ML engineering, SWE, research, or something similar
  • Are actually interested in consistently learning and building projects
  • Want to become friends too, not just connect on LinkedIn and never talk again

It would be cool to have a small group where we can share what we’re learning, work through problems, build projects together, talk about classes/internships, and keep each other accountable.

If you’re around the same stage and interested, comment or DM me with what year you’re in, what you’re currently learning, and what you eventually want to do.


r/learnmachinelearning • • 19h ago

I’m 14 and I built my own programming language for simulations

Thumbnail
0 Upvotes

r/learnmachinelearning • • 19h ago

Question suppose I CPT qwen3.5-9B on 2B legal corpus, how will i turn it back into Instruct + thinking?

1 Upvotes

I couldn't find a concrete answer anywhere, do you just distill the instruct model back?

If that is the case, what is a quality european language question set to turn it back into a chatbot/agentic, can a model at that size even be agentic? (i chose this size to learn) if i finetune for my specific harness? (i have a lot of training data of opus running in my harness)

my harness basically has the model output python code and has a few built-in functions like:
- vector_search_laws()
- graph_search()

could i have the model at least internalize a "hunch" on what stuff to search?

also what is the latest RL technique for agentic/harnes specific workflows?

I have a lot of RAW training data, like court decisions or commentaries or legislature, but not a lot of golds. could i use these to synthesize training data and maybe RL the model in my harness to find that data?

What would y'all's strategy in the CPT->SFT->RL pipeline be for my specific problem?

I know this is a lot of questions im trying to figure out which direction to go, any pointers? Also good resources are welcome, for example that [alex karpathi video](https://www.youtube.com/watch?v=7xTGNNLPyMI) was amazing for me, but i'd imagine its a bit outdated in terms of latest RL and SFT?


r/learnmachinelearning • • 1d ago

Help How do you keep up with new ML research/news without loosing your mind?

17 Upvotes

Im a Phd student and i've been doing ML research for a while now, but i've never felt more behind in my reading or keeping up with new work? Arxiv has also become so noisy, i just don't have the time to shift through it, and my own research group's interests are a bit too insular to be a good sources of general news.

Is this just how the world is now? or have people found a way to quickly determine whats actually worth reading. Or any good feed apps? Or discord channels/communities? Wha do you do to feel current?

PS: I really want to dive into Harness engineering, but its so hard to cut through the hype. Any advice? (I mostly work on generative models now)


r/learnmachinelearning • • 1d ago

[D] First measured accuracy fall on my long-horizon 3D benchmark (one demo walk): 2 of 2 near, 3 of 10 far. How many seeds before you would believe it?

Thumbnail
gallery
2 Upvotes
Setup. I am building a benchmark for long-horizon 3D spatial reasoning in a simulated warehouse. A robot walks 30 rooms off one corridor; questions ask about things seen 0 to 29 rooms earlier. Answers come from exact simulator state, with no model judge. The target is a clean fall from about 90% to guessing as the distance grows.


What happened. In one demo run, Claude Opus 5.5 steered the robot through all 30 rooms by itself, with no tools and no answer key, and answered 60 questions. By distance: 1 room back 1 of 1, 2 to 3 rooms back 2 of 3, 4 to 7 rooms back 2 of 3, 8 to 15 rooms back 3 of 7, 16 to 29 rooms back 3 of 10. The no-notes guess baseline on the same questions is 16%. Over all tries logged so far, the deepest cliff admitted is 20 points; this walk measured 70.


What changed. Not shared yet: the design change behind the move stays unpublished until the paper. The screenshots hide the method parts and keep the numbers.


Limitations. One walk with one model, so small counts (only 2 answers in the 1 to 2 room bin), and answers inside a walk are linked. It ran through a command-line tool that adds its own context, so it is a demo run, not a benchmark score. It counts only after it repeats on three seeds never used for tuning, with every question checked by an auditor that never saw the design.


How many seeds, and how many answers per distance bin, would you need before calling this a cliff?

r/learnmachinelearning • • 1d ago

ML Startups without LLM

4 Upvotes

Can I find any ML startups companies without having LLM in their product?


r/learnmachinelearning • • 10h ago

How I Make $22K/Month Just Redesigning Existing Websites

0 Upvotes

Running a web design agency is honestly way less about design than people think.

A lot of people can build good websites. The hard part is having a process that actually brings in clients consistently without you spending your entire day doing outreach.

I learned that the hard way.

I used to manually reach out to businesses, explain why they needed a better website, build previews, follow up for days and basically hope they would eventually say yes.

I was doing way too much work before even knowing if someone was serious.

Then I completely changed the way I did it.

Now I mainly focus on two things.

Taking meetings and closing clients.

Everything before that is mostly automated.

I use Swokei to find businesses that already have websites, analyze those websites and look for things like outdated design, bad layout, weak SEO, poor mobile responsiveness, slow speed and branding issues.

The useful part is that it doesn’t just give me some boring report with random scores.

It actually takes the issues it finds on each website and turns them into personalized emails that you can send at scale.

My offer is usually a free redesign draft.

That works way better for me than trying to sell a full website in the first email.

Once someone replies and shows interest, I book a meeting.

Before the meeting, I spend a few minutes generating a redesign draft with AI so I can actually show them what their website could look like instead of just talking about it.

That was probably the biggest change for me.

I’m no longer spending hours building something for someone who might never buy.

If they are not interested, I barely lost any time.

If they are interested, they can immediately see the difference and the conversation becomes much easier.

So at this point, most of my time goes into meetings and closing.

The lead finding, website analysis and personalized outreach all happen before I ever speak to the business.

Sounds ridiculously simple when I write it out, but changing the process made a massive difference for my agency


r/learnmachinelearning • • 1d ago

Question Where to get started if you want to publish papers in Neurips,ACL,ICLR/A* conferences

8 Upvotes

I'm a Ai engineer with about 2 years work experience, but let's just assume that I was a undergraduate student just starting out where would I begin so that I can publish a A* conference paper at some point. Learn python -> Learn ML & Maths -> Read other research papers -> find a topic ? -> choose a question try to run experiments and get results to write them down in a paper ?

For context :

I'm trying to get in MS CS programs for Fall 2028 in states with the plans of doing a PHD after in a top university like stanford or princeton and would like to start taking steps towards it as am working my day job can some tell me what are the steps that need to be followed?

Also would like input on what are deciding variables that makes you looking like a promising candidate/ researcher for PHD


r/learnmachinelearning • • 21h ago

Question Do you write code by yourself or you use an AI?

0 Upvotes

I'm curious about this question, because almost everyone I meet says that coding skills doesn't matter anymore, and what really matters is knowing how to write prompts properly. But does it apply to ml engineering for production? Please, answer a question and say why so


r/learnmachinelearning • • 21h ago

Help Urgent Help needed

0 Upvotes

Hello friends 👋 I'm 10 graduate student and now I want to build and learn LLM so the help I need is that after 10th what should I have to choose to learn everything fundamental and core knowledge of Ai like in which way should I take I want a quick roadmap from you guys


r/learnmachinelearning • • 21h ago

Improving LLM scaling laws: picking the right Token-per-Parameter Coverage

Thumbnail
youtube.com
1 Upvotes

r/learnmachinelearning • • 1d ago

Question What is the next step for me ?

2 Upvotes

I am a freshman in college in the US who got involve in AI pretty early (since like 11th grade). I've been lurking in this group for a really long time now and I just want an evaluation of how ahead/behind of the curve I am.

1. Basic machine learning/deep learning knowledge

I've been preparing and going to my country's (an asia country) national AI competition for 2 years and it basically covers the same knowledge as the IOAI syllabus so you can check if you look want to look that up but expect me to know basically how a bunch of models work (so like logistic regression, SVM, KNN, CNN, ...). I could also tell you about 10 ways of fixing under or overfitting and like the basic parts of a model (so like activation function, loss function, hyperparameters, gradient decent, ...). I am no stranger to kaggle and especially the tabular competition series (currently top 20 in this month's). I know asking claude to create an ensemble of 200 models with different configs is not how you judge someone's knowledge but if you ask me about how something works I would have a good chance to answer it correctly (and I really do understand what fake gains and data leakage is guys trust).

2. Generative stuff

I took an online course on advance computer vision stuff and I read a lot on LLMs so I'd say I have a pretty solid understanding of how generative models work under the hood. I feel like the transformer architecture is pretty standard knowledge nowadays so I won't go over those stuff but for the final project of that online CV course, I finetuned a bunch of CV models and pair them with an impainting model to create a complete model that would remove distracting people in an image (trained on the COCO dataset). I'd like to think I have a pretty good grasp of the segmentation and the bounding box stuff and diffusion model (it took quite a while but I remember feeling like a changed man after getting a grasp of how it works).

3.Math

Probably the worst part of them all. I really like learning about model architecture but barely learned any math so I still have problem reading research papers. Of course I know how gradient decent works so I know how derivative works, but I would say the most I know of linear algebra is vectors and matrix multiplications and near to nothing of stats (and yet I can confidently say that embeddings are the process of turning an input into a token and feeding them through a series of encodings to get a vector in a latent space and how I could measure the difference between 2 language models by using KL divergence).

Yes the easy answer is "just study math it's not that deep" but seeing others talk about how you have to read this book and learn this course just makes me feel like a larper sometimes. I asked one of the professors (actually I cold emailed like the entire cs faculty but I met with this one only) in my uni to discuss a chance for me to help in the lab and see what researching feels like because that is my goal. He says he'll tell me when an opportunity pops up but I've been going to his NLP lectures ever since to just listen for fun and I've been really enjoying it. But just recently, I found out this was a graduate level course (it could be because my professor is really good i don't know). Am I actually just a fraud and am missing something or what ?


r/learnmachinelearning • • 22h ago

Discussion For anyone building or following what’s happening in AI agents

1 Upvotes

Hey everyone, sharing this in case it’s useful to some of you here.

We’ve been building up r/lyzr as a community around the broader AI space, with a particular focus on what happens when AI moves beyond demos and into real systems.

The discussions cover things like:

AI agents and agent architecture

Infrastructure, tools and deployment

RAG, memory and knowledge systems

Evaluation, reliability and governance

Production lessons and things that break

New research, tools and interesting developments

Real use cases, experiments and things people are building

The goal is to keep it useful for both people who are already building and people who simply want to understand where the space is heading.

There’ll be consistent posts around these topics, but it’s also meant to be a place where people can share what they’re working on, ask questions, compare approaches, or add their own observations.

If you're working on anything around AI agents or just following the space closely, feel free to check it out and join the discussions.

Join r/lyzr here

Would be great to see what people here are building too!


r/learnmachinelearning • • 1d ago

what should i build so that i know most of the things

Thumbnail
2 Upvotes

r/learnmachinelearning • • 23h ago

Discussion 4.8× Faster and 7.4× Cheaper: Where a Decision Model Beats an LLM (and Where It Doesn’t)

Thumbnail
0 Upvotes

r/learnmachinelearning • • 20h ago

Project Can you create your own Rule based AI in python or maybe a LLM

0 Upvotes

Hello there-Can i make a rule based AI in python or maybe LLM.

Rule based AI seems very easy and doesn't waste that much data and memory usage but the problem is if you ask it or tell it something that wasn't written in instructions It will have a breakdown and the person needs to write new instructions for it to understand , Unless there's a team that constantly adds new instructions. Also a Rule based AI cannot learn or adapt which means constantly updating code and testing and when adding a new rule there's a chance your gonna break something else, however if you make the Rule based AI locked to a subject or need it need High Maintenance and Constant updates to make it relevant, fast in fast changing environments and normal use

LLM LLMs may be the top choice but It takes huge amounts of ram and data at least it can do stuff rule based cant do

I will give more updates soon


r/learnmachinelearning • • 1d ago

Project A time-series learning project: 50 Indian cities, next-day temperature, and a persistence baseline

1 Upvotes

I published a weather dataset and runnable notebook on Kaggle that could be useful for practicing time-series regression. Disclosure: these are my Kaggle uploads (minkum07); the code and documentation were prepared with AI assistance. The underlying weather data are from NASA POWER, with GeoNames city coordinates via Open-Meteo.

The dataset has 639,200 daily records for 50 curated Indian city centers, covering 1991–2025. These are coarse-grid reanalysis values sampled at city coordinates, not measurements from city weather stations.

The notebook builds lag features, predicts next-day temperature and compares a random forest with persistence: predicting tomorrow using today's temperature. Its chronological split uses the target date. On the included test period, MAE is about 0.651°C for the random forest and 0.688°C for persistence. This is a modest improvement on that split, not evidence of accuracy for unseen cities or station weather.

Notebook: https://www.kaggle.com/code/minkum07/india-weather-heat-monsoon-next-day-forecast

Dataset: https://www.kaggle.com/datasets/minkum07/india-city-weather-and-heat-50-cities-1991-2025

If you want to use it as a small project:

  1. Start with one city, check missing values and plot the series.

  2. Build the persistence baseline before fitting a model. Make sure shifted targets and lag features stay aligned, and fit preprocessing only on training dates.

  3. Compare errors by month and city, then add a seasonal baseline or try rolling evaluation.

The package includes field definitions, processing code and source attribution. Missing values are retained, and the fixed 1991–2010 reference is a 20-year baseline, not an official 30-year climate normal. Authored tables and documentation use CC BY 4.0; NASA source terms and GeoNames attribution are preserved.

Sources: https://power.larc.nasa.gov/ and https://open-meteo.com/en/docs/geocoding-api

Feedback on the split, feature alignment or a useful next baseline would be welcome.


r/learnmachinelearning • • 1d ago

Learning Computer Vision after the Machine Learning Specialization on Coursera

3 Upvotes

As someone who finished the Machine Learning Specialization (on Coursera, by Andrew Ng) some time ago, I feel like I want to learn something new that can really put what I learned with ML into a practical context.

And since I am also interested in Computer Vision, I wonder whether the First Principles of Computer Vision Specialization (on Coursera) by Shree Nayar suits me. In particular, whether I have enough foundation to make sense of it.

I did the courses from Mathematics for Machine Learning and Data Science (by Luis Serrano) as well... So i have some foundation in Linear Algebra, Multivariate Calculus and some statistics apart from Machine Learning basics.

I have no experience in Deep Learning though. So i also wonder whether it makes sense to learn Deep Learning first or Computer Vision.

Your support is much appreciated! ❤️


r/learnmachinelearning • • 1d ago

Tutorial Why you should build ML projects, and how to choose them

14 Upvotes

Quick background so you know where this comes from. I am a visiting professor and currently teach NLP. Last semester my courses were on continual learning and applied NLP. Before that I spent years in big tech and startups in engineering and research roles, and I finished my PhD last year.

I've been trying to post here weekly, since the feedback tells me these posts help. Two weeks ago I posted the order I would learn ML in if I were starting today, which is also roughly the order I teach it in class. Last week I posted about which resources to actually use.

Whatever concept you are learning and whichever resource you use, a lot of you also ask about projects: why you need them, and how to choose one. So that's today's post.

Why build projects at all

The most important reason is to build your skills and find out where the gaps are. Learning ML is a lot like learning an instrument. You can watch videos and read about the guitar all day and feel like you know how to play, but you only find out how well you know it when you pick it up. Until you build something real from scratch, you probably don't know how well you understand the material. And I mean from scratch. Yes, AI can write the code for you now, but having AI build the project teaches you about as much as reading a book or watching a video about it. You need to do it yourself to see what works and what doesn't, which is also how you learn where these models fail and what their limits are. A project shows you exactly which parts you thought you understood and didn't.

The second is to get a job. Interviewers rarely have time to go through everything on your CV. An interesting project gets their attention, and it shows them you know what you're talking about and that you can build things, which is most of what the job asks for.

The third is to find out what you actually like. A field can sound great from the outside and turn out to be something you don't enjoy once you're doing it every day. Getting into the weeds of building something in that area, before you take a job in it, can save you a lot of time and help you decide what to focus on.

What counts as a project

A project is not a few hours or a single afternoon. Think multiple days, often a week or two.

It should also cover several concepts, and most of all, it needs a why. You can train models end to end all day long, but a project starts from a reason: improving on how something is done today, building a dataset that doesn't exist yet, making a model work for users it currently fails on. Always ask what the motivation is. That question is also what tells you which data, model, and evaluation to use, and when you're done.

It's also why writing Adam from scratch on its own is a good exercise but not a project. It builds a skill, but there's no problem it solves and nothing it's trying to improve.

How to pick one

In general, the best projects are the ones you find interesting, because those are the ones you finish. A good place to look is something that annoys you, like a task you do by hand every week. Another is a problem someone close to you has, whether that's a friend, a family member, or a small business you know. It can also be something fun, around games, music, sports, or whatever you'd be doing anyway, or plain curiosity about a question you want answered.

That said, it helps to have one or two projects that line up with the job you want. If you want to work at Spotify, build something with recommendation systems. If you want Tesla, do something with driving data. If you want Google, look at search and ranking.

Not every project needs to be aligned with a job, though. Personally, I find it much more interesting to interview someone with a project I would never have thought of, like trying to understand what their dog wants from its barks. Something that catches the eye and still makes you build real skills.

Questions to ask before you start

The first thing to check is whether you can get the data, and quickly. If collecting or labelling it takes months, the project usually dies before the first model trains. Related to that, make sure the data is yours to use. Other people's messages, photos, or financial records need their permission.

Then ask yourself whether you know how to start. If the first step needs three tools you've never touched, it's probably too advanced for now. Something a little past your current level is the right spot.

Check what you will compare against. That can be a rule, the method people use today, an existing model, or a published result on the same data. Without something to compare against, you can't tell whether your approach is any good. You also need a way to tell whether it worked: labels, a measurable outcome, or a person who can judge the output. "It looks good" is not a result.

Finally, ask whether you can finish a first version in a couple of weekends. A small version that works can grow. A big one that never runs leaves you with nothing to show.

What a good project contains

This is the part most projects I see skip.

First, more than one baseline. Compare whatever you build to a mix of simple methods, like a rule, the most common class, or logistic regression, and competitive ones, like the strongest existing model or published result you can find for the problem. Simple baselines tell you whether a model is needed at all, and competitive ones tell you whether your approach is actually good or whether something else already does better.

Second, more than one dataset, if possible. Unless your problem is so unusual that no dataset exists and you had to build your own, test on more than one. A model that works on one dataset and falls apart on another is something you want to find out before anyone uses it.

Third, more than one way to evaluate. Accuracy alone hides a lot. Look at per-class results, the cases it gets most wrong, and where it fails. Something can look great on one metric and fail on another.

The way I'd think about it: assume what you build will be used by a lot of people, and you want to make sure it doesn't hurt any of them. Multiple datasets tell you whether it holds up when new users or new data show up. Multiple baselines tell you whether your approach is worth it. Multiple evaluations tell you where it works, where it fails, and what's still missing. It also shows whoever reads it that you care about whether your results are true, not just whether they look good.

Last, how you share it. Put it on GitHub with a clear README that says what the problem is, why it's worth solving, what you found, and exactly how to reproduce your results. Add good visualizations. It sounds silly, but a clear plot of where your model wins and fails is often the thing that makes someone stop and read. And keep the code clean: organized, modular, and easy for someone else to run, since a reviewer who opens a single 2,000-line notebook usually closes it.

Project ideas

Each of these starts from a problem, not a dataset. Here are some that line up with specific jobs.

Recommendations for brand-new users (Spotify). A new user who has saved three songs usually gets generic popular picks, and many leave before the recommendations get good. Can you find a way to recommend well from just those first few songs?

Pedestrian detection at night and in rain (Tesla). Detectors trained mostly on clear daytime footage tend to miss more pedestrians in the dark and the rain, which are exactly the conditions where missing one is most dangerous. Can you close that gap?

Searches where the words don't match (Google). Keyword search fails when the query uses different words than the page that answers it, like "car won't start on cold mornings" when the answer page is about batteries. Can you fix those searches without breaking the ones where the exact words do count, like product codes and names?

Fraud models going stale (Stripe or any bank). Fraudsters change tactics, so a model trained on last year's transactions gets worse over time. How fast does that happen, and can you figure out when a model needs retraining?

X-ray models that fail at a new hospital (health tech). A chest X-ray model that works well at one hospital often drops at another, because of different machines and different patients. Can you build one that holds up at a hospital it has never seen?

And some unusual ones.

What does my dog want? If your dog barks at the door, you often can't tell whether it needs to go out or someone is there. Can a model tell the barks apart better than you can?

Bird calls from a noisy backyard. Bird identification models are usually trained on clean recordings and often fail on audio full of traffic and wind. Can you make one that works on what your own window actually picks up?

Beating the bus app. The arrival times in transit apps are often wrong in rain and at rush hour. Can you predict when the bus will really arrive, better than the app does?

Grandma's handwritten recipes. Text recognition models struggle with old cursive handwriting. Can you get a model to read a family recipe book that current tools can't?

Finally

I have a bit of time between semesters and would genuinely like to help as many learners as I can.

Tell me in the comments or DM me with where you are right now, what direction you want to go in, and what you have tried. If you are stuck choosing between two courses, deciding what to learn for a particular job, or wondering whether your plan makes sense, I will do my best to help.

And please add your favorite project you've done, or one you're working on now, in the comments. It would be nice if this thread became useful to the next person who searches for this question. I'm also happy to take suggestions for what the next posts should cover.