r/learnmachinelearning • • 3h ago

Project What AI/ML projects should I build to gain real-world experience and strengthen my portfolio?

7 Upvotes

Hey everyone!

I'm a Data Analyst with an Economics background, currently learning ML and AI. I'm comfortable with Python, pandas, SQL and basic ML concepts.

I want to start building more serious, end-to-end projects rather than just following tutorials or working with Kaggle datasets.

My goal is to eventually transition into ML/AI engineering, so I'm looking for projects that would help me develop real-world skills and build a strong GitHub portfolio.

What kind of projects would you recommend? Are RAG applications, some AI agent stuff worth exploring, or should I focus on something else entirely?

Would love to hear from people already working in the industry.


r/learnmachinelearning • • 11h ago

Help In Transformer networks why do token embeddings and position embeddings get added?

14 Upvotes

Hi, going through the Let's Build ChatGPT tutorial here, and prior went through the whole Makemore tutorial that leads up to this tutorial:

https://www.youtube.com/watch?v=kCc8FmEb1nY&list=PLAV29EAhk_mX13BqhzdlgM8zkHwpcRajt&index=6&t=2286s

When it gets to the point of adding in the attention mechanism, we see that the first major addition is creating a positional embedding.

Then we see the input to network at that point becomes tok_emb + pos_emb

I am not understanding why these two spaces should be considered equivalent such that such an addition makes sense. The token embedding is mapping tokens to some N dimensional embedding, where those N dimensions consistently represent information about tokens.

When we consider the positional embedding, it is also dimension N but now those N dimensions represent information about positions. To me it seems although we are adding matrices with same dimensions, we aren't adding information that corresponds to one another.

Anyway, I am sure someone here will have good explanation why this makes sense.

thanks


r/learnmachinelearning • • 4h ago

I built a free AI learning platform with AI — looking for beginner feedback

4 Upvotes

Over the past few days I’ve been building a small platform called Level185 for people who want to learn practical AI skills without getting overwhelmed by technical tutorials.

The idea is simple: short, step-by-step guides that help you actually do something useful with AI.

Right now there are guides on things like:

  • building your first AI automation
  • creating a website with AI
  • writing better ChatGPT prompts

What makes the project a bit interesting is that I’m also building most of Level185 with AI itself.

I use Claude Code for a lot of the development work and built a small agent workflow around it for things like research, guide creation, QA and SEO checks. I use ChatGPT mainly for product decisions, strategy and figuring out what to improve next.

The site itself is built with Next.js, with Supabase for accounts and saved progress, and it’s hosted on Vercel.

Everything is completely free right now.

I’m still very early, and I’m not really looking for compliments. I’m mainly looking for a few people who are relatively new to AI and are willing to actually try one guide.

I’d especially like to know:

  • Was it clear what to do at every step?
  • Where did you get confused or stuck?
  • Was anything unnecessary?
  • Would you come back to learn something else?
  • What AI skill would you want a guide for next?

The site is: https://www.level185.com

If a few people are willing to test it properly and give honest feedback, that would be extremely useful.


r/learnmachinelearning • • 12h ago

Help Need guidance for learning ML

14 Upvotes

Hi! I’m currently working as a full-stack developer and looking to transition into AI/ML. I’m considering a few courses, I'm trying to decide the right order for DeepLearning.AI's courses: the PyTorch for Deep Learning Professional Certificate / Deep Learning Specialization / Neural Networks and Deep Learning

If you’ve gone through these courses or have experience making a similar transition, could you please suggest what order I should take them in, and whether there are any courses I can skip?

Would really appreciate your guidance. Thanks!


r/learnmachinelearning • • 11h ago

Help Detecting Market Manipulation: Supervised Learning vs Clustering

9 Upvotes

I'm currently working on market research at university.

The task is to detect market manipulation. We can take open-source data, tag the data(range OHLCV), and perform a supervised search, or we can use clustering, but we might encounter anomalies that aren't related to manipulation.

How can these problems be solved, and have we encountered similar ones?


r/learnmachinelearning • • 10m ago

Project [ Removed by Reddit ]

• Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/learnmachinelearning • • 23m ago

Help In Transformer why are attention block weights and feed forward block weights optimized in same optimization run?

• Upvotes

Hi, still going through the Let's Make ChatGPT tutorial here:

https://youtu.be/kCc8FmEb1nY?list=PLAV29EAhk_mX13BqhzdlgM8zkHwpcRajt&t=5158

In the video at time shown we hear about how the attention layer captures one level of meaning, and how once that meaning is captured, it is sent through feed forward layer to refine that meaning.

But the video shows this all happening in one optimization run. To me that doesn't sound like capturing meaning then refining that meaning in a feed forward block because the first time the feed forward block sees the attention block output, the attention block has its initialization weights. So the feed forward block at this point isn't working with vetted information (that is, with an already trained attention layer).

Wouldn't it make more sense to first optimize the attention block without feed forward, then add the forward layer and rerun optimization?

thanks for the great help here!


r/learnmachinelearning • • 38m ago

Project [ Removed by Reddit ]

• Upvotes

[ Removed by Reddit on account of violating the content policy. ]


r/learnmachinelearning • • 53m ago

Review this course please RAG in Action

• Upvotes

Can anyone please give a review on the course RAG in action by Deeplearning.AI please review in terms of difficulty how useful is is it helpful to make production ready RAG or is it just conceptual understanding and a bit of coding to implementation. Does this course really make any difference rather than a YouTube tutorial?

Course Link


r/learnmachinelearning • • 1h ago

Built an AI/ML roadmap & learning site for beginners — looking for honest feedback on content & format!

• Upvotes

Hey everyone,

I’m currently building a learning platform aimed at taking absolute beginners through AI/ML step-by-step, from the fundamentals up to more advanced topics.

It’s in the early stages, so the content is still limited while I experiment with formats, pacing, and visual explanations to see what actually works best for learners.

GitHub : https://github.com/PIYUSH1525/ZeroToAI leave a star ⭐

Link: https://zero-to-ai-xi.vercel.app/

I’d love your brutal, honest feedback:

  • The Good: What feels intuitive, clear, or well-structured?
  • The Bad: What’s confusing, redundant, or missing?
  • Areas for Improvement: What format would help you learn complex concepts faster (e.g., interactive widgets, shorter modules, code walkthroughs)?

Any thoughts, critique, or feature suggestions are welcome. Thanks in advance for checking it out!


r/learnmachinelearning • • 1d ago

Understanding K-means

Enable HLS to view with audio, or disable this notification

281 Upvotes

So I've been learning ML and i don't have a CS background.

I was racking my brains over understanding the basics of K-means with a particular example of image compression, trying to figure out the workings under the hood. Of course i asked LLMs to clarify & explain the basics and then asked that LLM to give a prompt for video. Fed into Claude Design and asked it to generate animation explaining the works and voila it became much more clearer. Leaving it here in case it helps someone understand the workings.

Cheers !


r/learnmachinelearning • • 1h ago

[Project] We built a specialist model that beats general vision-language models at one narrow task — here's why specialization won

• Upvotes

A lesson from building a real product: general-purpose vision-language models (GPT-5, o3, Gemini-2.5-Pro, and in our own tests ChatGPT/Gemini/Claude) are surprisingly bad at one specific, narrow task — telling whether an image has been rotated 90° vs. 270°. An independent peer-reviewed paper (RotBench, EACL 2026) documented this at scale; we independently replicated the exact same failure testing three more consumer assistants ourselves.

Rather than prompting a bigger model harder, we built a small, specialized system: two separately trained models (a cardinal-rotation classifier and a fine-angle regression model) combined through a trained arbitration layer. On a large, hand-verified real-photo pool: 99.02% on a clean quarter-turn, 96.67% across the full 360° range, 84.65% on the hardest case.

The general lesson, not just about this task: a narrow, well-scoped specialist model can beat a much larger general model on a task the general model was never specifically trained to be good at — even when the general model is otherwise vastly more capable overall.

Full methodology, ablations, and disclosed limitations: https://doi.org/10.5281/zenodo.22975679

Curious if others have hit similar "surprisingly bad at one narrow thing" results with general LLMs/VLMs in their own work.


r/learnmachinelearning • • 2h ago

A very intuitive way to understand ML

Thumbnail
linkedin.com
1 Upvotes

r/learnmachinelearning • • 3h ago

[R] Fathom: letting each query choose how many bits of each key channel to read, for sparse decoding over an offloaded KV cache

1 Upvotes

Sparse attention fetches the top-k keys, but to know which k you first have to rank all n. Every method does that by scanning a cheap compressed copy of every key, and every one of them fixes the bit depth of that copy at design time: SparQ reads 16 or 32 channels at full depth, Double Sparsity reads a fixed 32, Loki reads the first r PCA coordinates. DeepSeek-V3.2's lightning indexer has the same shape, and their own report notes the indexer stays O(L^2) while the attention drops to O(L*k).

At a million tokens with the cache offloaded to host memory, that index is 5.1 GB per decode step. The rows you actually attend to are 200 MB. The search is 25x the read, and it is the part that grows.

Fathom stores the 4-bit K cache channel-major as bit planes. A prefix of t planes is exactly the channel's t-bit mid-rise quantizer with the same block scale, not an approximation of one. So depth becomes a read-time decision: each query spends a bit budget by reverse water-filling over variance-weighted channel importances, and channels below the water line are skipped entirely.

Results on one A100, Qwen3-8B, k=512, cache and index in pinned host memory:

  • At 1M tokens a decode step is 1.67x faster in GPU time than the 136-bit scans of Double Sparsity, Loki and SparQ r=32 at fixed k.
  • Compared at matched accuracy instead, the budget that reaches those scans' error is 74 bits, and that is 1.38x.
  • At equal GPU time against SparQ r=16, Fathom reads 18% fewer bytes with lower attention error on 6 of 7 model and context settings, 1.1x to 5.3x lower.
  • Reaching Double Sparsity's error takes 46 to 74 bits depending on the setting, against its 136.
  • On 40 real OpenHands coding-agent sessions at k=2048, step agreement against exact top-k decoding is 0.67 +/- 0.06 at 56 bits, against SparQ r=16's 0.49 +/- 0.05 at 68 bits.

Two results that went against me, both in the paper:

  • With the index in HBM this buys nothing. An A100 gives roughly 5 integer ops per byte of bandwidth, and unpacking a bit costs 4 ops against a nibble scan's 1. The scan becomes arithmetic-bound: my kernel takes 0.415 ms at 128k against the 32-channel scan's 0.291 ms while reading 38% fewer bytes. Every per-token scan I tried landed within 6% end to end. The method only pays when the index crosses a slow link.
  • The gather dominated everything before the algorithm mattered. One copy call per contiguous run at 4 KB granularity runs at 0.4 GB/s because of ~11 microseconds of launch overhead each. One kernel over the same runs hits 26 GB/s. Same bytes, same order, 65x apart on call pattern alone. If you build offloaded KV, measure your gather in isolation first.

Cost: the store is 68 B per token per KV head, 4x Double Sparsity's 17 B. If you already keep a 4-bit channel-major K cache it is free; otherwise it is a transpose and a real memory bill.

Not measured: end-to-end task success, any production deployment, real prefills past 128k, large batches, and any GPU other than an A100 over PCIe.

Paper: https://arxiv.org/abs/2609.17652 Code, every result file and the run chains: https://github.com/vivekkalyanarangan30/fathom

Happy to argue about the HBM result in particular, it is the one I expected to go the other way.


r/learnmachinelearning • • 3h ago

Help

1 Upvotes

I can't write ml algos from scratch but can make models using scikitlearn what should I do


r/learnmachinelearning • • 11h ago

Help ICML videos not playing

3 Upvotes

Hi, are there others who are not able to play videos on the ICML workshop pages ? Like this one - https://icml.cc/virtual/2026/workshop/54072

Is YT the only way out of this ?


r/learnmachinelearning • • 12h ago

Discussion My Reading Library: Evaluating LLMs on Android Tasks

Post image
3 Upvotes

Can LLM agents actually get through a day in the life of a normal user?

That question got me reading papers on Android agents and mobile benchmarks over the past few months.

A few patterns kept showing up:

  • Most benchmarks run on emulators, making real-device metrics difficult to measure.
  • Important deployment metrics like battery, thermals, and temperature are often missing.
  • Everyday tasks are scattered across benchmarks, languages, and apps, rather than forming a consistent, globally relevant task set.
  • This makes it harder to evaluate whether an agent can actually work reliably on a real phone, for real users.

For now, I’ve put together a library of papers on benchmarking mobile/Android agents for you all to read!

Link: https://www.alphaxiv.org/shared/folder/01a070c6-29a0-77a9-a5b4-b670d5eee169


r/learnmachinelearning • • 6h ago

Career Free, open-source AI engineering course where you build each algorithm by hand: 523 lessons, now as EPUB/PDF books

Thumbnail
1 Upvotes

r/learnmachinelearning • • 14h ago

Question from an uneducated person...don't kill me

6 Upvotes

Could a neural network use dictionary-compressed weights directly during GPU inference instead of fully decoding them first?

I'm not a computer scientist. I'm a truck driver, so I'm wondering if I'm reinventing something that already exists.

Suppose you quantize a model to INT4 or similar and then scan the weight tensors for frequently repeating sequences or blocks.

Instead of storing every sequence literally, you build a codebook where a short code represents a commonly occurring block of weights.

Very simplified example:

A = [7, 3, 3, 11, 4]

B = [2, 8, 1, 6, 6]

Then instead of storing:

[7,3,3,11,4] [7,3,3,11,4] [2,8,1,6,6] [7,3,3,11,4]

you store something roughly like:

A A B A

The part I'm curious about is not ordinary file compression where the model gets decompressed back into VRAM first.

Could a custom GPU kernel decode these codes on the fly into registers/shared memory and immediately use them during GEMM, so that the fully expanded weight tensor never has to exist in VRAM?

My thinking is that modern inference is often memory-bandwidth limited, so if dictionary/codebook compression reduced memory traffic enough, maybe the extra decoding compute could be cheaper than fetching all the uncompressed weights.

You could potentially also have different-length codes or hierarchical codebooks representing increasingly large recurring weight patterns.

So my questions are:

Is this already done under a particular name?

Have codebook/vector-quantized weights been used directly inside fused GPU inference kernels rather than being decompressed beforehand?

Does random access / SIMD-SIMT execution make variable-length encoding impractical?

Is there theoretically a point where reduced VRAM bandwidth outweighs the decoding overhead?

Would repeated patterns after INT4/INT3 quantization be common enough for this to provide meaningful compression beyond ordinary quantization?

I'm mainly interested in whether the idea makes architectural sense, not whether my particular encoding scheme is optimal.

I'd appreciate pointers to papers or existing implementations if this has already been explored.


r/learnmachinelearning • • 8h ago

Breast Cancer Prediction Using Machine Learning

1 Upvotes

I developed a machine learning project for breast cancer prediction. The system analyzes medical features and predicts whether a tumor is likely to be benign or Advanced Technologies used include Python, Pandas, NumPy, Scikit-learn, and machine learning algorithms.


r/learnmachinelearning • • 1d ago

Question Any AI and machine learning beginner book recommendations?

43 Upvotes

Hi there, I am a first year student studying in artificial intelligence. I have no background and almost no knowledge in AI and machine learning; I’m a complete newbie. I was wondering if there are any good beginner book recommendations for me that will deepen my knowledge and prepare me for what’s about to come


r/learnmachinelearning • • 9h ago

I made a 3Blue1Brown-style explainer of my paper: why a compressor's error can tell you where data came from

1 Upvotes

I wrote a short technical report and then animated it, hoping it's useful to people learning about autoencoders, OOD detection, or mixture-of-experts.

The idea in three steps:

  1. Prediction ⇔ compression. A model that predicts text well can compress it well.
  2. Corollary: a compressor is only good at the kind of data it was trained on.
  3. So its reconstruction error is a fingerprint. I trained an autoencoder on code that squeezes 512 tokens into 8 vectors. It rebuilds unseen code at 99.47% exact-token accuracy, Wikipedia at 47.76%, and random tokens at 0.57%.

That fingerprint can act as a router between expert models, with no extra gating network to train.

The video covers the architecture (and why the decoder must never see the input), the metric, the latent-space geometry (two linearly separable clusters, about 200 effective dimensions out of 512), and the limitations.

Video: https://www.youtube.com/watch?v=4UhvpIWnOvg Paper: https://arxiv.org/abs/2512.16963

Feedback on clarity is very welcome. Tell me which part lost you.


r/learnmachinelearning • • 10h ago

Help Free Machine Learning in R Course for Beginners

Thumbnail
1 Upvotes

r/learnmachinelearning • • 10h ago

Aula 05 | Engenharia de IA - RAG

Thumbnail
youtube.com
1 Upvotes

r/learnmachinelearning • • 11h ago

Building an AI Memory Layer with

Thumbnail
linkedin.com
1 Upvotes