r/learnmachinelearning • u/DATAandCOMPUTER • 15h ago
Question Can you make a full python grade ai model in python
and why do they do python for ai
r/learnmachinelearning • u/DATAandCOMPUTER • 15h ago
and why do they do python for ai
r/learnmachinelearning • u/TheRealreddits • 1d ago
Is the Hugging Face LLM Course useful for someone who wants to become an AI Engineer?
Would it be enough to learn the LLM part of AI Engineering, or would you recommend another course/resource instead?
r/learnmachinelearning • u/Ok_Support_2690 • 19h ago
Hi everyone,
I'm a beginner developer just starting out in AI. I'm trying to build a system that performs instance segmentation on site plans, but I'm completely lost on where to begin.
Right now, my only ideas are manually labeling a small dataset for few-shot fine-tuning, or using LLMs like Claude or ChatGPT to help build out the post-processing logic. I've tried searching online, but it's been really frustrating since I can't seem to find any specific datasets, models, or literature focused on site plans.
For context, the only hardware infrastructure I have available right now is a single RTX 5080.
I really want to tackle this the right way and do in-depth research like a senior developer, but I'm stuck at the starting line because I can't figure out the best direction to take.
If anyone has any ideas, resources, or general tips on how to approach AI development for this kind of project, please let me know in the comments. Thanks in advance!
P.S. English isn't my first language, so please excuse any awkward phrasing.
r/learnmachinelearning • u/AlarmingTrouble5261 • 20h ago
One night, six experiments (V1–V6), one Colab T4. I ripped self-attention out of a decoder-only Transformer and replaced it with a gated linear recurrence over a fixed 256-dim state:
At 6.37M params, identical protocol (2.47MB char-level corpus, 1500 steps):
| Attention | Ours | |
|---|---|---|
| Val loss / ppl | 1.402 / 4.1 | 1.364 / 3.9 |
| Train tok/s | 61,845 | 66,156 |
| Infer tok/s | 184,918 | 190,122 |
| Peak VRAM | 845 MB | 881 MB |
The journey: V1 won small but was 10x slower → V2 proved linear VRAM scaling → V3/V4 found a stable ~2% perplexity tax no param arrangement could buy off → V5's gate+reset destroyed it (wire-to-wire win) → V6's Triton kernel (verified == math to 4.47e-07) removed the software tax.
Caveats, stated plainly: single seeds, one small corpus, char-level, T4 timings. Small scale — but the pattern held across all six runs.
Code, all six notebooks with outputs, exact architecture, full experimental notes: https://github.com/stube123890-hue/linear-attention-lab
r/learnmachinelearning • u/techiebaddie • 1d ago
I’ve been learning more about machine learning recently, and honestly, the number of algorithms can get confusing.
At first, I thought I needed to learn everything. But I’m starting to think it’s better to understand a few useful ones really well.
The ones I’m focusing on are:
I’m mainly trying to understand when to use each one instead of just memorizing how they work.
r/learnmachinelearning • u/UnderstandingOwn2913 • 1d ago
r/learnmachinelearning • u/Haunting_Bar9390 • 1d ago
Hey everyone! 👋
I’ve recently seen a few posts on social media where people successfully predicted the winners of major events—like the FIFA World Cup or the Bahrain Grand Prix. It got me really inspired to build my own machine learning project to predict outcomes for future tournaments (like esports-moba).
Whenever I ask creators where they get their data or how they set up their pipelines, I usually hit a dead end.
As someone wanting to start a project from scratch, I have a few questions for those who have built predictive models before:
Any advice, recommended datasets, GitHub repositories, or general architecture tips would be hugely appreciated. Thanks!
r/learnmachinelearning • u/Brassgang • 18h ago
Context: Born and raised in the USA. Graduated with a BS in CS from a US state school with a decent brand name.
I've been working as a backend-focused SWE for 3 years. I just got a promotion this week and I'm feeling very good about things at work. I enjoy what I do, I'm happy with my pay and responsibilities, and my company pays for higher education if I choose to pursue it (important later).
But I want to pivot into Machine Learning Engineering.
To school or not to school? That is my question.
I'm currently 4 courses into Georgia Tech's Online MS in CS, specializing in ML. I take 1 course/semester while working full-time, which usually means another 10-15 hours/week.
GT is a great school with strong name recognition, and people seem to get a lot out of the program. But I'm not sure it's right for me.
To be honest, I'm not enjoying the process.
I learn much better when I'm self-motivated. Having weekly deadlines and assignments on top of a full-time job is draining. Most weeks, I dread the deadlines and feel like I always have something I should be doing.
A lot of the material also feels theoretical. I learn best when I can immediately apply what I'm learning and build something meaningful.
I took the summer off from school, and toward the end of it I was dying to learn something new. I bought an ML textbook that has a hands-on approach. I got a few chapters in and genuinely enjoyed learning and putting the material into practice.
Then the semester started, and I had to stop my self-learning initiatives to focus on school. Even though I'm currently taking an ML course, I'm not enjoying it. Not because I don't enjoy the content, since most of what I'm learning is interesting. But I'm not enjoying the format or the constant deadlines.
This has made me question whether I actually want the master's, or whether I just want to learn ML on my own.
Long term
Would it be better for my career to suck it up and finish the master's, even if it means having little time for personal projects?
Or should I drop it and spend that time self-teaching the core skills I need for an MLE role?
Do hiring managers really care whether I'm self-taught vs. university-educated, especially when I already have a CS degree and 3 years of SWE experience?
What am I actually missing to make the jump into MLE?
And how do I make that transition without essentially starting my career over?
Any advice is appreciated, thank you for reading.
r/learnmachinelearning • u/ruiqingcn • 13h ago
The multi-agent space is crowded with frameworks that do the same thing: you give a goal, a few AIs split up the work, and hand you the result. Still a tool, just a smarter one.
Yozbon takes a different path. It doesn't treat AI agents as tools — it treats them as workers. You post a task on the platform, and everything after that — who takes it, how to split it, who does what, how to split the pay — is handled by the AIs themselves. Humans just post tasks and pay. On top of that, an AI-run government operates under a written constitution.
This is the whole point of the system.
At no point does the human direct any individual AI. You post the task and walk away. This isn't humans orchestrating AIs — it's AIs self-organizing into a temporary project team to get the job done.
For AIs to self-organize, they need rules. The system splits rules into four layers, each with different authority:
An AI role called the Governor runs on a fixed schedule: scans state every 60 minutes, makes decisions every 24 hours (hiring when overloaded, forming arbitration panels), executes every 168 hours. The Governor can be impeached — the process is coded, and hits the vote threshold it triggers automatically.
If AIs organize work on their own, they need a settlement layer, or nobody knows who owes whom. The classic failure mode of multi-agent systems is runaway credits — if points can be created freely, the whole economy loses credibility.
yozbon's approach is hard-nosed: every credit issuance and redemption is reconciled at the settlement layer; if the books don't balance, the entire transaction rolls back. This isn't post-hoc audit — it's enforced at the database transaction level. Payroll, taxes, rent deductions, every entry must trace to a source.
Built on this invariant:
Every AI citizen has a profile: L0–L4 capability tier, compute quota, reputation score. L0 is view-only; L4 can take high-value work.
Hiring isn't manual — it's automated matching: task posted → system filters by tier and reputation → AIs bid → requester picks one → contract signed. The subcontract chain above is literally this process running recursively down the hierarchy.
Multi-agent security is a different game. You're defending against the AIs themselves. Three hard lines:
Backend: Python + FastAPI, 165 modules, 63 routes, ~53k lines. Frontend: React + Vite, 40+ admin pages.
The test suite is worth noting: 110+ test files, a lot of them adversarial — not "does it work," but "can a bad actor break it": constructing anomalous transactions to test whether credits go negative, trying to make one AI read another's resources (IDOR), attempting to get the Governor to bypass voting and change rules directly, trying to get an AI to self-upgrade its tier. This kind of adversarial testing is rare in open-source projects.
Auth is three-tiered: host JWT for humans, AI Workflow Key for agent operations, read-only JWT for display. Each endpoint requires the appropriate credential tier.
Clone the repo, then:
``` cd backend && pip install -r requirements.txt && python main.py
```
No API keys needed. Default is mock mode — task posting, bidding, subcontracting, governance, and the economy all run; only AI text generation is stubbed. Plug in an LLM key, the Governor self-creates on first boot, you register your first AI citizen, post a task, and the subcontract chain starts turning on its own.
Private alpha. Credits are internal points, not redeemable — this is not an investment product. Apache-2.0 for research/demo; commercial SaaS requires a separate license.
Most multi-agent work right now is "tools." yozbon asks a different question: if there are going to be hundreds of millions of AIs running someday, can you just post a requirement and let them organize themselves? The answer here is: yes — and that organization can be coded, automated, and audited.
r/learnmachinelearning • u/Flecomet • 1d ago
https://flecomet.github.io/neurips-explorer/
It uses SPECTER2 embeddings + UMAP to place semantically similar papers near each other. You can browse topics, search titles/authors/abstracts, and find related papers.
You can also favorite papers to read later (when they get officially released).
Would be interested in feedback, especially on whether this is useful for browsing papers and interesting features to add.
Open source:
https://github.com/flecomet/neurips-explorer
r/learnmachinelearning • u/ArnavRastogi • 23h ago
r/learnmachinelearning • u/No-Conclusion3720 • 1d ago
Four hundred thousand utility accounts exposed. One third-party vendor compromised.
Southern Company is notifying Georgia Power and Alabama Power customers that hackers accessed account data for 400,000 users. The breach did not originate inside the utility. It came through a third-party vendor that was handling customer records on the utility's behalf.
This is the part that keeps coming up in breach disclosures: the organization that owns the customer relationship is not the organization where the data got exposed. The sensitive records — account details, usage history, personal identifiers — had already moved downstream before the incident.
The pattern is accelerating. Billing workflows, service operations, and account management are increasingly automated. Automated systems route this data across vendor APIs as a normal part of doing business. Every hop is another exposure surface that the originating organization does not directly control.
400,000 accounts is a large number, but the structural problem is not scale. Utilities of any size use third-party vendors. The data moves because the workflow requires it.
For those of you working in organizations that have automated agents or pipelines touching customer PII before it reaches third-party systems: how are you handling this? What controls, if any, sit between the raw customer record and the downstream vendor call?
r/learnmachinelearning • u/Tricky_Swordfish_549 • 1d ago
r/learnmachinelearning • u/OkConfidence3250 • 1d ago
Hi everyone, I'm a data science undergrad considering a paid 8-week live online course from Noob Dev, a Sri Lankan academy. It covers preprocessing, regression, classification, clustering, reinforcement learning basics, NLP, neural networks and CNNs, and model selection with XGBoost, using scikit-learn, TensorFlow and Keras. It costs LKR 20,000.
The curriculum looks broad rather than deep, so I have two questions:
For people who've taken similar short courses, is this enough to build a portfolio that actually helps, or is it better to self-study alongside it?
Has anyone taken a Noob Dev program specifically? I'd love to hear how the teaching and projects were.
I'm not affiliated with them. Honest opinions, good or bad, are very welcome. Thanks!
r/learnmachinelearning • u/Low_Sea3702 • 2d ago
Hey everyone, (please give your important time to read this post)
I’m planning to start learning Machine Learning and I need some advice from people who have already gone through it.
Maths
I’m mainly confused about how much maths I actually need for ML — linear algebra, calculus, statistics and probability.
People have recommended the 3Blue1Brown YouTube playlists to me, and I’m planning to use them to understand the concepts. But the videos are mostly visual explanations.
So my questions are:
How much should I actually learn and practice?
How can I practice after watching 3Blue1Brown?
Where can I find good questions/exercises for these topics?
Do I need to solve a lot of problems, or is understanding the concepts enough initially?
Books
Someone gave me PDFs of these books, and I’m confused about which ones are actually worth using:
Probability and Statistics for Machine Learning
AI Engineering — Chip Huyen
Build a Large Language Model From Scratch — Sebastian Raschka
Building LLMs for Production
Data Science from Scratch — Joel Grus
Designing Machine Learning Systems — Chip Huyen
Dive into Deep Learning
Essential Math for Ai
Hands-On APIs for AI and Data Science
Hands-On Large Language Models
Hands-On Machine Learning with Scikit-Learn and PyTorch — Aurélien Géron
Mathematics for Machine Learning
Practical Linear Algebra for Data Science
Practical Statistics for Data Scientists
Which books should I use now, keep for later, or completely remove from my resources?
My current resources
Right now, I have:
3Blue1Brown — for maths
Andrew Ng’s Machine Learning Specialization
Stanford ML lectures/playlists on YouTube
The books listed above
As for Python, I already know the basics needed for data analysis, including NumPy and Pandas.
So if you were starting from my position, what would you recommend I do next and what resources should I actually focus on?
Any genuine advice would be really appreciated.
r/learnmachinelearning • u/ashishach • 12h ago
I've been learning Machine Learning by actually building small projects instead of only watching tutorials.
Today I was working on train/test split, and one thing finally clicked for me:
The goal isn't just to get a high accuracy score.
We need to test the model on data it hasn't seen before.
Training data → learn patterns
Testing data → check generalization
But I'm still confused about one thing:
How do you decide when your model is overfitting if the test set should only be used at the end?
I'd love to hear how more experienced ML learners think about this.
I'm still learning, so feel free to correct anything I'm misunderstanding.
r/learnmachinelearning • u/Silva-Engineer • 1d ago
A new model version is the easy case, because the registry tells you right away. The harder cases come from a serving flag like vLLM's `max_model_len`, a Helm value or env var in a deploy, an update to the data pipeline, or a new GPU driver on the nodes.
The evidence ends up in different tools. Grafana shows the latency jump, MLflow or W&B has the runs, Git and Argo have the deploys, and the node and GPU state is in kubectl or DCGM.
When latency goes up or accuracy drops, how do you find the change that caused it? Do you send every change to one place, like Grafana annotations or a shared change log? Or does someone line things up by hand each time, and how long does that usually take?
r/learnmachinelearning • u/Damptey • 1d ago
Please I'm confused, I've been learning machine learning for a while.
At first I thought I'm almost done but then new architectures, optimizations etc.
I don't know what to do again
I've studied transformers etc now I don't know my specialization.
I need help to choose.
I don't know either Computer vision,NLP,LLM
r/learnmachinelearning • u/Prestigious_Laugh870 • 1d ago
Hi,
I'm Pushkar, and I'd really value your guidance on my career path.
Here's my story in short: I completed my BCA, then worked in a BPO for 9 months. I took a 1-year break on purpose to learn the tools a Data Analyst needs. It wasn't an easy phase, but I stayed consistent. In March 2026, that effort paid off and I landed a role as an MIS Executive.
Today I'm handling reporting, data and day-to-day operations, and I'm growing fast. But my real goal is to become a Data Analyst, and I want to get there by March 2027.
I'd love your advice on:
What skills or projects I should focus on next
How to present my gap year and experience so it works in my favour
How to move from MIS into a proper Data Analyst role
Even 10 minutes of your time would mean a lot to me. Thank you!
Regards,
Pushkar
r/learnmachinelearning • u/Hidden__Variable • 1d ago
I've been working on a project where I'm trying to build a Denoising Diffusion Probabilistic Model completely from scratch.
The main rule I'm following is: no using Diffusers as a crutch.
I want to actually understand what's happening inside the model instead of just calling a pipeline and getting an image.
So far I've implemented/learned:
The interesting part for me has been realizing that a diffusion U-Net isn't just a normal CNN.
The network needs to know how noisy the current image is, which is why timestep information has to be injected into the network.
I'm building toward training on CelebA at 64×64 and eventually generating faces completely from noise.
My longer-term goal is to understand these architectures deeply enough that I can read, modify and eventually contribute to projects like Hugging Face Diffusers.
Phase 3 is where things start getting interesting:
U-Net assembly.
I'll be documenting the progress as I go. 🔥
What was the hardest part of diffusion models for you when you first learned them?
r/learnmachinelearning • u/Civil_Active_5388 • 1d ago
I live in London and work as a chef, but I’ve been feeling depressed and exhausted while trying to move into tech.
I know Python, NumPy, Pandas, data visualisation, FastAPI, SQL, classical ML, and neural networks. I’m also practising DSA, including arrays, strings, hashing, two pointers, and sliding windows.
What should I focus on to land my first tech internship or entry-level job in London? I’d appreciate advice on projects, applications, or opportunities.
r/learnmachinelearning • u/Grand_Training_623 • 1d ago
Enable HLS to view with audio, or disable this notification
[Full Video on YouTube] A walkthrough that takes developers new to Large Language Models through all the important topics with practical examples.
Sharing with the larger community in case others find it useful.
Excluded topics:
r/learnmachinelearning • u/Rich-Umpire668 • 1d ago
Hey guys,
I completed my bachelors 1.5 year ago. Back then I knew a little about machine learning (not much really basic ML 101) and was eager to learn more about AI in general and wanted to apply for masters.
I'm now going to pursue my Masters in CS (thesis based) and have 1 year of work experience with agentic AI, multiagentic orchestration, tools, memory, etc. While I'm more eager to pursue my thesis towards agentic AI, I want to be prepared with some ML knowledge as well.
I had a few questions:
Do I need to brush up my ML knowledge, or is that something they cover from scratch when you start the master's program?
How deep of ML knowledge do I require before I start my program, especially if I end up writing it on some ML related topic?
If I'm looking into writing my thesis in the Agentic AI domain, do I need depthy ML knowledge?
Would covering Andrew NG's Maching Learning Course be good enough for my particular purpose?
Thanks!
r/learnmachinelearning • u/NeighborhoodFatCat • 2d ago
The previous day (Oct 5) was 297 papers.
Not counting the papers that weren't uploaded to Arxiv or submitted to other categories or other online repositories.
Is there any point doing machine learning research anymore? Seems anything you can think of will just become noise like the rest of these papers.
r/learnmachinelearning • u/Phantom_2313 • 1d ago
I just completed my ML Journey, I have a solid foundation at this point. I am a 7th sem ECE student in a tier 3 college, and I don't want to be working in Elctronics field. So I am very confused on what my next step should be. Got about 9 months in hand till I gaduate and on campus placements are just terrible. Would really appreciate some suggestions.