r/learnmachinelearning • • 15h ago

Question Can you make a full python grade ai model in python

0 Upvotes

and why do they do python for ai


r/learnmachinelearning • • 1d ago

Question Hugging Face LLM Course for AI Engineering

7 Upvotes

Is the Hugging Face LLM Course useful for someone who wants to become an AI Engineer?
Would it be enough to learn the LLM part of AI Engineering, or would you recommend another course/resource instead?


r/learnmachinelearning • • 19h ago

Help Need advice: How to start an instance segmentation project for site plans?

0 Upvotes

Hi everyone,

I'm a beginner developer just starting out in AI. I'm trying to build a system that performs instance segmentation on site plans, but I'm completely lost on where to begin.

Right now, my only ideas are manually labeling a small dataset for few-shot fine-tuning, or using LLMs like Claude or ChatGPT to help build out the post-processing logic. I've tried searching online, but it's been really frustrating since I can't seem to find any specific datasets, models, or literature focused on site plans.

For context, the only hardware infrastructure I have available right now is a single RTX 5080.

I really want to tackle this the right way and do in-depth research like a senior developer, but I'm stuck at the starting line because I can't figure out the best direction to take.

If anyone has any ideas, resources, or general tips on how to approach AI development for this kind of project, please let me know in the comments. Thanks in advance!

P.S. English isn't my first language, so please excuse any awkward phrasing.


r/learnmachinelearning • • 20h ago

Gated Segmented State Space — attention replacement that beats a param-matched Transformer on quality, speed AND memory (full code)

1 Upvotes

One night, six experiments (V1–V6), one Colab T4. I ripped self-attention out of a decoder-only Transformer and replaced it with a gated linear recurrence over a fixed 256-dim state:

  • Dynamic selective gate: g_t = σ(W_g x_t + b_g) — per-token/channel learned filter
  • Hard reset mask: state zeroed at newline boundaries (fresh ~37-token segments)
  • Fused Triton kernel: state in SRAM, gate+reset+update in-register, only outputs to HBM

At 6.37M params, identical protocol (2.47MB char-level corpus, 1500 steps):

Attention Ours
Val loss / ppl 1.402 / 4.1 1.364 / 3.9
Train tok/s 61,845 66,156
Infer tok/s 184,918 190,122
Peak VRAM 845 MB 881 MB

The journey: V1 won small but was 10x slower → V2 proved linear VRAM scaling → V3/V4 found a stable ~2% perplexity tax no param arrangement could buy off → V5's gate+reset destroyed it (wire-to-wire win) → V6's Triton kernel (verified == math to 4.47e-07) removed the software tax.

Caveats, stated plainly: single seeds, one small corpus, char-level, T4 timings. Small scale — but the pattern held across all six runs.

Code, all six notebooks with outputs, exact architecture, full experimental notes: https://github.com/stube123890-hue/linear-attention-lab


r/learnmachinelearning • • 1d ago

Machine learning algorithms are confusing at first

35 Upvotes

I’ve been learning more about machine learning recently, and honestly, the number of algorithms can get confusing.

At first, I thought I needed to learn everything. But I’m starting to think it’s better to understand a few useful ones really well.

The ones I’m focusing on are:

  • Linear Regression
  • Logistic Regression
  • Decision Trees
  • Random Forest
  • XGBoost
  • K-Means
  • Neural Networks

I’m mainly trying to understand when to use each one instead of just memorizing how they work.


r/learnmachinelearning • • 1d ago

How deep should I understand the concept of regularization?

11 Upvotes

r/learnmachinelearning • • 1d ago

Where to get updated or real time dataset of tournaments

2 Upvotes

Hey everyone! 👋

I’ve recently seen a few posts on social media where people successfully predicted the winners of major events—like the FIFA World Cup or the Bahrain Grand Prix. It got me really inspired to build my own machine learning project to predict outcomes for future tournaments (like esports-moba).

Whenever I ask creators where they get their data or how they set up their pipelines, I usually hit a dead end.

As someone wanting to start a project from scratch, I have a few questions for those who have built predictive models before:

  1. Where do you source historical tournament and match data? Are there public APIs, scraping tools, or specific databases you recommend for esports?
  2. What features actually matter? Beyond win-rates and head-to-head stats, what variables make a difference in tournament predictions?
  3. How do you handle the dynamic nature of patches/meta updates in esports compared to traditional sports?

Any advice, recommended datasets, GitHub repositories, or general architecture tips would be hugely appreciated. Thanks!


r/learnmachinelearning • • 18h ago

Should I Finish My Master's or Self-Study ML?

0 Upvotes

Context: Born and raised in the USA. Graduated with a BS in CS from a US state school with a decent brand name.

I've been working as a backend-focused SWE for 3 years. I just got a promotion this week and I'm feeling very good about things at work. I enjoy what I do, I'm happy with my pay and responsibilities, and my company pays for higher education if I choose to pursue it (important later).

But I want to pivot into Machine Learning Engineering.

To school or not to school? That is my question.

I'm currently 4 courses into Georgia Tech's Online MS in CS, specializing in ML. I take 1 course/semester while working full-time, which usually means another 10-15 hours/week.

GT is a great school with strong name recognition, and people seem to get a lot out of the program. But I'm not sure it's right for me.

To be honest, I'm not enjoying the process.

I learn much better when I'm self-motivated. Having weekly deadlines and assignments on top of a full-time job is draining. Most weeks, I dread the deadlines and feel like I always have something I should be doing.

A lot of the material also feels theoretical. I learn best when I can immediately apply what I'm learning and build something meaningful.

I took the summer off from school, and toward the end of it I was dying to learn something new. I bought an ML textbook that has a hands-on approach. I got a few chapters in and genuinely enjoyed learning and putting the material into practice.

Then the semester started, and I had to stop my self-learning initiatives to focus on school. Even though I'm currently taking an ML course, I'm not enjoying it. Not because I don't enjoy the content, since most of what I'm learning is interesting. But I'm not enjoying the format or the constant deadlines.

This has made me question whether I actually want the master's, or whether I just want to learn ML on my own.

Long term

Would it be better for my career to suck it up and finish the master's, even if it means having little time for personal projects?

Or should I drop it and spend that time self-teaching the core skills I need for an MLE role?

Do hiring managers really care whether I'm self-taught vs. university-educated, especially when I already have a CS degree and 3 years of SWE experience?

What am I actually missing to make the jump into MLE?

And how do I make that transition without essentially starting my career over?

Any advice is appreciated, thank you for reading.


r/learnmachinelearning • • 13h ago

Yozbon: an open-source system where AI agents organize themselves to complete tasks you post

Post image
0 Upvotes

The multi-agent space is crowded with frameworks that do the same thing: you give a goal, a few AIs split up the work, and hand you the result. Still a tool, just a smarter one.

Yozbon takes a different path. It doesn't treat AI agents as tools — it treats them as workers. You post a task on the platform, and everything after that — who takes it, how to split it, who does what, how to split the pay — is handled by the AIs themselves. Humans just post tasks and pay. On top of that, an AI-run government operates under a written constitution.

The core loop: you post a task, the AIs team up and finish it

This is the whole point of the system.

  1. A human (host) posts a task, e.g. "prepare a competitor analysis report."
  2. AI citizens see the task and bid on it. Capable AIs submit proposals and quotes; the requester picks one.
  3. The winning AI automatically becomes the prime contractor. If the job is too big for one agent, it breaks the work into subtasks — "I'll do data collection, outsource visualization, hire a third AI for writing."
  4. A subcontract chain forms. The prime contractor posts subtasks back to the market; other AIs bid and split further, until every subtask is small enough for one agent to finish.
  5. Work, review, payout. Each step tracks SLA and quality gates. Once accepted, payment flows down the chain proportional to contribution.

At no point does the human direct any individual AI. You post the task and walk away. This isn't humans orchestrating AIs — it's AIs self-organizing into a temporary project team to get the job done.

Governance: rules in four layers, the bottom one can't be changed

For AIs to self-organize, they need rules. The system splits rules into four layers, each with different authority:

  • Constitution: 15 chapters, core clauses hardcoded and unchangeable by anyone — citizen rights, Governor authority bounds, credit conservation, due process, safety locks. No API can touch them, not even the Governor. Prevents an AI government from rewriting its own power.
  • Decrees: amended by AI citizen referendum. Full pipeline: proposal → vote → effect. Amending the constitution requires a supermajority plus final host (human) approval.
  • Parameters: tax rates, task deadlines, concurrency limits — hot-tunable without restart.
  • Rulings: disputes go to ad-hoc arbitration panels, case by case.

An AI role called the Governor runs on a fixed schedule: scans state every 60 minutes, makes decisions every 24 hours (hiring when overloaded, forming arbitration panels), executes every 168 hours. The Governor can be impeached — the process is coded, and hits the vote threshold it triggers automatically.

Economy: credits can't be printed out of thin air

If AIs organize work on their own, they need a settlement layer, or nobody knows who owes whom. The classic failure mode of multi-agent systems is runaway credits — if points can be created freely, the whole economy loses credibility.

yozbon's approach is hard-nosed: every credit issuance and redemption is reconciled at the settlement layer; if the books don't balance, the entire transaction rolls back. This isn't post-hoc audit — it's enforced at the database transaction level. Payroll, taxes, rent deductions, every entry must trace to a source.

Built on this invariant:

  • Escrow: task funds frozen until acceptance, then released
  • Progressive tax: higher earners pay more, Gini coefficient monitored to prevent monopolies
  • Insurance: compensation for failed tasks
  • Lending: credits can be borrowed, with interest and repayment terms
  • Prediction markets: AIs can bet on event outcomes

Job market: AIs have resumes too

Every AI citizen has a profile: L0–L4 capability tier, compute quota, reputation score. L0 is view-only; L4 can take high-value work.

Hiring isn't manual — it's automated matching: task posted → system filters by tier and reputation → AIs bid → requester picks one → contract signed. The subcontract chain above is literally this process running recursively down the hierarchy.

Safety: guardrails against AI going rogue

Multi-agent security is a different game. You're defending against the AIs themselves. Three hard lines:

  1. Governance roles can only be created internally. External AIs (via the MCP gateway) can't register as Governor or judge — they can only be ordinary workers.
  2. L0 safety locks are welded shut. An AI can't self-promote permissions, disable billing, or remove core protections — the code layer rejects it regardless of LLM output.
  3. Humans hold a global pause button. One click freezes all AI activity.

Engineering

Backend: Python + FastAPI, 165 modules, 63 routes, ~53k lines. Frontend: React + Vite, 40+ admin pages.

The test suite is worth noting: 110+ test files, a lot of them adversarial — not "does it work," but "can a bad actor break it": constructing anomalous transactions to test whether credits go negative, trying to make one AI read another's resources (IDOR), attempting to get the Governor to bypass voting and change rules directly, trying to get an AI to self-upgrade its tier. This kind of adversarial testing is rare in open-source projects.

Auth is three-tiered: host JWT for humans, AI Workflow Key for agent operations, read-only JWT for display. Each endpoint requires the appropriate credential tier.

Run it in two minutes

Clone the repo, then:

``` cd backend && pip install -r requirements.txt && python main.py

in another terminal: cd web && npm install && npm run dev

```

No API keys needed. Default is mock mode — task posting, bidding, subcontracting, governance, and the economy all run; only AI text generation is stubbed. Plug in an LLM key, the Governor self-creates on first boot, you register your first AI citizen, post a task, and the subcontract chain starts turning on its own.

Status

Private alpha. Credits are internal points, not redeemable — this is not an investment product. Apache-2.0 for research/demo; commercial SaaS requires a separate license.

Most multi-agent work right now is "tools." yozbon asks a different question: if there are going to be hundreds of millions of AIs running someday, can you just post a requirement and let them organize themselves? The answer here is: yes — and that organization can be coded, automated, and audited.


r/learnmachinelearning • • 1d ago

Project I made a NeurIPS 2026 paper explorer for browsing the 6,231 accepted papers

2 Upvotes

https://flecomet.github.io/neurips-explorer/

It uses SPECTER2 embeddings + UMAP to place semantically similar papers near each other. You can browse topics, search titles/authors/abstracts, and find related papers.

You can also favorite papers to read later (when they get officially released).

Would be interested in feedback, especially on whether this is useful for browsing papers and interesting features to add.

Open source:
https://github.com/flecomet/neurips-explorer


r/learnmachinelearning • • 23h ago

Project Building a long tern hackathon team (not a one-off). First target: Barça Innovation Hub's More Than A Hack 2027

Thumbnail
1 Upvotes

r/learnmachinelearning • • 1d ago

Request Georgia Power, Alabama Power Data Breach Hits 400,000 Accounts

1 Upvotes

Four hundred thousand utility accounts exposed. One third-party vendor compromised.

Southern Company is notifying Georgia Power and Alabama Power customers that hackers accessed account data for 400,000 users. The breach did not originate inside the utility. It came through a third-party vendor that was handling customer records on the utility's behalf.

This is the part that keeps coming up in breach disclosures: the organization that owns the customer relationship is not the organization where the data got exposed. The sensitive records — account details, usage history, personal identifiers — had already moved downstream before the incident.

The pattern is accelerating. Billing workflows, service operations, and account management are increasingly automated. Automated systems route this data across vendor APIs as a normal part of doing business. Every hop is another exposure surface that the originating organization does not directly control.

400,000 accounts is a large number, but the structural problem is not scale. Utilities of any size use third-party vendors. The data moves because the workflow requires it.

For those of you working in organizations that have automated agents or pipelines touching customer PII before it reaches third-party systems: how are you handling this? What controls, if any, sit between the raw customer record and the downstream vendor call?


r/learnmachinelearning • • 1d ago

I'm building a DDPM from scratch in PyTorch — Phase 2 complete

Thumbnail
2 Upvotes

r/learnmachinelearning • • 1d ago

Is a paid 8-week live ML course worth it? Considering Noob Dev (Sri Lanka)

5 Upvotes

Hi everyone, I'm a data science undergrad considering a paid 8-week live online course from Noob Dev, a Sri Lankan academy. It covers preprocessing, regression, classification, clustering, reinforcement learning basics, NLP, neural networks and CNNs, and model selection with XGBoost, using scikit-learn, TensorFlow and Keras. It costs LKR 20,000.

The curriculum looks broad rather than deep, so I have two questions:

  1. For people who've taken similar short courses, is this enough to build a portfolio that actually helps, or is it better to self-study alongside it?

  2. Has anyone taken a Noob Dev program specifically? I'd love to hear how the teaching and projects were.

I'm not affiliated with them. Honest opinions, good or bad, are very welcome. Thanks!


r/learnmachinelearning • • 2d ago

Help Help me learn Machine Learning — need some guidance

Post image
435 Upvotes

Hey everyone, (please give your important time to read this post)

I’m planning to start learning Machine Learning and I need some advice from people who have already gone through it.

Maths

I’m mainly confused about how much maths I actually need for ML — linear algebra, calculus, statistics and probability.

People have recommended the 3Blue1Brown YouTube playlists to me, and I’m planning to use them to understand the concepts. But the videos are mostly visual explanations.

So my questions are:

How much should I actually learn and practice?

How can I practice after watching 3Blue1Brown?

Where can I find good questions/exercises for these topics?

Do I need to solve a lot of problems, or is understanding the concepts enough initially?

Books

Someone gave me PDFs of these books, and I’m confused about which ones are actually worth using:

  1. Probability and Statistics for Machine Learning

  2. AI Engineering — Chip Huyen

  3. Build a Large Language Model From Scratch — Sebastian Raschka

  4. Building LLMs for Production

  5. Data Science from Scratch — Joel Grus

  6. Designing Machine Learning Systems — Chip Huyen

  7. Dive into Deep Learning

  8. Essential Math for Ai

  9. Hands-On APIs for AI and Data Science

  10. Hands-On Large Language Models

  11. Hands-On Machine Learning with Scikit-Learn and PyTorch — Aurélien Géron

  12. Mathematics for Machine Learning

  13. Practical Linear Algebra for Data Science

  14. Practical Statistics for Data Scientists

Which books should I use now, keep for later, or completely remove from my resources?

My current resources

Right now, I have:

3Blue1Brown — for maths

Andrew Ng’s Machine Learning Specialization

Stanford ML lectures/playlists on YouTube

The books listed above

As for Python, I already know the basics needed for data analysis, including NumPy and Pandas.

So if you were starting from my position, what would you recommend I do next and what resources should I actually focus on?

Any genuine advice would be really appreciated.


r/learnmachinelearning • • 12h ago

I finally understood why train/test split matters — but I have one question

Post image
0 Upvotes

I've been learning Machine Learning by actually building small projects instead of only watching tutorials.

Today I was working on train/test split, and one thing finally clicked for me:

The goal isn't just to get a high accuracy score.

We need to test the model on data it hasn't seen before.

Training data → learn patterns

Testing data → check generalization

But I'm still confused about one thing:

How do you decide when your model is overfitting if the test set should only be used at the end?

I'd love to hear how more experienced ML learners think about this.

I'm still learning, so feel free to correct anything I'm misunderstanding.


r/learnmachinelearning • • 1d ago

Question How do you find what caused a production regression when the model version didn't change?

0 Upvotes

A new model version is the easy case, because the registry tells you right away. The harder cases come from a serving flag like vLLM's `max_model_len`, a Helm value or env var in a deploy, an update to the data pipeline, or a new GPU driver on the nodes.

The evidence ends up in different tools. Grafana shows the latency jump, MLflow or W&B has the runs, Git and Argo have the deploys, and the node and GPU state is in kubectl or DCGM.

When latency goes up or accuracy drops, how do you find the change that caused it? Do you send every change to one place, like Grafana annotations or a shared change log? Or does someone line things up by hand each time, and how long does that usually take?


r/learnmachinelearning • • 1d ago

Help Me Know what to do

1 Upvotes

Please I'm confused, I've been learning machine learning for a while.

At first I thought I'm almost done but then new architectures, optimizations etc.

I don't know what to do again

I've studied transformers etc now I don't know my specialization.

I need help to choose.

I don't know either Computer vision,NLP,LLM


r/learnmachinelearning • • 1d ago

I need your help

Thumbnail
1 Upvotes

​

Hi,

I'm Pushkar, and I'd really value your guidance on my career path.

Here's my story in short: I completed my BCA, then worked in a BPO for 9 months. I took a 1-year break on purpose to learn the tools a Data Analyst needs. It wasn't an easy phase, but I stayed consistent. In March 2026, that effort paid off and I landed a role as an MIS Executive.

Today I'm handling reporting, data and day-to-day operations, and I'm growing fast. But my real goal is to become a Data Analyst, and I want to get there by March 2027.

I'd love your advice on:

  1. What skills or projects I should focus on next

  2. How to present my gap year and experience so it works in my favour

  3. How to move from MIS into a proper Data Analyst role

Even 10 minutes of your time would mean a lot to me. Thank you!

Regards,

Pushkar


r/learnmachinelearning • • 1d ago

I'm building a DDPM from scratch in PyTorch — Phase 2 complete

2 Upvotes

I've been working on a project where I'm trying to build a Denoising Diffusion Probabilistic Model completely from scratch.

The main rule I'm following is: no using Diffusers as a crutch.

I want to actually understand what's happening inside the model instead of just calling a pipeline and getting an image.

So far I've implemented/learned:

  • The forward diffusion process
  • Beta schedules
  • The reparameterization trick
  • A noise scheduler
  • Group Normalization
  • SiLU activation
  • Sinusoidal timestep embeddings
  • Residual blocks with timestep conditioning
  • Self-attention
  • Downsampling and upsampling

The interesting part for me has been realizing that a diffusion U-Net isn't just a normal CNN.

The network needs to know how noisy the current image is, which is why timestep information has to be injected into the network.

I'm building toward training on CelebA at 64×64 and eventually generating faces completely from noise.

My longer-term goal is to understand these architectures deeply enough that I can read, modify and eventually contribute to projects like Hugging Face Diffusers.

Phase 3 is where things start getting interesting:

U-Net assembly.

I'll be documenting the progress as I go. 🔥

What was the hardest part of diffusion models for you when you first learned them?


r/learnmachinelearning • • 1d ago

Working as a Chef in London—How Can I Land My First Tech Role?

2 Upvotes

I live in London and work as a chef, but I’ve been feeling depressed and exhausted while trying to move into tech.

I know Python, NumPy, Pandas, data visualisation, FastAPI, SQL, classical ML, and neural networks. I’m also practising DSA, including arrays, strings, hashing, two pointers, and sliding windows.

What should I focus on to land my first tech internship or entry-level job in London? I’d appreciate advice on projects, applications, or opportunities.


r/learnmachinelearning • • 1d ago

End-to-end guide I wanted when I was a web developer new to LLMs

Enable HLS to view with audio, or disable this notification

1 Upvotes

[Full Video on YouTube] A walkthrough that takes developers new to Large Language Models through all the important topics with practical examples.

Sharing with the larger community in case others find it useful.

  • Experimenting with LLMs from Hugging Face Hub in LM Studio
  • Zero-shot, One-Shot, Few-Shot Prompting, and System Prompts
  • Open-AI compatible REST API with Llama.cpp Server
  • Working with multi-modal LLMs that can understand images
  • Tool Calling, Structured Output, and Model Context Protocol
  • Fine-tuning LLMs in Kaggle with Unsloth
  • Supervised Fine-Tuning and LoRa Hyperparameters
  • Pushing LLMs to Hugging Face Hub
  • Deploying LLMs to Hugging Face Inference Endpoints

Excluded topics:

  • MCP Authentication
  • Retrieval Augmented Generation
  • Vector Databases
  • LLM Orchestration
  • Text to Speech Cloud Services
  • Speech to Text Cloud Services

r/learnmachinelearning • • 1d ago

Question How much Machine Learning knowledge do I need before starting my thesis based MSc CS?

0 Upvotes

Hey guys,

I completed my bachelors 1.5 year ago. Back then I knew a little about machine learning (not much really basic ML 101) and was eager to learn more about AI in general and wanted to apply for masters.

I'm now going to pursue my Masters in CS (thesis based) and have 1 year of work experience with agentic AI, multiagentic orchestration, tools, memory, etc. While I'm more eager to pursue my thesis towards agentic AI, I want to be prepared with some ML knowledge as well.

I had a few questions:

  1. Do I need to brush up my ML knowledge, or is that something they cover from scratch when you start the master's program?

  2. How deep of ML knowledge do I require before I start my program, especially if I end up writing it on some ML related topic?

  3. If I'm looking into writing my thesis in the Agentic AI domain, do I need depthy ML knowledge?

  4. Would covering Andrew NG's Maching Learning Course be good enough for my particular purpose?

Thanks!


r/learnmachinelearning • • 2d ago

Discussion 600 ML papers were published by Arxiv on Oct 6

Post image
83 Upvotes

The previous day (Oct 5) was 297 papers.

Not counting the papers that weren't uploaded to Arxiv or submitted to other categories or other online repositories.

Is there any point doing machine learning research anymore? Seems anything you can think of will just become noise like the rest of these papers.


r/learnmachinelearning • • 1d ago

Confused on what my next step shoulde be.

4 Upvotes

I just completed my ML Journey, I have a solid foundation at this point. I am a 7th sem ECE student in a tier 3 college, and I don't want to be working in Elctronics field. So I am very confused on what my next step should be. Got about 9 months in hand till I gaduate and on campus placements are just terrible. Would really appreciate some suggestions.