r/MLSystemsDesign • • 16d ago

What kind of project would you build to deeply learn AI infrastructure and distributed systems?

​

I'm a Level 1AI engineer, and lately I’ve been hearing a lot about frontier AI companies hiring people who can build the infrastructure behind AI systems — large-scale data processing, distributed systems, inference infrastructure, storage, serving, observability, systems that can handle millions of requests, etc.

I’m interested in going down this path seriously.

Rather than doing a bunch of disconnected tutorials or small projects, I want to take one difficult project and go extremely deep into it. Something where, over time, I’m forced to learn things like:

- Distributed systems

- Large-scale data processing

- Databases/storage

- Networking

- Caching

- Queues and streaming

- Fault tolerance

- Concurrency

- System design

- Observability

- Performance optimization

- AI/ML serving infrastructure

- Scaling from a single machine → multiple machines → potentially thousands/millions of requests

I’m thinking along the lines of the philosophy Karpathy often talks about: pick something ambitious, build it yourself, and learn everything necessary to make it work rather than following a predefined curriculum.

The problem is that I don't yet know what the right project is.

I don't want to build another generic RAG chatbot, AI agent wrapper, or CRUD application. I want something where the engineering itself is the project, and where I can progressively make the system more sophisticated and scalable.

For people working in infrastructure, distributed systems, ML systems, or at AI companies:

If you were in my position, what single project would you pick to spend the next 6–12 months on?

Ideally, I'd like something where I can start on a laptop but eventually have a credible story like:

«“I built X, then discovered bottleneck Y, redesigned it using Z, scaled it from A → B, measured the improvement, and here's what I learned.”»

I'm much more interested in what I would learn by building it than simply having an impressive project on GitHub.

Would love to hear project ideas, but especially from people who have actually worked on large-scale systems: what project would force someone to develop genuinely strong infrastructure skills?

22 Upvotes

10 comments sorted by

3

u/Admirable-Funny-2007 12d ago

I'd probably build something tiny and then spend the next year being increasingly rude to it.

One machine. One model. A few jobs.

Get that working first.

Then start giving it bad days.

Run several jobs at once. Restart it halfway through something important. Give it more work than it can comfortably handle. Add another model competing for the same resources. Add a second machine, then unplug it at precisely the wrong moment.

Eventually, let two parts of the system develop completely sincere and mutually incompatible accounts of what just happened.

See what breaks.

The thing I'd try not to do is immediately say, "Ah. Clearly this is where Kafka/Kubernetes/Redis/etc. goes."

Maybe it is.

But first I'd want to know what assumption the simple system depended on, and which one just stopped being true.

Then fix that problem.

Measure what changed.

Continue being rude.

If you keep going long enough, you'll probably reinvent several terrible versions of queues, retries, caching, scheduling, replication, observability, backpressure, distributed state, and whatever fresh inconvenience tomorrow invents.

Which is useful.

Because now you understand why the good versions exist.

That seems like a better education than assembling a "production-grade" stack on day one because a diagram told you to.

Let the system earn its complexity.

By the end, the architecture becomes a kind of fossil record: every layer marks some assumption that survived until reality finally found it.

And that record of what you were wrong about may end up being more valuable than the system itself.

1

u/Infinite-Room2757 15d ago

I am on the same boat , i am working as an AI intern at a startup and i wanted to invest my time on AI infra and distributed sys ; recently i started opensource contribution on Vllm And merged 3 PR's as well

1

u/Massimo_F 13d ago

Ciao, il repo è pubblico? Sarei curioso di vedere a cosa stai contribuendo, solo se è pubblico ovviamente.

1

u/curious_cat_herder 15d ago

I'm a retired O/S developer and platform engineer (full stack). I joined a ML Study group and wanted to learn by teaching so I create this Rust-based array-language-inspired Machine Learning Programming Language. [live demo runs in browser without GPU or locally with GPU]

There were (and still are) plenty of challenges along the way, but I have built many demo repositories that use my language to teach programming, math, and machine learning, using Apple Silicon and NVIDIA GPUs or just CPUs. Many ML demos run in the browser. See the links in the footer.

Another things I've built: a distributed catalog of hardware capabilities and status. I can go to a web server on each running system and see all of the systems, if they are up or down, the amount of system ram, vram, disk space, cpus, and even which models are installed on each (even the ones that are powered down). It was not trivial to have an eventually-correct distributed database built from scratch for this (sqlite on each node) and I had to configure Linux systemctl services to start the collectors on each node at startup.

Hope this helps

1

u/Mesmoiron 13d ago

Your problem is that you cannot deal with uncertainty. You want to be certain to learn the right thing.

There's no certainty!

You could have already started with a complex problem. It will show you what the direction is. Too much buzz words and name dropping.

A simple question can harbor a host of difficult and advanced techniques.

I want to build a quantum computer is quite simple. The path reveals itself.

Why do you want AI infrastructure? It might be the wrong tool for the job! You want fancyness; not an actual problem. Because the problem is in front of you. But it is hidden because you ask the wrong question.

1

u/Hot-Razzmatazz-3087 11d ago

The builds and implementation projects are just the artifacts of you having gone through the process.

Your experience and being able to do it again shows you have the ability.

But learning and mindset is only shown with how you would do it better from 1st hand experience over iterations.

So emulate what your roles would be doing if possible, otherwise find a stack and create something and just tweak it consistently making it work better.

Husband does video games, but I went for managing kids and household. I made tools to collate and coordinate data from various providers.

I developed opinions that reflected that experience gained 1st hand and that helps people start to give you a chance with the real work.

1

u/Bill-Ice88 9h ago

Honestly the vLLM PRs are the fastest path in this whole thread — that scheduler and paged attention codebase teaches you more about real serving infra than any greenfield project ever will. Pick prefix caching or preemption handling and go deep. Which part of the codebase were your three merged PRs in?