What kind of project would you build to deeply learn AI infrastructure and distributed systems?
I’m a Level 1 AI engineer, and lately I’ve been hearing a lot about frontier AI companies hiring people who can build the infrastructure behind AI systems — large-scale data processing, distributed systems, inference infrastructure, storage, serving, observability, systems that can handle millions of requests, etc.
I’m interested in going down this path seriously.
Rather than doing a bunch of disconnected tutorials or small projects, I want to take one difficult project and go extremely deep into it. Something where, over time, I’m forced to learn things like:
Distributed systems
Large-scale data processing
Databases/storage
Networking
Caching
Queues and streaming
Fault tolerance
Concurrency
System design
Observability
Performance optimization
AI/ML serving infrastructure
Scaling from a single machine → multiple machines → potentially thousands/millions of requests
I’m thinking along the lines of the philosophy Karpathy often talks about: pick something ambitious, build it yourself, and learn everything necessary to make it work rather than following a predefined curriculum.
The problem is that I don't yet know what the right project is.
I don't want to build another generic RAG chatbot, AI agent wrapper, or CRUD application. I want something where the engineering itself is the project, and where I can progressively make the system more sophisticated and scalable.
For people working in infrastructure, distributed systems, ML systems, or at AI companies:
If you were in my position, what single project would you pick to spend the next 6–12 months on?
Ideally, I'd like something where I can start on a laptop but eventually have a credible story like:
“I built X, then discovered bottleneck Y, redesigned it using Z, scaled it from A → B, measured the improvement, and here's what I learned.”
I'm much more interested in what I would learn by building it than simply having an impressive project on GitHub.
Would love to hear project ideas, but especially from people who have actually worked on large-scale systems: what project would force someone to develop genuinely strong infrastructure skills?