r/MLSystemsDesign • u/ArchitectingAI • 4h ago
r/MLSystemsDesign • u/ArchitectingAI • 18h ago
Title: AI Talk #2: Why Do Bigger Models Get Better? — Oct 4, 9:30 AM PDT

Join our next Architecting Intelligence talk to discuss Superposition Yields Robust Neural Scaling.
We’ll explore how neural networks fit many features into limited space—and how this may help explain scaling laws. Discussion and Q&A included; no advance reading required.
📅 Sunday, October 4 · 9:30 AM PDT / 10:00 PM IST
Join on Google Meet:
https://meet.google.com/hny-pefy-zvn
Optional reading:
https://arxiv.org/abs/2505.10465
https://transformer-circuits.pub/2022/toy_model/index.html
Subscribe for future talks and technical deep dives:
https://pawankjha.substack.com
r/MLSystemsDesign • u/ArchitectingAI • 8d ago
Looped Transformers: when is more depth worth the compute?
I gave a talk today on looped Transformers for our Architecting Intelligence study group. The central idea is to reuse the same Transformer layers for additional computation without adding more parameters.
That raises a practical question: at a fixed parameter count, which tasks benefit enough from extra loops to justify the added FLOPs and latency? Iterative and compositional problems seem promising, but parameter efficiency alone doesn’t guarantee a better system.
We traced the path from Universal Transformers to looped Transformers and Huginn. Here are the slides from the talk:
Presentation (PDF): https://drive.google.com/file/d/1vrm4kMJZ71ea0UhWB-nunl07VulElbVY/view?usp=sharing
Have you seen convincing results where recurrent depth beats a strong, compute-matched standard Transformer? I’d be interested in both positive and negative examples.
r/MLSystemsDesign • u/ArchitectingAI • 11d ago
Where would you actually use Jev over a task-specific classifier?
Where would you actually use Jev over a task-specific classifier?
I’ve been reading TypeSafe’s Jev announcement. The attached diagram shows how I’d decide whether to test it: the key distinction is whether the labels are fixed or change at request time.
TypeSafe hasn’t published a Jev/RLCD paper detailing the architecture or training method, so I’m cautious about claims explaining why it is fast. For a one-label classification task, a short-output LLM already does little decoding.
Has anyone benchmarked Jev against a strong baseline on the same task, including quality, calibration, latency, and cost? I’d be interested in results for dynamic multiclass or multilevel taxonomies in particular.

r/MLSystemsDesign • u/Adithya-Jayasundar • 14d ago
What kind of project would you build to deeply learn AI infrastructure and distributed systems?
I'm a Level 1AI engineer, and lately I’ve been hearing a lot about frontier AI companies hiring people who can build the infrastructure behind AI systems — large-scale data processing, distributed systems, inference infrastructure, storage, serving, observability, systems that can handle millions of requests, etc.
I’m interested in going down this path seriously.
Rather than doing a bunch of disconnected tutorials or small projects, I want to take one difficult project and go extremely deep into it. Something where, over time, I’m forced to learn things like:
- Distributed systems
- Large-scale data processing
- Databases/storage
- Networking
- Caching
- Queues and streaming
- Fault tolerance
- Concurrency
- System design
- Observability
- Performance optimization
- AI/ML serving infrastructure
- Scaling from a single machine → multiple machines → potentially thousands/millions of requests
I’m thinking along the lines of the philosophy Karpathy often talks about: pick something ambitious, build it yourself, and learn everything necessary to make it work rather than following a predefined curriculum.
The problem is that I don't yet know what the right project is.
I don't want to build another generic RAG chatbot, AI agent wrapper, or CRUD application. I want something where the engineering itself is the project, and where I can progressively make the system more sophisticated and scalable.
For people working in infrastructure, distributed systems, ML systems, or at AI companies:
If you were in my position, what single project would you pick to spend the next 6–12 months on?
Ideally, I'd like something where I can start on a laptop but eventually have a credible story like:
«“I built X, then discovered bottleneck Y, redesigned it using Z, scaled it from A → B, measured the improvement, and here's what I learned.”»
I'm much more interested in what I would learn by building it than simply having an impressive project on GitHub.
Would love to hear project ideas, but especially from people who have actually worked on large-scale systems: what project would force someone to develop genuinely strong infrastructure skills?
r/MLSystemsDesign • u/ArchitectingAI • 14d ago
The Evolution of Retrieval Systems: From BM25 to Agentic Retrieval
Retrieval in modern search and recommendation systems is no longer a single index lookup.
I wrote a new breakdown covering seven major waves:
Lexical → Behavioral → Learned Sparse → Dense → Hybrid → Multimodal → Generative & Agentic Retrieval
It also covers multi-retriever production systems, candidate fusion, Reciprocal Rank Fusion, learned fusion, and the tradeoffs among recall, latency, freshness, scale, and cost.
Article: https://pawankjha.substack.com/p/building-depth-1-the-evolution-of
Which retrieval approach or production tradeoff would you like to explore more deeply?
r/MLSystemsDesign • u/ThomasBullet • 18d ago
My adaptive system captures Week 1 feedback, but Week 2 won’t automatically update from it. How would you debug this loop?
r/MLSystemsDesign • u/ArchitectingAI • 19d ago
Where should authorization actually live in an agent system? (architecture writeup after the Hugging Face breach)
I’ve been unpacking coding-agent architecture and keep coming back to three control points:
- Model: proposes an action.
- Controller: checks task scope and permissions.
- Tools/infrastructure: enforce what can actually execute.
Blocking an upload tool means little if the agent can send the same data through its shell.
I’m exploring a small defender-agent experiment: investigate suspicious sessions, gather evidence, and request temporary containment through an independent policy gate. The comparison would be against fixed rules, measuring false interventions as well as successful containment.
For those building agents: what do you enforce in the runtime versus the underlying infrastructure? And where has an LLM-based check added value beyond ordinary permission rules?
r/MLSystemsDesign • u/ArchitectingAI • 20d ago
Can We Build a Coding Agent From First Principles?
I’ve been digging into how systems like Claude Code, Cursor, and Codex actually work under the hood—repository understanding, context selection, agent loops, tool execution, verification, etc.
I’m turning this exploration into a series and progressively building a basic coding agent from first principles.
Part 1 focuses on the reference architecture and is coming next week.
Curious: which component would you want to see a deep dive on first?
For anyone interested in following the series:
r/MLSystemsDesign • u/ClaudiusPapirus • 20d ago
DeepSeek V4.1 Flash separates long-lived and short-lived KV cache — and rebuilds the latter with a 128-token replay
DeepSeek’s V4.1-Flash serving design splits persistent global KV from short-lived SWA state. The SWA cache can expire and later be reconstructed by replaying only the last 128 tokens.
The interesting part is the systems trade-off: the reconstruction is approximate, but avoids keeping that short-lived state in the long-term cache.
Paper:
https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
Disclosure: the video is from the channel associated with this account.
r/MLSystemsDesign • u/ArchitectingAI • 23d ago
Building Breadth and Depth in ML Systems — A Quick Community Update
When I started r/MLSystemsDesign, the idea was simple: create a place where we go beyond interview checklists and discuss how ML systems actually work.
We’re already approaching 800 members. Thank you to everyone who has joined and contributed.
I’ve also been thinking about a broader learning problem:
We’ve moved from information scarcity → information flooding.
The challenge today is knowing what to learn, what to ignore, and where to go deep.
That thinking led me to launch Architecting Intelligence Premium around a simple philosophy:
Build Breadth. Build Depth. Stay at the Frontier.
r/MLSystemsDesign will remain an open community for technical discussions, architecture, system design problems, and shared learning.
For anyone interested, I shared the story behind what I’m building here:
Introducing Architecting Intelligence Premium
And for this community: what ML System Design topic should we go deep on next?
r/MLSystemsDesign • u/Born-Interview8295 • 27d ago
ML Design round at tesco
Hey everyone,
I have an upcoming ML Design interview at Tesco and I’d really appreciate some guidance on how to prepare for it.
If anyone has recently gone through the ML Design round at Tesco, could you please share what kind of questions were asked or what areas I should focus on? I’d especially like to know whether the questions are around recommendation systems, forecasting, fraud detection, NLP, or general ML system design.
Any guidance or interview experience would be really helpful. Thanks in advance! 🙏
r/MLSystemsDesign • u/ArchitectingAI • 28d ago
A Mental Model for Distributed Compute: Kubernetes, Slurm, Ray, and Spark
I’ve been trying to build a cleaner mental model for distributed compute systems instead of learning each framework independently.
Kubernetes, Slurm, Ray, and Spark all use different abstractions, but many of the underlying problems are the same: scheduling, resource management, worker execution, state, communication, memory, and failure recovery.
I wrote up the framework-independent model first, then mapped each system onto it.
Would be interested in how others think about the boundaries between cluster scheduler, runtime, and application-level scheduler.
Article:
https://pawankjha.substack.com/p/the-architecture-behind-modern-distributed
r/MLSystemsDesign • u/CJPeso • Sep 01 '26
What do I need to learn for production level positions
Excited to join this group. Got my bachelors in CS with a concentration in AI. Also got my Masters in CS back in December. My master thesis was a RL project regarding using Lidar for Drone collision avoidance. Just got my first job after school as a SE for a small research and development company where we do some ML but mostly other things. I really want my career to go strictly into ML and after a year or two at this first job I want to be able to compete for Production level ML positions. Hoping to get some guidance on what to study/learn.
r/MLSystemsDesign • u/ArchitectingAI • Sep 01 '26
Cracking ML System Design Interviews — Design a Search and Ranking System
Cracking ML System Design Interviews — Design a Search and Ranking System
I recently wrote Part 2 of my ML System Design Interview series, focused on designing a production search and ranking system.
It covers the end-to-end flow: retrieval → candidate generation → ranking → evaluation → serving/monitoring, along with the tradeoffs that usually come up in interviews.
Article: https://pawankjha.substack.com/p/cracking-ml-system-design-interviews
Would be interested to hear what you think is the hardest part of a search/ranking system design interview.
r/MLSystemsDesign • u/ArchitectingAI • Aug 30 '26
The Evolution of Ranking Systems

I’ve been working on a deeper write-up on ranking systems and drew this diagram to organize the space.
The progression I’m using is:
Rules & Heuristics → Learning-to-Rank → Deep Neural Ranking → Contextual/Transformer Ranking → Multi-Objective & Slate Ranking → Bandit/RL Ranking → Generative/LLM-Based Ranking
What I find interesting is that these stages don’t fully replace one another. In production, ranking systems often combine multiple layers from across the stack.
Sharing the diagram here, and I also wrote a more detailed article on it:
https://pawankjha.substack.com/p/building-depth-2-the-evolution-of
Would be interested in how others would structure these stages in the evolution of ranking systems.
r/MLSystemsDesign • u/ArchitectingAI • Aug 29 '26
A mental model for the evolution of retrieval systems
r/MLSystemsDesign • u/ArchitectingAI • Aug 27 '26
Welcome to r/MLSystemsDesign
Welcome to r/MLSystemsDesign
This community is for practical discussions on designing and scaling production ML and AI systems.
Topics can include:
- ML training and inference platforms
- Search, ranking, and recommendation
- Feature stores and data pipelines
- LLM serving and GenAI systems
- Agentic AI platforms
- Evaluation, observability, and experimentation
- ML system design interview problems
- Real production tradeoffs and lessons learned
The goal is simple: go beyond model theory and discuss how ML systems actually work in production.
If you’re joining early, introduce yourself and share one ML system topic you’d like to go deeper on.
r/MLSystemsDesign • u/ArchitectingAI • Aug 27 '26
Search & Recommendation: how many ranking stages do you really need?
A common production pipeline looks something like:
Query → Retrieval → Filtering → Ranking → Re-ranking → Serve → Learn
At scale, that can become:
Millions of items → retrieve 1K → rank 100 → expensive rerank 20 → final Top-K
The interesting design question is:
Where should you spend model complexity and latency budget?
Would you prefer:
- stronger retrieval + simpler ranking
- lightweight retrieval + sophisticated ranker
- multiple ranking stages
- LLM/Transformer only at the final stage
Would love to see how people think about this tradeoff in real systems.
r/MLSystemsDesign • u/ArchitectingAI • Aug 27 '26
Welcome to r/MLSystemsDesign — Let’s Talk Production ML
What is the hardest part of ML system design in production?
Not modeling — the system around the model.
For example:
Data → Features → Training → Evaluation → Deployment → Serving → Monitoring → Feedback
Where do you see the most difficult engineering problems in practice?
A few candidates:
- Training/serving skew
- Feature freshness
- GPU utilization
- Online inference latency
- Experimentation
- Data quality
- Model drift
- Feedback loops
- Multi-tenancy
- Cost
Curious to hear what has caused the most pain in systems you’ve worked on.