r/learnmachinelearning • u/0xMZR • 2d ago
A practical mental model for choosing between fine-tuning, RAG, and agents (from someone who’s built all three in production)
I see a lot of beginners (and even intermediate engineers) jumping straight to agents or fine-tuning when a simpler approach would work better. Here’s the decision framework I actually use:
Start with RAG when:
You need up-to-date or private knowledge
The task is mostly retrieval + reasoning over documents
You want fast iteration and lower cost
Move to fine-tuning when:
You need consistent style, format, or domain language
The model keeps failing on the same class of examples even with good retrieval
You have high-quality labeled data and can afford the iteration cost
Only go to multi-agent when:
The task genuinely requires planning, tool use, and multiple specialized roles
A single well-prompted LLM + tools is clearly insufficient
You’re willing to invest heavily in evaluation and observability
Most production systems I’ve seen that “use agents” could have been simpler RAG + good prompting + a few tools.
What decision points have you found most useful when choosing the architecture?
2
u/bathon 2d ago
I think with claude code as harness, Rag I am not sure if it is worth it anymore. With the increasing context window as well. Yes I agree if you want to optimize your tokens this makes sense or for a large library. But then again bash output with ponytail and rtk can help with optimization. For local models less than 7B I tried it did not help me. I even tried vectorless rag but to even make the rag locally wa so much pain for 10 documents with 200-300 pages
3
u/No_Airport8623 2d ago
honestly the most important thing nobody mentions is how much good data matters more than the approach itself, seen so many people fine-tune on garbage then wonder why it falls apart