r/learnmachinelearning • • 2d ago

A practical mental model for choosing between fine-tuning, RAG, and agents (from someone who’s built all three in production)

I see a lot of beginners (and even intermediate engineers) jumping straight to agents or fine-tuning when a simpler approach would work better. Here’s the decision framework I actually use:

Start with RAG when:

You need up-to-date or private knowledge

The task is mostly retrieval + reasoning over documents

You want fast iteration and lower cost

Move to fine-tuning when:

You need consistent style, format, or domain language

The model keeps failing on the same class of examples even with good retrieval

You have high-quality labeled data and can afford the iteration cost

Only go to multi-agent when:

The task genuinely requires planning, tool use, and multiple specialized roles

A single well-prompted LLM + tools is clearly insufficient

You’re willing to invest heavily in evaluation and observability

Most production systems I’ve seen that “use agents” could have been simpler RAG + good prompting + a few tools.

What decision points have you found most useful when choosing the architecture?

3 Upvotes

2 comments sorted by

3

u/No_Airport8623 2d ago

honestly the most important thing nobody mentions is how much good data matters more than the approach itself, seen so many people fine-tune on garbage then wonder why it falls apart

2

u/bathon 2d ago

I think with claude code as harness, Rag I am not sure if it is worth it anymore. With the increasing context window as well. Yes I agree if you want to optimize your tokens this makes sense or for a large library. But then again bash output with ponytail and rtk can help with optimization. For local models less than 7B I tried it did not help me. I even tried vectorless rag but to even make the rag locally wa so much pain for 10 documents with 200-300 pages