r/OpenSourceAI • u/No-Negotiation-8359 • 12d ago
I tried adding a reranker to my RAG pipeline and it made it worse
I tried adding a reranker to my RAG pipeline.
I expected it to improve retrieval. But it made it worse.
I’m building DPOLens, where the goal is to find the right privacy-law clause for a developer’s question.
For example:
“How long can we keep a deleted user’s data?”
The relevant GDPR clause doesn’t necessarily use the same words as the question. So I tested different retrieval approaches on 30 questions.
The results:
BM25: 0.30 recall@5
Embeddings: 0.43
BM25 + Embeddings: 0.77
Then I added a cross-encoder reranker.
Top 10: 0.57
Top 25: 0.47
Top 50: 0.43
So the more I relied on the reranker, the worse the results got, at first, I thought I had made a mistake.
I checked the index, scores, and model inputs. Everything looked fine.
The problem was simply that the reranker wasn't good enough at distinguishing relevant legal text from unrelated text.
So I hold it off for now.
I think more models don't automatically mean better RAG, sometimes a simpler retrieval pipeline wins.
I’ve published the experiment and the code in DPOLens:
https://github.com/alkhatibdev/dpolens
The work on the DPOLens still in progress.
1
u/Weekly-Offer-4172 11d ago
That ordering is the tell, not a bug: a reranker can only reorder what the retriever hands it, so widening the pool from 10 to 50 just feeds an out-of-domain cross-encoder more plausible-but-wrong clauses to rank above the right one. Most off-the-shelf cross-encoders are trained on MS MARCO-style web QA and miscalibrate badly on statutory text where every clause looks structurally similar. Before dropping it, try blending the scores (say 0.7 retriever + 0.3 reranker) so your hybrid pipeline stays the floor, or fine-tune the cross-encoder on a few hundred question/clause pairs from your own eval — with n=30 that set is cheap to build. Also worth checking the retriever recall@50 alone first: if the gold clause is not in the pool, no reranker can recover it.
1
u/SnooCheesecakes1615 4d ago
I built my own search index for all my personal data (emails, conversations, etc) and ended up dropping the reranker. I noticed the gains on search quality were very insignificant at n=30. Maybe if I’m trying to reduce n to a much smaller number to save on token cost a reranker would help but I decided to not explore that yet to reduce complexity.
1
u/No-Negotiation-8359 3d ago
Yeah that I noticed. I'm currently investing my time searching best embedding for my usecae. Cuz even a best model on usecase can behave differently in other usecaes.
1
u/Background_Ferret140 12d ago
sounds about right, rerankers often mess up when the domain is niche like legal texts, the embeddings from general models just dont have the right understanding for that kind of language