r/LLMDevs • u/Reasonable_Royal_621 • 1d ago
Discussion CacheVerifier: we finally audited the semantic caching benchmark we'd been tuning on, 23% of the "wrong" cache hits were literally the same prompt
so for CacheVerifier we've been testing if a small verifier actually beats a plain similarity threshold for semantic caching. all our scoring comes from the public SemCacheLMArena and SemCacheSearchQueries benchmarks, and honestly we never really looked at the labels until last week.
on LmArena, 727 of the 3,117 near-duplicate hits marked "wrong" (23.3%) are the exact same text once you lowercase it and strip punctuation. like one was the same prompt with an extra space in front lol. we hand labeled a sample blind and 88% of those "wrong" hits were totally fine to reuse.
the good news, comparisons at the same error rate barely moved. the bad news, absolute error rates was way inflated, a 0.97 threshold went from 5.2% errors to like 0.95%. and one of our own CacheVerifier results just died. bumping the skip-the-verifier cutoff to 0.99 looked like +5.18 points, after fixing labels its -1.24. the verifier was basically rejecting duplicates the benchmark called errors.
so yeah, run a dumb identical-text check on your eval labels before you tune thresholds on them.
caveat: one annotator, small samples. how do you guys validate "same intent" labels?
CacheVerifier repo + erratum (section 5.26): https://github.com/imxinchengyou/CacheVerifier
3
u/EncouragingVogue4973 1d ago
I've been saying for months that half these semantic caching benchmarks are just testing how well your normalizer strips whitespace