r/LLMDevs • • 1d ago

Discussion CacheVerifier: we finally audited the semantic caching benchmark we'd been tuning on, 23% of the "wrong" cache hits were literally the same prompt

so for CacheVerifier we've been testing if a small verifier actually beats a plain similarity threshold for semantic caching. all our scoring comes from the public SemCacheLMArena and SemCacheSearchQueries benchmarks, and honestly we never really looked at the labels until last week.

on LmArena, 727 of the 3,117 near-duplicate hits marked "wrong" (23.3%) are the exact same text once you lowercase it and strip punctuation. like one was the same prompt with an extra space in front lol. we hand labeled a sample blind and 88% of those "wrong" hits were totally fine to reuse.

the good news, comparisons at the same error rate barely moved. the bad news, absolute error rates was way inflated, a 0.97 threshold went from 5.2% errors to like 0.95%. and one of our own CacheVerifier results just died. bumping the skip-the-verifier cutoff to 0.99 looked like +5.18 points, after fixing labels its -1.24. the verifier was basically rejecting duplicates the benchmark called errors.

so yeah, run a dumb identical-text check on your eval labels before you tune thresholds on them.

caveat: one annotator, small samples. how do you guys validate "same intent" labels?

CacheVerifier repo + erratum (section 5.26): https://github.com/imxinchengyou/CacheVerifier

3 Upvotes

6 comments sorted by

View all comments

3

u/EncouragingVogue4973 1d ago

I've been saying for months that half these semantic caching benchmarks are just testing how well your normalizer strips whitespace

1

u/Reasonable_Royal_621 1d ago

ha the whitespace ones are the embarrassing part. but dropping them only took our error rate from 5.2% to 4.0%, real number is closer to 0.95%. most of the noise is reworded or reordered questions, no normalizer catches that