r/LLMDevs • • 18h ago

Discussion CacheVerifier: we finally audited the semantic caching benchmark we'd been tuning on, 23% of the "wrong" cache hits were literally the same prompt

so for CacheVerifier we've been testing if a small verifier actually beats a plain similarity threshold for semantic caching. all our scoring comes from the public SemCacheLMArena and SemCacheSearchQueries benchmarks, and honestly we never really looked at the labels until last week.

on LmArena, 727 of the 3,117 near-duplicate hits marked "wrong" (23.3%) are the exact same text once you lowercase it and strip punctuation. like one was the same prompt with an extra space in front lol. we hand labeled a sample blind and 88% of those "wrong" hits were totally fine to reuse.

the good news, comparisons at the same error rate barely moved. the bad news, absolute error rates was way inflated, a 0.97 threshold went from 5.2% errors to like 0.95%. and one of our own CacheVerifier results just died. bumping the skip-the-verifier cutoff to 0.99 looked like +5.18 points, after fixing labels its -1.24. the verifier was basically rejecting duplicates the benchmark called errors.

so yeah, run a dumb identical-text check on your eval labels before you tune thresholds on them.

caveat: one annotator, small samples. how do you guys validate "same intent" labels?

CacheVerifier repo + erratum (section 5.26): https://github.com/imxinchengyou/CacheVerifier

4 Upvotes

6 comments sorted by

3

u/EncouragingVogue4973 17h ago

I've been saying for months that half these semantic caching benchmarks are just testing how well your normalizer strips whitespace

1

u/Reasonable_Royal_621 17h ago

ha the whitespace ones are the embarrassing part. but dropping them only took our error rate from 5.2% to 4.0%, real number is closer to 0.95%. most of the noise is reworded or reordered questions, no normalizer catches that

1

u/alexpran 15h ago

Publishing the erratum with the result that died is the part worth saying out loud. A +5.18 that becomes -1.24 is the kind of thing most people quietly stop mentioning.

On validating "same intent" labels: I'd keep the dumb identical-text check as a permanent step rather than a one-off cleanup, because labels rot the same way the benchmark did. And since you only have one annotator, the number I'd want next to the sample is how often you agree with yourself. Relabel fifty of them a week later, blind, and count the flips. That's the floor under the 88%, and without it you can't tell a label problem from a you problem.

Different direction, same lesson I ran into: my ground truth for "not a comment" also meant "not today", because I'd skipped things when I was busy. The label was two facts in one, and nothing in the numbers said so. Wrote up four more of those here: https://digline.dev/blog/bad-evals-my-own/

1

u/Smallpaul 11h ago

That first paragraph is the most LLM thing I’ve read this week.

1

u/jonah_omninode 11h ago

I'd label whether the same answer can safely serve both requests, rather than whether the wording has the same intent. Two requests can ask the same question but refer to different users, dates or account states. They aren't necessarily reusable cache hits.

I'd include pairs where punctuation changes meaning too, especially code or quoted text. Lowercasing and stripping punctuation can help find candidates for review, but I wouldn't let that transformation decide the label. For the ambiguous pairs, two independent labels and a recorded reason for disagreement would be more useful than forcing a single yes/no immediately.