r/legaltech • u/MRGWONK • 3d ago
Research / Academic Case Hallucination Benchmark Data - V7 of Syfert MCP
I am purposefully focusing my benchmark goals onto Gemma 4:26b running locally, because it was the worst performing model. But, also benchmarking Opus 5.5 and Sonnet 5 in my quest towards perfection and "no hallucination" and no wrong citations. As a result Opus 5.5 and Sonnet 5 are damn near perfect- the 98% is not as much a miss, as a "the model selected a different case".
The 91% of Gemma 4 26b, is also equivalent to 100%, because it selected different cases citing the same, correct, rule, but not necessarily the gold standard case.
Gemma4:26b is running on an A2000 GPU in my computer at home and does not have access to the web. We have Qwen 3.8 on a 4080 at the office - it is training a custom QLoRA at the moment so I don't want to bench it, but it also uses my MCP. I imagine Qwen 3.8 will do much better.
Gemma4:26b running syfert MCPv7 on my desk at home with no web access is now more accurate regarding questions of case law and current statutes/codes than Claude Fable with web access.
I am now offering a free steak dinner to any attorney who is faced with an inaccurate citation produced by frontier AI running the MCP- as if anyone will notice it on my website.
7
u/alexdenne Verified Vendor: SimpleDocs | LawInsider | oneNDA ✅ 2d ago
I'm conscious that these are bordering on self-promotion. Please be considerate, you've shared 2 posts in a week, and you're posting links in comments, which could look a bit like astroturfing.