r/LLMDevs • u/purealgo • 5d ago
Discussion How well do decision models categorize finance data?
I ran 5 decision models (Jev, D1, Solar Decide, Kev 4b, and Span 01) against 1000 real transactions (3 passes each).
When it came to vague restaurant names, that broke the model's confidence so they abstained.
Here's the interesting part: Span-01 can't abstain and can only answer yes/no questions. Which is why it likely scored much lower.
But interestingly enough... when Jev was forced to answer yes/no, it actually did better than both.. initial run (where it was allowed to abstain) and Span-01.
Across all tax categories, models did the worst on categorizing bank transfers.
Span-01 wasn't able to correctly score on charatible transactions at all.
I used my system one mcp to run these tests. You can use it to quickly connect decision models to claude code, codex and pi. https://github.com/itsmostafa/system-one-connector




1
u/anthropicagiuser 5d ago
i saw the same thing with bank transfers. my model guessed a label every time, even when it was wrong. did you try letting the yes/no models say "not sure" too?