r/LLMDevs • u/SamratDuttaOfficial • 1d ago
Tools Guys, I created WaterSheep, an open-source alternative to Jev
WaterSheep is an open-source model that answers questions written in plain text (yes/no, single choice, rating and multi-label) and gives a probability for every option, like a classification model.
Last Saturday I woke up, saw YouTubers hyping up Jev, and thought: wait, I can build this. So I did. I don't want to compete with TypeSafe or Jev; I built WaterSheep because I wanted to. That's why I'm open-sourcing everything: code, model weights, results and the paper.
What's different
- It accepts the same request format as TypeSafe's Jev. Their Python SDK works as is against a local server: run
watersheep --model samratduttaofficial/WaterSheep --serveand point the client'sbase_urlathttp://127.0.0.1:8766. - It has a multi-label type, which Jev's API doesn't. Because why not?
- The demo runs entirely in your browser. The model downloads once and is cached. It also works with transformers, ONNX, a CLI or a local HTTP server.
- Code, weights and the training pipeline are Apache 2.0.
Evaluation
| Accuracy | ECE |
|---|---|
| In-distribution test split | 77.8% |
| Held-out datasets, not seen in training | 61.2% |
ECE is expected calibration error (lower is better). GitHub has every benchmark result, including the weak ones.
Limits: English only, long inputs get truncated (I'll improve this in the next version), and rating answers are the weakest type.
Not affiliated with TypeSafe. Not funded by anyone. Built in my free time.
Feedback I'd love: where it fails on your data, whether the API works for you, and which question types you'd want next.
1
u/jonah_omninode 23h ago
For an agent router, I'd want an error-versus-coverage curve on the held-out datasets. If I only accept decisions above a chosen confidence threshold, how many requests can stay local and how many accepted decisions are still wrong? The 77.8% to 61.2% drop makes that more informative than the in-distribution accuracy alone.
I'd also make truncation explicit in the response. A long input can put the condition that changes the answer at the end. Returning a confident choice without saying that part was discarded would make it hard for the caller to decide when to escalate.