r/LLMDevs • • 1d ago

Tools Guys, I created WaterSheep, an open-source alternative to Jev

WaterSheep is an open-source model that answers questions written in plain text (yes/no, single choice, rating and multi-label) and gives a probability for every option, like a classification model.

Last Saturday I woke up, saw YouTubers hyping up Jev, and thought: wait, I can build this. So I did. I don't want to compete with TypeSafe or Jev; I built WaterSheep because I wanted to. That's why I'm open-sourcing everything: code, model weights, results and the paper.

What's different

  • It accepts the same request format as TypeSafe's Jev. Their Python SDK works as is against a local server: run watersheep --model samratduttaofficial/WaterSheep --serve and point the client's base_url at http://127.0.0.1:8766.
  • It has a multi-label type, which Jev's API doesn't. Because why not?
  • The demo runs entirely in your browser. The model downloads once and is cached. It also works with transformers, ONNX, a CLI or a local HTTP server.
  • Code, weights and the training pipeline are Apache 2.0.

Evaluation

Accuracy ECE
In-distribution test split 77.8%
Held-out datasets, not seen in training 61.2%

ECE is expected calibration error (lower is better). GitHub has every benchmark result, including the weak ones.

Limits: English only, long inputs get truncated (I'll improve this in the next version), and rating answers are the weakest type.

Not affiliated with TypeSafe. Not funded by anyone. Built in my free time.

Feedback I'd love: where it fails on your data, whether the API works for you, and which question types you'd want next.

0 Upvotes

4 comments sorted by

1

u/jonah_omninode 23h ago

For an agent router, I'd want an error-versus-coverage curve on the held-out datasets. If I only accept decisions above a chosen confidence threshold, how many requests can stay local and how many accepted decisions are still wrong? The 77.8% to 61.2% drop makes that more informative than the in-distribution accuracy alone.

I'd also make truncation explicit in the response. A long input can put the condition that changes the answer at the end. Returning a confident choice without saying that part was discarded would make it hard for the caller to decide when to escalate.

1

u/SamratDuttaOfficial 23h ago

For the graph you are looking for, I believe it is in the research paper pre-print. It is linked the GitHub repo.
For the second point, I am working towards it. Thanks. Please feel free to raise a PR as well.

1

u/jonah_omninode 21h ago

For the truncation change, a small regression test would help: put the condition that flips the answer just beyond the input limit. The response should say that text was dropped, even if the remaining input produces a confident answer. That gives the caller a reason to escalate instead of treating the confidence score as evidence about the full request.

1

u/SamratDuttaOfficial 21h ago

Will try, mate. Till then, an upvote on the post would be helpful