Calibrated decisions for any text
Yes/no, single-choice, rating and multi-label answers, each with a reliable probability.
Capabilities
Options are defined per request, so one model covers routing, triage, moderation, scoring and tagging.
Yes / No
Probability of yes.
Single choice
The option, with a probability for each.
Rating
The expected level, with a probability for each.
Multi-label
Every option above the threshold, with probabilities.
Design
- Structured. Typed answers, no text to parse.
- Calibrated. Probabilities you can use as confidence scores.
- Open. Code and weights under Apache 2.0, trained on openly licensed data. See NOTICE.
- Portable. Transformers, JavaScript and ONNX.
Get started
The model downloads from Hugging Face on first use, then runs locally.
Python
pip install transformers torch
from transformers import pipeline ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True) ws("I was charged twice.", question="Which team should handle this?", options=["billing", "shipping", "support"])
Several questions at once
With the Python package, decide returns the answer, its confidence and a probability for every option. ask answers several questions about one text:
pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
from watersheep import WaterSheep ws = WaterSheep.load("samratduttaofficial/WaterSheep") ws.ask({ "state": {"customer": "Priya (premium plan)", "message": "Charged twice for order #4411 and the package is 12 days late."}, "questions": { "escalate": {"type": "noul", "instructions": "Should a human agent take over now?"}, "team": {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"billing": "payments, refunds", "shipping": "delivery problems"}}, "frustration": {"type": "score", "instructions": "How frustrated is the customer?", "criteria": ["calm", "annoyed", "frustrated", "furious"]}, "issues": {"type": "multi", "instructions": "Which issues are reported?", "criteria": ["double charge", "late delivery", "damaged item"]}, }, })
Download
hf download samratduttaofficial/WaterSheep --local-dir WaterSheep
Or with Git (requires Git LFS):
git clone https://huggingface.co/samratduttaofficial/WaterSheep
Then load it from the folder, offline:
ws = pipeline(model="WaterSheep", trust_remote_code=True)
API
Deploy it as a Hugging Face Inference Endpoint, then:
curl https://YOUR-ENDPOINT -H "Authorization: Bearer $HF_TOKEN" -H "Content-Type: application/json" -d '{"inputs": "I was charged twice.", "parameters": {"question": "Which team should handle this?", "options": ["billing", "shipping", "support"]}}'
JavaScript
<script type="module"> import { decide } from "https://samratduttaofficial.github.io/WaterSheep/watersheep.js"; console.log(await decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"])); </script>
No install. With a downloaded copy on your web server, call load({ base: "WaterSheep/" }) first.
Evaluation
ECE is the expected calibration error (lower is better).
| Evaluation | Accuracy | ECE |
|---|---|---|
| In-distribution test split | 77.8% | 0.026 |
| Held-out datasets, not seen in training | 61.2% | 0.043 |
Benchmarks
| Benchmark | Suite | Questions | Accuracy | ECE | In training data |
|---|---|---|---|---|---|
| goemotions | sentiment | 2,000 | 22.4% | 0.023 | other split |
| hatecheck | safety | 2,000 | 75.1% | 0.139 | no |
| legal_abercrombie | legal | 95 | 21.1% | 0.316 | no |
| legal_contract_nli_confidentiality_of_agreement | legal | 82 | 69.5% | 0.177 | no |
| legal_corporate_lobbying | legal | 490 | 68.4% | 0.216 | no |
| legal_cuad_audit_rights | legal | 1,216 | 86.3% | 0.041 | no |
| legal_definition_classification | legal | 1,337 | 56.9% | 0.279 | no |
| legal_function_of_decision_section | legal | 367 | 24.3% | 0.245 | no |
| legal_hearsay | legal | 94 | 56.4% | 0.307 | no |
| legal_overruling | legal | 2,000 | 62.5% | 0.151 | no |
| legal_personal_jurisdiction | legal | 50 | 50.0% | 0.160 | no |
| legal_privacy_policy_qa | legal | 2,000 | 58.9% | 0.274 | no |
| legal_proa | legal | 95 | 51.6% | 0.379 | no |
| legal_ucc_v_common_law | legal | 94 | 62.8% | 0.171 | no |
| prompt_injection | safety | 116 | 91.4% | 0.079 | other split |
| xstest | safety | 450 | 73.6% | 0.140 | no |
Training
- Base model: answerdotai/ModernBERT-base, fine-tuned with a decision head.
- Data: openly licensed public datasets (listed in NOTICE) and synthetic decisions from Qwen3.5-4B.
- Calibration: a temperature per question type, fitted on a validation split.
Limitations
- English only.
- Long inputs are truncated.
- Rating-scale answers are less accurate than the other types.
- Probabilities are calibrated on data like the training data; validate them on your own.
- Not for high-stakes decisions (medical, legal, financial, hiring) on its own.