WaterSheep

Calibrated decisions for any text

Yes/no, single-choice, rating and multi-label answers, each with a reliable probability.

Apache 2.0 · Open weights

Live demo

Runs entirely in your browser. Data stays on your device.

Also on Hugging Face Spaces.

Capabilities

Options are defined per request, so one model covers routing, triage, moderation, scoring and tagging.

noul

Yes / No

Probability of yes.

choice

Single choice

The option, with a probability for each.

score

Rating

The expected level, with a probability for each.

multi

Multi-label

Every option above the threshold, with probabilities.

Design

  • Structured. Typed answers, no text to parse.
  • Calibrated. Probabilities you can use as confidence scores.
  • Open. Code and weights under Apache 2.0, trained on openly licensed data. See NOTICE.
  • Portable. Transformers, JavaScript and ONNX.

Get started

The model downloads from Hugging Face on first use, then runs locally.

Python

pip install transformers torch
from transformers import pipeline

ws = pipeline(model="samratduttaofficial/WaterSheep", trust_remote_code=True)
ws("I was charged twice.", question="Which team should handle this?",
   options=["billing", "shipping", "support"])

Several questions at once

With the Python package, decide returns the answer, its confidence and a probability for every option. ask answers several questions about one text:

pip install git+https://github.com/SamratDuttaOfficial/WaterSheep
from watersheep import WaterSheep

ws = WaterSheep.load("samratduttaofficial/WaterSheep")
ws.ask({
    "state": {"customer": "Priya (premium plan)",
              "message": "Charged twice for order #4411 and the package is 12 days late."},
    "questions": {
        "escalate": {"type": "noul", "instructions": "Should a human agent take over now?"},
        "team": {"type": "choice", "instructions": "Which team should handle this?",
                 "criteria": {"billing": "payments, refunds", "shipping": "delivery problems"}},
        "frustration": {"type": "score", "instructions": "How frustrated is the customer?",
                        "criteria": ["calm", "annoyed", "frustrated", "furious"]},
        "issues": {"type": "multi", "instructions": "Which issues are reported?",
                   "criteria": ["double charge", "late delivery", "damaged item"]},
    },
})

Download

hf download samratduttaofficial/WaterSheep --local-dir WaterSheep

Or with Git (requires Git LFS):

git clone https://huggingface.co/samratduttaofficial/WaterSheep

Then load it from the folder, offline:

ws = pipeline(model="WaterSheep", trust_remote_code=True)

API

Deploy it as a Hugging Face Inference Endpoint, then:

curl https://YOUR-ENDPOINT -H "Authorization: Bearer $HF_TOKEN" -H "Content-Type: application/json" -d '{"inputs": "I was charged twice.", "parameters": {"question": "Which team should handle this?", "options": ["billing", "shipping", "support"]}}'

JavaScript

<script type="module">
  import { decide } from "https://samratduttaofficial.github.io/WaterSheep/watersheep.js";
  console.log(await decide("I was charged twice.", "Which team should handle this?", ["billing", "shipping", "support"]));
</script>

No install. With a downloaded copy on your web server, call load({ base: "WaterSheep/" }) first.

Evaluation

ECE is the expected calibration error (lower is better).

EvaluationAccuracyECE
In-distribution test split77.8%0.026
Held-out datasets, not seen in training61.2%0.043
Reliability diagrams: accuracy against confidence for each question type
Accuracy against confidence for each question type, before (raw) and after calibration.

Benchmarks

BenchmarkSuiteQuestionsAccuracyECEIn training data
goemotionssentiment2,00022.4%0.023other split
hatechecksafety2,00075.1%0.139no
legal_abercrombielegal9521.1%0.316no
legal_contract_nli_confidentiality_of_agreementlegal8269.5%0.177no
legal_corporate_lobbyinglegal49068.4%0.216no
legal_cuad_audit_rightslegal1,21686.3%0.041no
legal_definition_classificationlegal1,33756.9%0.279no
legal_function_of_decision_sectionlegal36724.3%0.245no
legal_hearsaylegal9456.4%0.307no
legal_overrulinglegal2,00062.5%0.151no
legal_personal_jurisdictionlegal5050.0%0.160no
legal_privacy_policy_qalegal2,00058.9%0.274no
legal_proalegal9551.6%0.379no
legal_ucc_v_common_lawlegal9462.8%0.171no
prompt_injectionsafety11691.4%0.079other split
xstestsafety45073.6%0.140no

Training

  • Base model: answerdotai/ModernBERT-base, fine-tuned with a decision head.
  • Data: openly licensed public datasets (listed in NOTICE) and synthetic decisions from Qwen3.5-4B.
  • Calibration: a temperature per question type, fitted on a validation split.
Training loss, learning rate and validation accuracy by question type
Training loss and learning rate (left); validation accuracy by question type (right).
Share of synthetic examples kept after verification
Share of synthetic examples kept after verification, by question type (left) and by family (right).

Limitations

  • English only.
  • Long inputs are truncated.
  • Rating-scale answers are less accurate than the other types.
  • Probabilities are calibrated on data like the training data; validate them on your own.
  • Not for high-stakes decisions (medical, legal, financial, hiring) on its own.

License and citation

Apache 2.0 (LICENSE). Attributions: NOTICE.

@misc{watersheep,
  author = {Samrat Dutta},
  title  = {WaterSheep: calibrated decisions for any text},
  year   = {2026},
  url    = {https://huggingface.co/samratduttaofficial/WaterSheep}
}