Instructions to use sthanika-ai/Sieve-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use sthanika-ai/Sieve-4B with PEFT:
from peft import PeftModel from transformers import AutoModel base_model = AutoModel.from_pretrained("Qwen/Qwen3.5-4B") model = PeftModel.from_pretrained(base_model, "sthanika-ai/Sieve-4B") - Notebooks
- Google Colab
- Kaggle
Sieve-4B
Sieve-4B is a decision model from the Sieve family. You give it a piece of text or JSON (the state) and typed questions, and it returns a calibrated probability for every option of every question. It reads the state once, scores the options directly, and never generates text.
It is a LoRA adapter (rank 32) and a 255-way answer head on Qwen/Qwen3.5-4B
(revision 851bf6e8), with one calibration temperature per question type. Its readout differs from the pointer head
of Sieve-27B and Sieve-9B-Plus, so it ships its own inference code (sieve4b/, below).
Earlier weights. This revision replaces the first Sieve-4B weights (same architecture and recipe, retrained on
newer data). They remain available at revision
e49deceb (Decision Index
0.3 public index 43.69).
How it works
- Backbone. The post-trained Qwen3.5-4B, a hybrid of Gated DeltaNet and full-attention layers. Only its text part is used; the model cannot produce text.
- Adapter. LoRA r=32, alpha 64, on every linear layer: attention, MLP and the Gated DeltaNet projections.
- Answer head. Each option is listed under a code (A, B, ..., AA, ...). A 255-way linear head, initialised from the base model's own output rows for those codes, scores all codes at the question's decision point. The head reads the whole option list.
- Isolation. The state is encoded once; each question runs as its own branch from that cache, so questions never see each other.
- Yes/no in both orders. A yes/no question is asked with the options in both orders and the answers are averaged.
| type | options | returns |
|---|---|---|
choice |
the caller's keys, optionally described (up to 255) | a probability per key |
noul |
yes / no | P(yes) |
score |
ordered levels (up to 10) | a probability per level |
Use
hf download sthanika-ai/Sieve-4B --local-dir Sieve-4B && cd Sieve-4B
pip install -r requirements.txt # torch, transformers>=5.17, peft>=0.21, safetensors, pydantic
pip install flash-linear-attention==0.5.2 # recommended on GPU: the Gated DeltaNet kernels our runs used
from sieve4b import load, decide
engine = load(".", device="cuda:0") # this folder; fetches Qwen/Qwen3.5-4B at its pinned revision on first use
decide(engine, {"subject": "Duplicate charge on invoice 4411", "body": "Billed twice. Refund today or we cancel."},
{"department": {"type": "choice", "instructions": "Which team handles this?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages"}},
"churn_risk": {"type": "noul", "instructions": "Does the customer threaten to cancel?"}})
- Server.
python -m sieve4b.server --ckpt . --port 8123 --name sieve-4bserves the same thing as a/v1/systemoneHTTP endpoint (needsfastapianduvicorn). The Decision Index run below used it. - GPU memory. About 10 GB in bf16. Longer states need more.
- Token limits. Up to 32,768 tokens for the state plus the longest question, 65,536 per request, 255 choice options and 10 score levels. Longer requests are refused, never truncated. Training records reached 16,806 tokens.
- Temperatures.
sieve_config.jsonstores one temperature per question type: choice 1.161, yes/no 1.440 and score 1.000. A temperature never changes an answer.
Results
Decision Index 0.3
Sieve-4B scores 50.06 (raw index 62.10, breadth 48.32) on the public 0.3 suite of the
Decision Index, on our own run with the official
kit at 9eb2dbe (0.3) and its scorer.
- Coverage. All 140,178 scoreable requests over 43 benchmarks were answered, with no truncation, no unsupported request and no error. The index averages 37 of the benchmarks.
- How the run was made. A complete run of the 0.2.1 suite, extended to 0.3 with the 2,638 rebuilt GSM8K requests, on two A100 80GB PCIe GPUs with three server processes per GPU.
- Results. sthanika-ai/Sieve-4B-decision-index-results.
- Not on the leaderboard yet. The score is self-reported. On the 0.3 board the ranking uses the Full score, where this public suite counts 20% and the maintainers' private tests count 80%.
| Decision Index 0.3 | Knowledge & Reasoning | Language Understanding | Retrieval & Classification | Tools & Automation | Arts & Human Taste |
|---|---|---|---|---|---|
| 50.06 | 33.9 | 53.9 | 61.6 | 66.8 | 28.3 |
The index and area scores are chance-corrected: 0 is random guessing and 100 is perfect. Below, the score is in each benchmark's own metric and the skill is chance-corrected. On edition 0.2.1 the same run scores 49.21.
All 43 benchmarks
| benchmark | metric | score | skill (×100) | requests |
|---|---|---|---|---|
| ACOS | per-review F1 | 0.206 | 18.1 | 1,565 |
| Amazon ESCI | macro-F1 | 0.491 | 36.2 | 5,000 |
| ANLI | macro-F1 | 0.635 | 45.3 | 3,200 |
| API-Bank | accuracy | 0.738 | 73.3 | 508 |
| BANKING77 | macro-F1 | 0.828 | 82.6 | 3,080 |
| BBH fixed-option tasks | accuracy | 0.682 | 53.9 | 5,507 |
| BFCL | case exact accuracy | 0.944 | 92.4 | 1,694 |
| BPoMP | accuracy | 0.804 | 61.0 | 5,000 |
| BRIGHT | nDCG@10 | 0.444 | 37.1 | 220 |
| cfcolor | accuracy | 0.613 | 17.8 | 5,000 |
| ChessBench | accuracy | 0.138 | 6.1 | 5,000 |
| CLadder | accuracy | 0.639 | 27.8 | 5,000 |
| CLINC150+OOS | macro-F1 | 0.902 | 90.1 | 5,500 |
| ContractNLI | macro-F1 | 0.798 | 70.8 | 123 |
| CRUXEval | accuracy | 0.558 | 29.9 | 570 |
| FinEntity | macro-F1 | 0.886 | 83.2 | 979 |
| GPQA Diamond | accuracy | 0.423 | 23.1 | 196 |
| GSM8K | accuracy | 0.645 | 57.4 | 2,638 |
| Habermas Machine | accuracy | 0.400 | 12.9 | 1,676 |
| HellaSwag | accuracy | 0.950 | 93.4 | 10,042 |
| HLE | accuracy | 0.094 | 0.0 | 501 |
| Home appliance simulator | case exact accuracy | 0.261 | 26.1 | 88 |
| HoVer claim verification | accuracy | 0.750 | 50.0 | 4,000 |
| Humicroedit | accuracy | 0.595 | 18.9 | 2,628 |
| iSarcasmEval | Sarcasm F1 · track A, English | 0.461 | 30.6 | 4,600 |
| MMLU-Pro | accuracy | 0.550 | 49.4 | 12,032 |
| MuSR | accuracy | 0.550 | 28.5 | 752 |
| New Yorker caption matching | accuracy | 0.616 | 51.9 | 528 |
| NLI4CT | macro-F1 | 0.768 | 54.9 | 5,500 |
| PhishNChips phishing decisions | accuracy | 0.843 | 68.5 | 2,000 |
| POP909-CL | accuracy | 0.079 | 7.2 | 2,000 |
| RAGTruth response-level hallucination | F1 on hallucinated class | 0.779 | 54.1 | 2,700 |
| SATA-Bench | case exact accuracy | 0.206 | 19.6 | 1,650 |
| ToolRet | nDCG@10 | 0.673 | 62.3 | 685 |
| VAST | macro-F1 | 0.524 | 28.6 | 3,006 |
| When2Call MCQ | accuracy | 0.799 | 73.2 | 3,652 |
| WinoGrande | accuracy | 0.858 | 71.6 | 1,267 |
| ARC-Challenge (not in the index) | accuracy | 0.944 | 92.5 | 1,172 |
| ARC-Easy (not in the index) | accuracy | 0.979 | 97.1 | 2,376 |
| MMLU (not in the index) | accuracy | 0.757 | 67.7 | 14,033 |
| RouterBench (not in the index) | selected quality (quality objective) | 0.790 | 50.8 | 10,000 |
| SGD/SGD-X (not in the index) | macro-F1 | 0.382 | 0.0 | 2,500 |
| SimpleBench (not in the index) | accuracy | 0.100 | 0.0 | 10 |
Calibration
| choice | yes/no | |
|---|---|---|
| served temperature | 1.161 | 1.440 |
| ECE on held-out sources, out of fold | 0.039 | 0.030 |
- Fit. The served temperatures were fitted by NLL on 1,456 held-out records from 12 public sources that are not in
the training data, and reported 2-fold out of fold. A fit on the in-domain calibration split gives lower temperatures
(choice 0.914, yes/no 0.862); both reports are in
calibration/. - On the Decision Index. Over the 0.2.1 rows, with our reimplementation of the board's calibration panel (benchmarks with a per-field gold answer, equally weighted; confidence = the probability of the chosen option): ECE 0.019, and 0.5% of answers wrong with confidence of 0.95 or more.
Latency
HTTP wall time as recorded in the Decision Index run, with three server processes sharing each A100 80GB PCIe: median 302 ms, mean 433 ms, 80th percentile 400 ms, 95th percentile 817 ms.
Training
- Recipe.
- LoRA r=32, α=64, dropout 0.05 on the attention, MLP and Gated DeltaNet projections, and the 255-way answer head.
- AdamW, lr 5e-5 for the adapter and 1e-4 for the head, 3% warm-up then cosine, 1 epoch, 128 records per step, bf16.
- Loss: cross-entropy on the target distribution, an option-order consistency term and a small replay term that keeps the model near its base.
- 1,295 steps, 11.6 hours on two A100 80GB PCIe GPUs. No record was skipped or truncated.
- Selection. The last checkpoint. The benchmarks were never used to choose or tune anything.
- Data. 165,832 training records.
- Sources: public datasets and generated decision records.
- Not published. The training files are not part of this release; the SHA-256 of each file is in
training_config.json.
Files
| file | contents |
|---|---|
adapter_model.safetensors, adapter_config.json |
the LoRA adapter (rank 32, PEFT format) |
head.safetensors |
the 255-way answer head |
sieve_config.json |
option codes, prompt preamble and the served temperatures |
calibration/ |
the two temperature fits, with out-of-fold metrics |
sieve4b/ |
inference code: model, request contract, engine and HTTP server |
tokenizer.json, tokenizer_config.json |
Qwen3.5-4B's tokenizer, unchanged |
training_config.json, training_metrics.json, train.log |
the recipe, the run and its log |
result.json |
the Decision Index, calibration and latency numbers above |
provenance.json |
SHA-256 of every file and of the training data, and library versions |
License
- Downloads last month
- 53
Model tree for sthanika-ai/Sieve-4B
Collection including sthanika-ai/Sieve-4B
Evaluation results
- Decision Index 0.3, public index (chance-corrected; self-reported, not yet on the leaderboard) on Decision Index 0.3 public suite (140,178 scored requests; the index averages 37 benchmarks)self-reported50.060