Sieve-4B

Sieve-4B is a decision model from the Sieve family. You give it a piece of text or JSON (the state) and typed questions, and it returns a calibrated probability for every option of every question. It reads the state once, scores the options directly, and never generates text.

It is a LoRA adapter (rank 32) and a 255-way answer head on Qwen/Qwen3.5-4B (revision 851bf6e8), with one calibration temperature per question type. Its readout differs from the pointer head of Sieve-27B and Sieve-9B-Plus, so it ships its own inference code (sieve4b/, below).

Earlier weights. This revision replaces the first Sieve-4B weights (same architecture and recipe, retrained on newer data). They remain available at revision e49deceb (Decision Index 0.3 public index 43.69).

How it works

  • Backbone. The post-trained Qwen3.5-4B, a hybrid of Gated DeltaNet and full-attention layers. Only its text part is used; the model cannot produce text.
  • Adapter. LoRA r=32, alpha 64, on every linear layer: attention, MLP and the Gated DeltaNet projections.
  • Answer head. Each option is listed under a code (A, B, ..., AA, ...). A 255-way linear head, initialised from the base model's own output rows for those codes, scores all codes at the question's decision point. The head reads the whole option list.
  • Isolation. The state is encoded once; each question runs as its own branch from that cache, so questions never see each other.
  • Yes/no in both orders. A yes/no question is asked with the options in both orders and the answers are averaged.
type options returns
choice the caller's keys, optionally described (up to 255) a probability per key
noul yes / no P(yes)
score ordered levels (up to 10) a probability per level

Use

hf download sthanika-ai/Sieve-4B --local-dir Sieve-4B && cd Sieve-4B
pip install -r requirements.txt            # torch, transformers>=5.17, peft>=0.21, safetensors, pydantic
pip install flash-linear-attention==0.5.2  # recommended on GPU: the Gated DeltaNet kernels our runs used
from sieve4b import load, decide

engine = load(".", device="cuda:0")        # this folder; fetches Qwen/Qwen3.5-4B at its pinned revision on first use
decide(engine, {"subject": "Duplicate charge on invoice 4411", "body": "Billed twice. Refund today or we cancel."},
       {"department": {"type": "choice", "instructions": "Which team handles this?",
                        "criteria": {"billing": "invoices, payments, refunds", "technical": "bugs and outages"}},
        "churn_risk": {"type": "noul", "instructions": "Does the customer threaten to cancel?"}})
  • Server. python -m sieve4b.server --ckpt . --port 8123 --name sieve-4b serves the same thing as a /v1/systemone HTTP endpoint (needs fastapi and uvicorn). The Decision Index run below used it.
  • GPU memory. About 10 GB in bf16. Longer states need more.
  • Token limits. Up to 32,768 tokens for the state plus the longest question, 65,536 per request, 255 choice options and 10 score levels. Longer requests are refused, never truncated. Training records reached 16,806 tokens.
  • Temperatures. sieve_config.json stores one temperature per question type: choice 1.161, yes/no 1.440 and score 1.000. A temperature never changes an answer.

Results

Decision Index 0.3

Sieve-4B scores 50.06 (raw index 62.10, breadth 48.32) on the public 0.3 suite of the Decision Index, on our own run with the official kit at 9eb2dbe (0.3) and its scorer.

  • Coverage. All 140,178 scoreable requests over 43 benchmarks were answered, with no truncation, no unsupported request and no error. The index averages 37 of the benchmarks.
  • How the run was made. A complete run of the 0.2.1 suite, extended to 0.3 with the 2,638 rebuilt GSM8K requests, on two A100 80GB PCIe GPUs with three server processes per GPU.
  • Results. sthanika-ai/Sieve-4B-decision-index-results.
  • Not on the leaderboard yet. The score is self-reported. On the 0.3 board the ranking uses the Full score, where this public suite counts 20% and the maintainers' private tests count 80%.
Decision Index 0.3 Knowledge & Reasoning Language Understanding Retrieval & Classification Tools & Automation Arts & Human Taste
50.06 33.9 53.9 61.6 66.8 28.3

The index and area scores are chance-corrected: 0 is random guessing and 100 is perfect. Below, the score is in each benchmark's own metric and the skill is chance-corrected. On edition 0.2.1 the same run scores 49.21.

All 43 benchmarks
benchmark metric score skill (×100) requests
ACOS per-review F1 0.206 18.1 1,565
Amazon ESCI macro-F1 0.491 36.2 5,000
ANLI macro-F1 0.635 45.3 3,200
API-Bank accuracy 0.738 73.3 508
BANKING77 macro-F1 0.828 82.6 3,080
BBH fixed-option tasks accuracy 0.682 53.9 5,507
BFCL case exact accuracy 0.944 92.4 1,694
BPoMP accuracy 0.804 61.0 5,000
BRIGHT nDCG@10 0.444 37.1 220
cfcolor accuracy 0.613 17.8 5,000
ChessBench accuracy 0.138 6.1 5,000
CLadder accuracy 0.639 27.8 5,000
CLINC150+OOS macro-F1 0.902 90.1 5,500
ContractNLI macro-F1 0.798 70.8 123
CRUXEval accuracy 0.558 29.9 570
FinEntity macro-F1 0.886 83.2 979
GPQA Diamond accuracy 0.423 23.1 196
GSM8K accuracy 0.645 57.4 2,638
Habermas Machine accuracy 0.400 12.9 1,676
HellaSwag accuracy 0.950 93.4 10,042
HLE accuracy 0.094 0.0 501
Home appliance simulator case exact accuracy 0.261 26.1 88
HoVer claim verification accuracy 0.750 50.0 4,000
Humicroedit accuracy 0.595 18.9 2,628
iSarcasmEval Sarcasm F1 · track A, English 0.461 30.6 4,600
MMLU-Pro accuracy 0.550 49.4 12,032
MuSR accuracy 0.550 28.5 752
New Yorker caption matching accuracy 0.616 51.9 528
NLI4CT macro-F1 0.768 54.9 5,500
PhishNChips phishing decisions accuracy 0.843 68.5 2,000
POP909-CL accuracy 0.079 7.2 2,000
RAGTruth response-level hallucination F1 on hallucinated class 0.779 54.1 2,700
SATA-Bench case exact accuracy 0.206 19.6 1,650
ToolRet nDCG@10 0.673 62.3 685
VAST macro-F1 0.524 28.6 3,006
When2Call MCQ accuracy 0.799 73.2 3,652
WinoGrande accuracy 0.858 71.6 1,267
ARC-Challenge (not in the index) accuracy 0.944 92.5 1,172
ARC-Easy (not in the index) accuracy 0.979 97.1 2,376
MMLU (not in the index) accuracy 0.757 67.7 14,033
RouterBench (not in the index) selected quality (quality objective) 0.790 50.8 10,000
SGD/SGD-X (not in the index) macro-F1 0.382 0.0 2,500
SimpleBench (not in the index) accuracy 0.100 0.0 10

Calibration

choice yes/no
served temperature 1.161 1.440
ECE on held-out sources, out of fold 0.039 0.030
  • Fit. The served temperatures were fitted by NLL on 1,456 held-out records from 12 public sources that are not in the training data, and reported 2-fold out of fold. A fit on the in-domain calibration split gives lower temperatures (choice 0.914, yes/no 0.862); both reports are in calibration/.
  • On the Decision Index. Over the 0.2.1 rows, with our reimplementation of the board's calibration panel (benchmarks with a per-field gold answer, equally weighted; confidence = the probability of the chosen option): ECE 0.019, and 0.5% of answers wrong with confidence of 0.95 or more.

Latency

HTTP wall time as recorded in the Decision Index run, with three server processes sharing each A100 80GB PCIe: median 302 ms, mean 433 ms, 80th percentile 400 ms, 95th percentile 817 ms.

Training

  • Recipe.
    • LoRA r=32, α=64, dropout 0.05 on the attention, MLP and Gated DeltaNet projections, and the 255-way answer head.
    • AdamW, lr 5e-5 for the adapter and 1e-4 for the head, 3% warm-up then cosine, 1 epoch, 128 records per step, bf16.
    • Loss: cross-entropy on the target distribution, an option-order consistency term and a small replay term that keeps the model near its base.
    • 1,295 steps, 11.6 hours on two A100 80GB PCIe GPUs. No record was skipped or truncated.
  • Selection. The last checkpoint. The benchmarks were never used to choose or tune anything.
  • Data. 165,832 training records.
    • Sources: public datasets and generated decision records.
    • Not published. The training files are not part of this release; the SHA-256 of each file is in training_config.json.

Files

file contents
adapter_model.safetensors, adapter_config.json the LoRA adapter (rank 32, PEFT format)
head.safetensors the 255-way answer head
sieve_config.json option codes, prompt preamble and the served temperatures
calibration/ the two temperature fits, with out-of-fold metrics
sieve4b/ inference code: model, request contract, engine and HTTP server
tokenizer.json, tokenizer_config.json Qwen3.5-4B's tokenizer, unchanged
training_config.json, training_metrics.json, train.log the recipe, the run and its log
result.json the Decision Index, calibration and latency numbers above
provenance.json SHA-256 of every file and of the training data, and library versions

License

Apache-2.0, see LICENSE and NOTICE.

Downloads last month
53
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sthanika-ai/Sieve-4B

Finetuned
Qwen/Qwen3.5-4B
Adapter
(730)
this model

Collection including sthanika-ai/Sieve-4B

Evaluation results

  • Decision Index 0.3, public index (chance-corrected; self-reported, not yet on the leaderboard) on Decision Index 0.3 public suite (140,178 scored requests; the index averages 37 benchmarks)
    self-reported
    50.060