Pulse Decide 150M
RouterML's on-device decision model. 149.3M parameters (~570MB fp32). Non-generative: it scores typed decisions rather than generating text, which is why it runs on a laptop instead of a server.
Makes typed decisions โ choice, boolean, multi-label, ordinal โ over arbitrary runtime schemas, offline, on-device.
Status: v0.1 research preview. This is an early research model, not a production release.
One customer message, routed twice. Without governed context Pulse isn't confident enough, so it escalates rather than guessing. With governed context it decides โ 99.7% confident, on-device, in ~32ms.
What it does
Given a piece of text (the state) and a set of candidate labels supplied at runtime, Pulse scores each candidate and returns a decision plus a confidence. The label set does not need to be seen during training โ that is the "arbitrary schema" property.
Key facts
| Parameters | 149.3M (0.149B) |
| Encoder | answerdotai/ModernBERT-base |
| Architecture | cross-encoder decision scorer (non-generative) |
| Size on disk | ~570MB (fp32); ~150MB at int8 |
| Runs on | CPU / Apple MPS / CUDA โ offline, no data leaves the device |
| Decision types | choice, boolean, multi-label, ordinal |
Performance
Pulse is designed to run behind a confidence gate โ answering what it is confident about and escalating the rest. That selective mode is how it is intended to be deployed, so it is what we report here. Measured on held-out, schema-disjoint evaluation sets (label sets never seen in training) at confidence threshold 0.5:
| evaluation set | answering everything | confident subset only |
|---|---|---|
| MASSIVE (intent, 60 labels) | 47.6% | 71.1% at 47% coverage |
| macro (3 held-out sets) | ~51% | 59.5% at 63% coverage |
Accuracy on accepted decisions is meaningfully higher than on forced answers โ at the cost of escalating the remainder. This is the intended operating mode.
A same-split benchmark evaluation against other published decision models is forthcoming. We make no claim of matching or exceeding any other model at this time.
Intended use
- Research on small, non-generative decision models and governed routing
- On-device / offline decisioning where data cannot leave the machine
- As the local tier of an escalation system (cheap local decision โ escalate when unsure)
Limitations
- Multi-label is weak (~8% exact-set on FD multi-label heads). The model often identifies the right labels but selects the wrong number of them. Active work.
- Confidence calibration drifts between training runs; treat thresholds as needing per-deployment calibration rather than as universal constants.
- Not evaluated for fairness, safety, or high-stakes domains. Do not use for consequential decisions without your own evaluation.
- English only.
- Early research preview โ interfaces and weights may change.
Training data and licensing โ ๏ธ
Trained on a derived corpus assembled from 40 public text-classification datasets (contamination-audited against all evaluation sets; near-duplicate filtered).
The source datasets carry mixed licenses, including:
- ~24 permissive (Apache-2.0, CC-BY, MIT-like)
- several research/non-commercial-only datasets
- several derived from Twitter/X content, whose Terms of Service restrict redistribution
We publish the model weights, not the training data. Because the corpus mixes research-only and ToS-restricted sources, users are responsible for verifying that their intended use complies with the underlying dataset licenses. A cleanly-licensed (permissive-only) release is planned.
Usage
The repo is self-contained: pulse_model.py is the full model definition (113 lines, only
torch + transformers required).
pip install torch transformers
import torch
from pulse_model import PulseConfig, PulseUnified, load_tokenizer
blob = torch.load("pulse_decide_150m.pt", map_location="cpu", weights_only=True)
cfg = PulseConfig(**blob["cfg"])
tok = load_tokenizer(cfg)
model = PulseUnified(cfg)
model.load_state_dict(blob["state_dict"])
model.eval()
def decide(state, question, choices, context=None):
"""Score every choice and return (label, confidence). Options are read at RUNTIME,
so label sets never seen in training still work."""
ctx = " | ".join(context) + " || " if context else ""
left = [f"{ctx}{state} [SEP] {question}"] * len(choices)
enc = tok(left, choices, truncation=True, max_length=cfg.max_length,
padding=True, return_tensors="pt")
with torch.no_grad():
scores = model.score_pairs(enc["input_ids"], enc["attention_mask"])
probs = torch.softmax(scores, dim=0)
i = int(probs.argmax())
return choices[i], float(probs[i])
label, confidence = decide(
state="We were invoiced twice this month and the NET-60 terms were not applied.",
question="Which team should handle this?",
choices=["sales", "support", "billing", "success", "compliance"],
context=["NET-60 payment terms", "prior billing dispute on file"],
)
print(label, f"{confidence:.1%}") # -> billing 99.7%
# Governance: act only when confident, otherwise escalate.
if confidence < 0.60:
print("escalate โ not confident enough to act")
Download the two files you need:
from huggingface_hub import hf_hub_download
repo = "RouterML/pulse-decide-150m"
ckpt = hf_hub_download(repo, "pulse_decide_150m.pt")
code = hf_hub_download(repo, "pulse_model.py")
Model family
Pulse Decide is the decision model inside RouterML, which adds the governance layer โ eligibility, confidence gating, escalation and verification โ around it.
Citation
@software{pulse_decide_150m,
title = {Pulse Decide 150M: an on-device non-generative decision model},
author = {RouterML},
year = {2026},
note = {v0.1 research preview}
}
- Downloads last month
- 17
