banking77-minilm / CPU_QUICKSTART.md
Future-Labs's picture
Add verified ONNX-only Python loader and offline deployment quickstart
37f5cd0 verified
|
Raw History Blame Contribute Delete
3.29 kB

Lightweight Python inference on CPU

Use the 45.4 MB ONNX artifact without PyTorch, Transformers, scikit-learn, pickle, or an inference service. The Python helper downloads only the model, label list, and tokenizer. Runtime packages and tokenizer data are additional.

Install and run

Create an isolated Python 3.10+ environment, then obtain the short, reviewable helper and tested dependency list:

python3 -m venv .venv
# Linux/macOS; on Windows use .venv\Scripts\activate
source .venv/bin/activate
python -m pip install huggingface_hub==1.32.0
hf download Future-Labs/banking77-minilm onnx_predict.py requirements-onnx.txt --local-dir .
python -m pip install -r requirements-onnx.txt
python onnx_predict.py "How can I change my PIN?"

The model revision is pinned in the helper. The first call downloads the artifact; subsequent calls can use the Hugging Face cache. Classification runs on CPU. Scores are uncalibrated, and unrelated messages still get a banking intent.

from onnx_predict import ONNXIntentClassifier

router = ONNXIntentClassifier(threads=2)
results = router.predict([
    "How can I change my PIN?",
    "How do I close my account?",
], top_k=3)
for message_results in results:
    print(message_results)

A single string also returns a list containing one result list. Each ranked item has label and score. All 77 labels can be requested with top_k=77. Messages are processed individually to match the export's batch-1 evaluation. Create the classifier once and reuse it; initialization is not part of normal per-message inference. Long messages are truncated to 256 tokens.

Offline deployment

On a connected machine, fetch the three required artifact files:

hf download Future-Labs/banking77-minilm \
  onnx/model_fp16_storage.onnx onnx/labels.json encoder/tokenizer.json \
  --revision dd6fc3794da468cdc8ad833aedb5286d1f761eaf \
  --local-dir ./banking77_files

Copy that directory, the helper, and your installed dependencies to the target environment. Then use an explicit local path:

router = ONNXIntentClassifier("./banking77_files", local_files_only=True)
print(router.predict("How do I close my account?"))

For an already populated Hugging Face cache, use ONNXIntentClassifier(local_files_only=True) or CLI --offline. Missing cached files fail instead of being fetched. Dependency installation and first model download still need connectivity unless separately pre-provisioned.

Interpretation and validation

The ONNX graph includes the encoder, pooling, normalization, and classification head. Large matrices are stored as FP16 and cast to FP32 for computation; run-time memory is not necessarily half the original. See export evaluation for full precision/compression comparisons and helper verification for the clean-environment checks, full test-prediction agreement, invalid inputs, public loading, and cached offline loading.

The model has no out-of-domain detector or calibrated refusal threshold. A label must never authorize a banking action. Evaluate on representative, consented data and retain a human fallback. See the main model card for the benchmark methodology, overlap, failure cases, and licenses.