# Lightweight Python inference on CPU Use the 45.4 MB ONNX artifact without PyTorch, Transformers, scikit-learn, pickle, or an inference service. The Python helper downloads only the model, label list, and tokenizer. Runtime packages and tokenizer data are additional. ## Install and run Create an isolated Python 3.10+ environment, then obtain the short, reviewable helper and tested dependency list: ```sh python3 -m venv .venv # Linux/macOS; on Windows use .venv\Scripts\activate source .venv/bin/activate python -m pip install huggingface_hub==1.32.0 hf download Future-Labs/banking77-minilm onnx_predict.py requirements-onnx.txt --local-dir . python -m pip install -r requirements-onnx.txt python onnx_predict.py "How can I change my PIN?" ``` The model revision is pinned in the helper. The first call downloads the artifact; subsequent calls can use the Hugging Face cache. Classification runs on CPU. Scores are uncalibrated, and unrelated messages still get a banking intent. ```python from onnx_predict import ONNXIntentClassifier router = ONNXIntentClassifier(threads=2) results = router.predict([ "How can I change my PIN?", "How do I close my account?", ], top_k=3) for message_results in results: print(message_results) ``` A single string also returns a list containing one result list. Each ranked item has `label` and `score`. All 77 labels can be requested with `top_k=77`. Messages are processed individually to match the export's batch-1 evaluation. Create the classifier once and reuse it; initialization is not part of normal per-message inference. Long messages are truncated to 256 tokens. ## Offline deployment On a connected machine, fetch the three required artifact files: ```sh hf download Future-Labs/banking77-minilm \ onnx/model_fp16_storage.onnx onnx/labels.json encoder/tokenizer.json \ --revision dd6fc3794da468cdc8ad833aedb5286d1f761eaf \ --local-dir ./banking77_files ``` Copy that directory, the helper, and your installed dependencies to the target environment. Then use an explicit local path: ```python router = ONNXIntentClassifier("./banking77_files", local_files_only=True) print(router.predict("How do I close my account?")) ``` For an already populated Hugging Face cache, use `ONNXIntentClassifier(local_files_only=True)` or CLI `--offline`. Missing cached files fail instead of being fetched. Dependency installation and first model download still need connectivity unless separately pre-provisioned. ## Interpretation and validation The ONNX graph includes the encoder, pooling, normalization, and classification head. Large matrices are stored as FP16 and cast to FP32 for computation; run-time memory is not necessarily half the original. See [export evaluation](onnx/browser_evaluation.json) for full precision/compression comparisons and [helper verification](onnx/helper_verification.json) for the clean-environment checks, full test-prediction agreement, invalid inputs, public loading, and cached offline loading. The model has no out-of-domain detector or calibrated refusal threshold. A label must never authorize a banking action. Evaluate on representative, consented data and retain a human fallback. See the main model card for the benchmark methodology, overlap, failure cases, and licenses.