TitaNet-small (ONNX) โ browser speaker embeddings
ONNX export of NVIDIA NeMo TitaNet-small speaker-embedding model, packaged for in-browser speaker diarization via onnxruntime-web (WASM).
Used by the Silent Notetaker, an on-device meeting notetaker. Hosted here so the app can be served from static hosts with small per-file limits (e.g. Cloudflare Pages) and load the model from a CDN.
I/O
| Tensor | Shape | dtype | Notes |
|---|---|---|---|
input audio_signal |
[1, 80, T] |
float32 | 80-bin log-mel (slaney), 16 kHz |
input length |
[1] |
int64 | number of mel frames T |
output embs |
[1, 192] |
float32 | L2-normalizable speaker embedding |
The mel front-end is reimplemented in pure JavaScript and byte-validated against the reference Python (cosine similarity 1.000000). Compare embeddings with cosine similarity.
~40 MB, single-file. Loads with:
const ort = await import('https://cdn.jsdelivr.net/npm/onnxruntime-web/dist/ort.wasm.bundle.min.mjs');
const session = await ort.InferenceSession.create(
'https://huggingface.co/FluffyBunnies/titanet-small-onnx/resolve/main/titanet.onnx',
{ executionProviders: ['wasm'] });
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support