TitaNet-small (ONNX) โ€” browser speaker embeddings

ONNX export of NVIDIA NeMo TitaNet-small speaker-embedding model, packaged for in-browser speaker diarization via onnxruntime-web (WASM).

Used by the Silent Notetaker, an on-device meeting notetaker. Hosted here so the app can be served from static hosts with small per-file limits (e.g. Cloudflare Pages) and load the model from a CDN.

I/O

Tensor Shape dtype Notes
input audio_signal [1, 80, T] float32 80-bin log-mel (slaney), 16 kHz
input length [1] int64 number of mel frames T
output embs [1, 192] float32 L2-normalizable speaker embedding

The mel front-end is reimplemented in pure JavaScript and byte-validated against the reference Python (cosine similarity 1.000000). Compare embeddings with cosine similarity.

~40 MB, single-file. Loads with:

const ort = await import('https://cdn.jsdelivr.net/npm/onnxruntime-web/dist/ort.wasm.bundle.min.mjs');
const session = await ort.InferenceSession.create(
  'https://huggingface.co/FluffyBunnies/titanet-small-onnx/resolve/main/titanet.onnx',
  { executionProviders: ['wasm'] });
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support