Instructions to use NomaDamas/v-splade-efficient-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use NomaDamas/v-splade-efficient-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir v-splade-efficient-mlx NomaDamas/v-splade-efficient-mlx
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
v-splade-efficient-mlx
MLX (bfloat16) conversion of naver/v-splade-efficient
(V-SPLADE, arXiv:2605.30917) for Apple Silicon,
produced by NomaDamas/SPLADE-mlx.
V-SPLADE is an inference-free sparse retriever for visual document retrieval: document pages (rendered PDFs, slides, scans) are encoded by a ModernVBERT backbone (SigLIP vision tower + pixel-shuffle connector + ModernBERT text encoder) with a SPLADE MLM head into a 50,368-dim vocabulary-space sparse vector, while queries are resolved by a learned Bag-of-Words lookup with no neural encoding at all.
Contents: weights.safetensors (document encoder, bfloat16),
query_lookup.npy (inference-free query table, fp32), config.json,
plus tokenizer/processor configs for self-contained loading.
Changes from upstream: PyTorch checkpoint converted to MLX safetensors
(parameter re-mapping, conv weight transposed to NHWC, cast to bfloat16); the
query lookup table softplus(embedding @ projection + bias) is precomputed
with special tokens zeroed. No training or fine-tuning was performed.
Quality (see repo REPORT.md for methodology):
- Separate fp32 conversion parity vs the PyTorch reference: max |logit delta| 1.5e-04 on real document-page inputs, sparse-vector cosine 1.000000, top-64 term overlap 100%; the query table matches the shipped Sentence Transformers static embedding to 1.2e-07.
- ViDoRe
docvqa_test_subsamplednDCG@5 (fp32): 0.4098 (torch) -> 0.4098 (MLX), delta +0.0000 (gate: ±0.002).
This repository itself stores bfloat16 document-encoder weights. The fp32 numbers above describe a separate fp32 conversion and must not be attributed to this linked bfloat16 artifact.
Usage
from splade_mlx.convert_vsplade import load_vsplade
import mlx.core as mx
from PIL import Image
model, query_encoder, processor = load_vsplade("NomaDamas/v-splade-efficient-mlx")
# documents (page images)
enc = processor(text=["User:<image><end_of_utterance>\nAssistant:"],
images=[[Image.open("page.png")]], return_tensors="np")
d = model.encode(mx.array(enc["input_ids"]), mx.array(enc["attention_mask"]),
enc["pixel_values"]) # (1, 50368)
# queries: inference-free lookup, no neural network
q = processor.tokenizer(["total revenue 2023"], return_tensors="np")
qw = query_encoder.encode(q["input_ids"], q["attention_mask"]) # (1, 50368)
score = d @ qw.T
License
Apache-2.0, same as the upstream checkpoint (© NAVER Corp). This repository is not affiliated with or endorsed by NAVER.
- Downloads last month
- 31
Quantized
Model tree for NomaDamas/v-splade-efficient-mlx
Base model
naver/v-splade-efficient