Recipe

2× DGX Spark · vLLM · TP=2

  • NVFP4 weights, FP8 KV cache: 6 GiB/node
  • CUTLASS target / B12x draft
  • DSpark: 6 speculative tokens, probabilistic
  • Batch limit: 4096 · Max sequences: 4
  • Context limit: 1M · Reasoning: max

Recipe · Runtime build and setup

Results

43.82 ± 1.54 tok/s aggregate decode
25.98 ± 2.80 tok/s per request

llama-benchy 0.4.0 · 16K prefix + 2K prompt · concurrency 2 · 128-token output limit · 3 repetitions. Mean ± standard deviation; no OOM kills. Short speed test only; quality and 1M-input performance unverified.

Credits

DeepSeek · MiaAI-Lab · Anemll · vLLM · FlashInfer

Downloads last month
382
Safetensors
Model size
305B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
U8
·
I64
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4

Quantized
(25)
this model