Recipe
2× DGX Spark · vLLM · TP=2
- NVFP4 weights, FP8 KV cache: 6 GiB/node
- CUTLASS target / B12x draft
- DSpark: 6 speculative tokens, probabilistic
- Batch limit: 4096 · Max sequences: 4
- Context limit: 1M · Reasoning: max
Recipe · Runtime build and setup
Results
43.82 ± 1.54 tok/s aggregate decode
25.98 ± 2.80 tok/s per request
llama-benchy 0.4.0 · 16K prefix + 2K prompt · concurrency 2 · 128-token output limit · 3 repetitions. Mean ± standard deviation; no OOM kills. Short speed test only; quality and 1M-input performance unverified.
Credits
- Downloads last month
- 382
Model tree for s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4
Base model
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp