Instructions to use RedHatAI/DeepSeek-V4-Flash-0731-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use RedHatAI/DeepSeek-V4-Flash-0731-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="RedHatAI/DeepSeek-V4-Flash-0731-NVFP4") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("RedHatAI/DeepSeek-V4-Flash-0731-NVFP4") model = AutoModelForCausalLM.from_pretrained("RedHatAI/DeepSeek-V4-Flash-0731-NVFP4", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use RedHatAI/DeepSeek-V4-Flash-0731-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "RedHatAI/DeepSeek-V4-Flash-0731-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/DeepSeek-V4-Flash-0731-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/RedHatAI/DeepSeek-V4-Flash-0731-NVFP4
- SGLang
How to use RedHatAI/DeepSeek-V4-Flash-0731-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "RedHatAI/DeepSeek-V4-Flash-0731-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/DeepSeek-V4-Flash-0731-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "RedHatAI/DeepSeek-V4-Flash-0731-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "RedHatAI/DeepSeek-V4-Flash-0731-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use RedHatAI/DeepSeek-V4-Flash-0731-NVFP4 with Docker Model Runner:
docker model run hf.co/RedHatAI/DeepSeek-V4-Flash-0731-NVFP4
DeepSeek-V4-Flash-0731-NVFP4-FP8-BLOCK
Model Overview
Model Architecture: DeepseekV4ForCausalLM
- Input: Text
- Output: Text
Model Optimizations:
- Routed expert weight quantization: NVFP4
- Routed expert activation quantization: NVFP4
- Attention/shared expert weight quantization: FP8 block quantization
- Attention/shared expert activation quantization: FP8
Release Date: 2026-09-23
Version: 1.0
Model Developers: RedHatAI
This model is a mixed-precision quantized version of deepseek-ai/DeepSeek-V4-Flash-0731.
It was evaluated on GPQA to assess its quality in comparison to the base model.
Model Optimizations
This model uses mixed-precision NVFP4 and FP8 quantization for efficient inference with vLLM.
The routed MoE expert linear layers use NVFP4 weights and activations with a group size of 16. Selected attention projections, the attention indexer q_b_proj, and shared expert layers use FP8 weights with 128x128 block quantization and dynamically quantized FP8 activations.
Layers outside these quantization targets retain their original precision.
The resulting checkpoint is exported using compressed-tensors and is ready for inference with vLLM.
The NVFP4 expert quantization was applied using LLM Compressor.
Deployment
vLLM Serving
vllm serve /data/kylesayrs/hub/RedHatAI/DeepSeek-V4-Flash-0731-NVFP4-FP8-BLOCK \
--tensor-parallel-size 8 \
--enable-expert-parallel \
--kv-cache-dtype fp8 \
--served-model-name dsv4_nvfp4_fp8_gptq \
--reasoning-parser deepseek_v4 \
--enable-auto-tool-choice \
--tool-call-parser deepseek_v4
Creation
This model was created using LLM Compressor with mixed-precision NVFP4 and FP8 quantization.
Routed MoE experts use NVFP4 weight and activation quantization, while selected attention, indexer, and shared expert linear operators use FP8 quantization. The checkpoint is exported in the compressed-tensors format for optimized deployment with vLLM.
Evaluations
| Model | GPQA |
|---|---|
| deepseek-ai/DeepSeek-V4-Flash-0731 | 84.34% |
| RedHatAI/DeepSeek-V4-Flash-0731-NVFP4-FP8-BLOCK | 86.15% |
Note that GPQA is a non-deterministic eval on its own and quantization can sometimes act as a regularizer for individual evals.
- Downloads last month
- 116
Model tree for RedHatAI/DeepSeek-V4-Flash-0731-NVFP4
Base model
deepseek-ai/DeepSeek-V4-Flash-0731