Download runtime/SETUP.md from s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4: direct link, hf CLI and curl.
- Browser
- Download file 2.26 kB
-
https://huggingface.co/s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4/resolve/main/runtime/SETUP.md
- Command line
-
hf download hf://s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4/runtime/SETUP.md
-
curl -L -o SETUP.md https://huggingface.co/s-zaizen/DeepSeek-V4-Flash-Vision-Exp-NVFP4/resolve/main/runtime/SETUP.md
Build and run
Requires two DGX Spark hosts with working Docker GPU support and inter-node RDMA.
Build on each host with bash runtime/build.sh. This uses pinned public Mia sources,
the pinned Anemll base image, and the compatibility patches included here.
The standalone image was built successfully and its 3,258 vLLM source/binary files
matched the measured runtime byte-for-byte. Image IDs differ because build layers differ.
Download this model into a regular directory on each host. Separately download
deepseek-ai/DeepSeek-V4-Flash-Vision-Exp at revision
86f746b36186f0e567729a5c06a8c918caba82a9 for the original draft checkpoint.
Use hf download --local-dir rather than mounting symlink-based cache snapshots.
Both target and source directories must contain their full 48-shard checkpoints.
Install PyYAML for the launcher (python3 -m pip install pyyaml in a virtual environment).
Create a writable cache directory. Checkpoints must be directly owned by UID 1000,
and the cache writable by UID/GID 1000:1000, matching the container user.
The opt-in UMA helper requires regular owned shard files, not symlinks.
Start rank 1 on the second host, then rank 0 on the first host:
python3 runtime/launch.py \
--target /absolute/path/to/target \
--draft /absolute/path/to/original-draft \
--cache /absolute/path/to/cache \
--node-rank 1 \
--master-addr HEAD_RDMA_IP \
--host-ip THIS_NODE_RDMA_IP \
--interface RDMA_NETWORK_INTERFACE \
--hca RDMA_HCA
Use --node-rank 0 on the head. Replace the example paths and network identifiers;
the same head address and master port must be used on both hosts. The recipe's network
interface values are overridden by these arguments. Check GID index 3 matches your fabric.
Use --dry-run to inspect the Docker invocation without starting it.
The API listens on port 8000 with no authentication in this example. Use a trusted, firewalled network; do not expose the endpoint publicly. Ctrl-C stops the foreground container. Stop both ranks when finished. No existing services are stopped by the launcher.
KV budget is 6 GiB per node. The 1M context ceiling is not a claim that full 1M inputs were tested. Cold startup includes compilation. Avoid concurrent large-model workloads.