s-zaizen's picture
Update NVFP4 weights, standalone DSpark recipe and benchmark results
325721d verified
|
Raw History Blame Contribute Delete
2.26 kB

Build and run

Requires two DGX Spark hosts with working Docker GPU support and inter-node RDMA. Build on each host with bash runtime/build.sh. This uses pinned public Mia sources, the pinned Anemll base image, and the compatibility patches included here. The standalone image was built successfully and its 3,258 vLLM source/binary files matched the measured runtime byte-for-byte. Image IDs differ because build layers differ.

Download this model into a regular directory on each host. Separately download deepseek-ai/DeepSeek-V4-Flash-Vision-Exp at revision 86f746b36186f0e567729a5c06a8c918caba82a9 for the original draft checkpoint. Use hf download --local-dir rather than mounting symlink-based cache snapshots. Both target and source directories must contain their full 48-shard checkpoints.

Install PyYAML for the launcher (python3 -m pip install pyyaml in a virtual environment). Create a writable cache directory. Checkpoints must be directly owned by UID 1000, and the cache writable by UID/GID 1000:1000, matching the container user. The opt-in UMA helper requires regular owned shard files, not symlinks.

Start rank 1 on the second host, then rank 0 on the first host:

python3 runtime/launch.py \
  --target /absolute/path/to/target \
  --draft /absolute/path/to/original-draft \
  --cache /absolute/path/to/cache \
  --node-rank 1 \
  --master-addr HEAD_RDMA_IP \
  --host-ip THIS_NODE_RDMA_IP \
  --interface RDMA_NETWORK_INTERFACE \
  --hca RDMA_HCA

Use --node-rank 0 on the head. Replace the example paths and network identifiers; the same head address and master port must be used on both hosts. The recipe's network interface values are overridden by these arguments. Check GID index 3 matches your fabric. Use --dry-run to inspect the Docker invocation without starting it.

The API listens on port 8000 with no authentication in this example. Use a trusted, firewalled network; do not expose the endpoint publicly. Ctrl-C stops the foreground container. Stop both ranks when finished. No existing services are stopped by the launcher.

KV budget is 6 GiB per node. The 1M context ceiling is not a claim that full 1M inputs were tested. Cold startup includes compilation. Avoid concurrent large-model workloads.