# Build and run Requires two DGX Spark hosts with working Docker GPU support and inter-node RDMA. Build on each host with `bash runtime/build.sh`. This uses pinned public Mia sources, the pinned Anemll base image, and the compatibility patches included here. The standalone image was built successfully and its 3,258 vLLM source/binary files matched the measured runtime byte-for-byte. Image IDs differ because build layers differ. Download this model into a regular directory on each host. Separately download `deepseek-ai/DeepSeek-V4-Flash-Vision-Exp` at revision `86f746b36186f0e567729a5c06a8c918caba82a9` for the original draft checkpoint. Use `hf download --local-dir` rather than mounting symlink-based cache snapshots. Both target and source directories must contain their full 48-shard checkpoints. Install PyYAML for the launcher (`python3 -m pip install pyyaml` in a virtual environment). Create a writable cache directory. Checkpoints must be directly owned by UID 1000, and the cache writable by UID/GID 1000:1000, matching the container user. The opt-in UMA helper requires regular owned shard files, not symlinks. Start rank 1 on the second host, then rank 0 on the first host: ```bash python3 runtime/launch.py \ --target /absolute/path/to/target \ --draft /absolute/path/to/original-draft \ --cache /absolute/path/to/cache \ --node-rank 1 \ --master-addr HEAD_RDMA_IP \ --host-ip THIS_NODE_RDMA_IP \ --interface RDMA_NETWORK_INTERFACE \ --hca RDMA_HCA ``` Use `--node-rank 0` on the head. Replace the example paths and network identifiers; the same head address and master port must be used on both hosts. The recipe's network interface values are overridden by these arguments. Check GID index 3 matches your fabric. Use `--dry-run` to inspect the Docker invocation without starting it. The API listens on port 8000 with no authentication in this example. Use a trusted, firewalled network; do not expose the endpoint publicly. Ctrl-C stops the foreground container. Stop both ranks when finished. No existing services are stopped by the launcher. KV budget is 6 GiB per node. The 1M context ceiling is not a claim that full 1M inputs were tested. Cold startup includes compilation. Avoid concurrent large-model workloads.