Presets

Preconfigured presets and Docker-based build/run scripts for llama.cpp.

Hardware

Linux host has:

  • RTX 2080 SUPER: CUDA device 0, compute capability sm_75, ~8 GB VRAM
  • GTX 980: CUDA device 1, compute capability sm_52, ~4 GB VRAM
  • Proprietary NVIDIA driver 580
  • CUDA 12.9.1

llama.cpp CUDA builds target both GPUs with:

CMAKE_CUDA_ARCHITECTURES=52;75

The host keeps CUDA 13.4. The llama.cpp build uses CUDA 12.9 inside an Ubuntu 24.04 container because CUDA 12.9 is incompatible with host Ubuntu 26 headers.

Preset files

Preset file defines model configurations with server parameters, quantization, and speculative decoding settings:

File Purpose
preset.ini Standard llama.cpp presets

GPU settings in preset.ini:

device = CUDA0,CUDA1
split-mode = layer
main-gpu = 0
fit = on

Automatic fitting handles model VRAM overflow. layer split minimizes PCIe traffic; the GTX 980 is connected at PCIe x1.

Configuration

Scripts read .env. Copy template for another machine:

cp .env.example .env

Edit paths, image names, ports, CUDA architectures, and benchmark repetitions in .env. Keep .env local; commit .env.example only.

Key files:

File Purpose
.env Local machine configuration; ignored by Git
.env.example Portable configuration template

Docker setup

Host requirements:

  • Docker
  • NVIDIA Container Toolkit
  • Proprietary NVIDIA driver 580

Build image once after changing Dockerfile:

./create-build-image.sh

Files:

File Purpose
Dockerfile CUDA 12.9.1 build image with CMake, OpenSSL, GCC, and ccache
docker-compose.yml GPU-enabled llama.cpp build service
build-llama.cpp.sh Pull source and run containerized build

Unsloth Studio

docker-compose.unsloth.yml runs the official NVIDIA image with Unsloth Studio and JupyterLab. Configure the UNSLOTH_* values in .env, set UNSLOTH_STUDIO_PASSWORD and JUPYTER_PASSWORD, then start it with:

./unsloth.sh start

The script defaults to start; use ./unsloth.sh stop, ./unsloth.sh restart, ./unsloth.sh status, or ./unsloth.sh logs for container lifecycle management.

Studio is available at http://<server-host>:8000; JupyterLab is available at http://<server-host>:8888. The service persists the Hugging Face cache, Studio data, Triton kernels, and files under UNSLOTH_WORK_DIR. The published ports default to 0.0.0.0; set UNSLOTH_BIND_ADDRESS="127.0.0.1" when using SSH port forwarding instead of exposing the services to the LAN.

The image includes libssl-dev, so llama.cpp can download Hugging Face models over HTTPS.

Scripts

Build (Linux)

./create-build-image.sh   # First time or after Dockerfile changes
./build-llama.cpp.sh         # Build standard llama.cpp

Server

./server.sh                 # Start in the background with preset.ini
./server.sh start            # Same as above
./server.sh start custom.ini
./server.sh status
./server.sh restart custom.ini
./server.sh stop

start and daemon launch the Docker container in detached mode, so the script returns while the server continues running in the background. The container name is used for lifecycle tracking; no host PID file is needed. The legacy form ./server.sh custom.ini is also supported.

server.sh uses the prebuilt CUDA 12.9 image, exposes port 8080, and serves both GPUs. From another computer:

http://<host-ip>:8080

The server runs in router mode and loads models on demand. Set LLAMA_API_KEY before exposing the legacy server beyond a trusted LAN.

llama-swap gateway

llama-swap.sh runs the five text models and the Stable Diffusion image model from llama-swap.yaml, swapping the corresponding inference server on demand:

  • Qwen3.6-35B-MTP
  • Ornith-1.5-35B-A3B-MTP
  • gemma-4-12B-it-qat
  • gemma-4-26B-A4B-it-qat-MTP
  • Qwopus3.6-35B-A3B-Coder-MTP
  • Z-Image-Turbo

Z-Image-Turbo uses the locally built sd-server from STABLE_DIFFUSION_DIR, mounted read-only into the gateway container. It uses the same model paths and multi-GPU settings as stable.sh, and is placed in the exclusive single-model group so image generation unloads any active text model before using the GPUs. Build stable-diffusion.cpp first with ./build-from-source.sh stable-diffusion.cpp.

On the target server, from the presets directory configured in .env:

./llama-swap.sh start
./llama-swap.sh status
./llama-swap.sh restart
./llama-swap.sh stop

./llama-swap.sh and ./llama-swap.sh daemon are aliases for start. The gateway runs as a detached Docker container and remains active after the shell script exits.

The gateway is available at http://<server-host>:<port>/v1. Its API key is read from LLAMA_API_KEY; the key is passed to llama-swap only, while the upstream llama-server processes remain bound to the container loopback interface. Each model receives an available internal port from llama-swap, and the proxy uses that same assigned port.

Authentication is API-key based; there is no separate llama-swap username or password. The web UI and API endpoints require the key:

  • In a browser prompt, enter any non-empty username and use the value of LLAMA_API_KEY as the password.
  • For OpenAI-compatible clients, set the base URL to http://<server-host>:<port>/v1 and use the same value as the API key. The client should send it as Authorization: Bearer <API key>.
  • http://<server-host>:<port>/health is unauthenticated and returns OK.

The derived single-model presets are in llama-swap-presets/. The llama.cpp model logs are written to ${PRESETS_DIR}/llama-swap-logs/; ZImage logs remain attached to llama-swap's monitored upstream stream. Activity history and metrics are stored in ${PRESETS_DIR}/llama-swap-state/activity.sqlite, so the llama-swap Activity page retains previous requests and runs across restarts. The request/response capture opened from an Activity row is different: current llama-swap keeps that compressed capture in memory only, so the body of a /chat/completions request is lost when the gateway process restarts even though the Activity row remains. There is currently no supported captureDirectory or store.captures setting. Use client-side logging or a separate reverse proxy if the full request and response bodies must survive a restart.

The state directory is created automatically by llama-swap.sh and is mounted into the container through the existing /presets volume. The currently loaded model process is still intentionally restarted and must be loaded again after a container restart; only historical activity and metrics are persisted.

Models are automatically unloaded after 30 minutes of inactivity because globalTTL is set to 1800 seconds. Increase this value to reduce reloads, or set it to 0 to keep the active model loaded indefinitely.

The gateway is configured with an exclusive single-model group. The Pi agent selects the model by sending its exact model ID in each OpenAI-compatible request. When that ID changes, llama-swap unloads the current model before starting the requested one; the Pi agent must not call /models/sse.

To discover the available IDs, call GET /v1/models with the API key. A chat request should use the selected ID in its model field:

{
  "model": "Qwen3.6-35B-MTP",
  "messages": [
    {"role": "user", "content": "Hello"}
  ]
}

Quick Start

cd <presets-directory>
./create-build-image.sh
./build-llama.cpp.sh
./server.sh start
./server.sh status

Then use the OpenAI-compatible API on port 8080. Run server.sh and llama-swap.sh separately; they both use port 8080 by default. To use the gateway instead, stop the standalone server and run:

./server.sh stop
./llama-swap.sh start

Then use the authenticated gateway at http://<server-host>:<port>/v1.

For image generation, select Z-Image-Turbo as the model and use the OpenAI-compatible images endpoint:

{
  "model": "Z-Image-Turbo",
  "prompt": "a watercolor painting of a mountain cabin",
  "size": "1024x1024"
}

The Stable Diffusion WebUI-compatible endpoints are also available through the same gateway, including /sdapi/v1/txt2img and /sdapi/v1/img2img.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support