Presets
Preconfigured presets and Docker-based build/run scripts for llama.cpp.
Hardware
Linux host has:
- RTX 2080 SUPER: CUDA device 0, compute capability
sm_75, ~8 GB VRAM - GTX 980: CUDA device 1, compute capability
sm_52, ~4 GB VRAM - Proprietary NVIDIA driver 580
- CUDA 12.9.1
llama.cpp CUDA builds target both GPUs with:
CMAKE_CUDA_ARCHITECTURES=52;75
The host keeps CUDA 13.4. The llama.cpp build uses CUDA 12.9 inside an Ubuntu 24.04 container because CUDA 12.9 is incompatible with host Ubuntu 26 headers.
Preset files
Preset file defines model configurations with server parameters, quantization, and speculative decoding settings:
| File | Purpose |
|---|---|
preset.ini |
Standard llama.cpp presets |
GPU settings in preset.ini:
device = CUDA0,CUDA1
split-mode = layer
main-gpu = 0
fit = on
Automatic fitting handles model VRAM overflow. layer split minimizes PCIe traffic; the GTX 980 is connected at PCIe x1.
Configuration
Scripts read .env. Copy template for another machine:
cp .env.example .env
Edit paths, image names, ports, CUDA architectures, and benchmark repetitions in .env. Keep .env local; commit .env.example only.
Key files:
| File | Purpose |
|---|---|
.env |
Local machine configuration; ignored by Git |
.env.example |
Portable configuration template |
Docker setup
Host requirements:
- Docker
- NVIDIA Container Toolkit
- Proprietary NVIDIA driver 580
Build image once after changing Dockerfile:
./create-build-image.sh
Files:
| File | Purpose |
|---|---|
Dockerfile |
CUDA 12.9.1 build image with CMake, OpenSSL, GCC, and ccache |
docker-compose.yml |
GPU-enabled llama.cpp build service |
build-llama.cpp.sh |
Pull source and run containerized build |
Unsloth Studio
docker-compose.unsloth.yml runs the official NVIDIA image with Unsloth Studio
and JupyterLab. Configure the UNSLOTH_* values in .env, set
UNSLOTH_STUDIO_PASSWORD and JUPYTER_PASSWORD, then start it with:
./unsloth.sh start
The script defaults to start; use ./unsloth.sh stop,
./unsloth.sh restart, ./unsloth.sh status, or ./unsloth.sh logs for
container lifecycle management.
Studio is available at http://<server-host>:8000; JupyterLab is available at
http://<server-host>:8888. The service persists the Hugging Face cache,
Studio data, Triton kernels, and files under UNSLOTH_WORK_DIR. The published
ports default to 0.0.0.0; set UNSLOTH_BIND_ADDRESS="127.0.0.1" when using
SSH port forwarding instead of exposing the services to the LAN.
The image includes libssl-dev, so llama.cpp can download Hugging Face models over HTTPS.
Scripts
Build (Linux)
./create-build-image.sh # First time or after Dockerfile changes
./build-llama.cpp.sh # Build standard llama.cpp
Server
./server.sh # Start in the background with preset.ini
./server.sh start # Same as above
./server.sh start custom.ini
./server.sh status
./server.sh restart custom.ini
./server.sh stop
start and daemon launch the Docker container in detached mode, so the
script returns while the server continues running in the background. The
container name is used for lifecycle tracking; no host PID file is needed.
The legacy form ./server.sh custom.ini is also supported.
server.sh uses the prebuilt CUDA 12.9 image, exposes port 8080, and serves both GPUs. From another computer:
http://<host-ip>:8080
The server runs in router mode and loads models on demand. Set LLAMA_API_KEY
before exposing the legacy server beyond a trusted LAN.
llama-swap gateway
llama-swap.sh runs the five text models and the Stable Diffusion image model
from llama-swap.yaml, swapping the corresponding inference server on demand:
Qwen3.6-35B-MTPOrnith-1.5-35B-A3B-MTPgemma-4-12B-it-qatgemma-4-26B-A4B-it-qat-MTPQwopus3.6-35B-A3B-Coder-MTPZ-Image-Turbo
Z-Image-Turbo uses the locally built sd-server from
STABLE_DIFFUSION_DIR, mounted read-only into the gateway container. It uses
the same model paths and multi-GPU settings as stable.sh, and is placed in
the exclusive single-model group so image generation unloads any active text
model before using the GPUs. Build stable-diffusion.cpp first with
./build-from-source.sh stable-diffusion.cpp.
On the target server, from the presets directory configured in .env:
./llama-swap.sh start
./llama-swap.sh status
./llama-swap.sh restart
./llama-swap.sh stop
./llama-swap.sh and ./llama-swap.sh daemon are aliases for start. The
gateway runs as a detached Docker container and remains active after the
shell script exits.
The gateway is available at http://<server-host>:<port>/v1. Its API key is read
from LLAMA_API_KEY; the key is passed to llama-swap only, while the upstream
llama-server processes remain bound to the container loopback interface. Each
model receives an available internal port from llama-swap, and the proxy uses
that same assigned port.
Authentication is API-key based; there is no separate llama-swap username or password. The web UI and API endpoints require the key:
- In a browser prompt, enter any non-empty username and use the value of
LLAMA_API_KEYas the password. - For OpenAI-compatible clients, set the base URL to
http://<server-host>:<port>/v1and use the same value as the API key. The client should send it asAuthorization: Bearer <API key>. http://<server-host>:<port>/healthis unauthenticated and returnsOK.
The derived single-model presets are in llama-swap-presets/. The llama.cpp
model logs are written to ${PRESETS_DIR}/llama-swap-logs/; ZImage logs remain
attached to llama-swap's monitored upstream stream. Activity history and metrics
are stored in ${PRESETS_DIR}/llama-swap-state/activity.sqlite, so the
llama-swap Activity page retains previous requests and runs across restarts.
The request/response capture opened from an Activity row is different: current
llama-swap keeps that compressed capture in memory only, so the body of a
/chat/completions request is lost when the gateway process restarts even
though the Activity row remains. There is currently no supported
captureDirectory or store.captures setting. Use client-side logging or a
separate reverse proxy if the full request and response bodies must survive a
restart.
The state directory is created automatically by llama-swap.sh and is mounted
into the container through the existing /presets volume. The currently loaded
model process is still intentionally restarted and must be loaded again after a
container restart; only historical activity and metrics are persisted.
Models are automatically unloaded after 30 minutes of inactivity because
globalTTL is set to 1800 seconds. Increase this value to reduce reloads, or
set it to 0 to keep the active model loaded indefinitely.
The gateway is configured with an exclusive single-model group. The Pi agent
selects the model by sending its exact model ID in each OpenAI-compatible
request. When that ID changes, llama-swap unloads the current model before
starting the requested one; the Pi agent must not call /models/sse.
To discover the available IDs, call GET /v1/models with the API key. A chat
request should use the selected ID in its model field:
{
"model": "Qwen3.6-35B-MTP",
"messages": [
{"role": "user", "content": "Hello"}
]
}
Quick Start
cd <presets-directory>
./create-build-image.sh
./build-llama.cpp.sh
./server.sh start
./server.sh status
Then use the OpenAI-compatible API on port 8080. Run server.sh and
llama-swap.sh separately; they both use port 8080 by default. To use the
gateway instead, stop the standalone server and run:
./server.sh stop
./llama-swap.sh start
Then use the authenticated gateway at http://<server-host>:<port>/v1.
For image generation, select Z-Image-Turbo as the model and use the
OpenAI-compatible images endpoint:
{
"model": "Z-Image-Turbo",
"prompt": "a watercolor painting of a mountain cabin",
"size": "1024x1024"
}
The Stable Diffusion WebUI-compatible endpoints are also available through the
same gateway, including /sdapi/v1/txt2img and /sdapi/v1/img2img.