glm-4-7-flash-p150
Runs on p150 (mesh P150) โ 202,752-token context, up to 32 concurrent sequences.
Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).
Quickstart
tt-model pull stisiTT/glm-4-7-flash-p150 --with-weights
tt-model serve stisiTT/glm-4-7-flash-p150
pull --with-weights downloads the Docker image and the zai-org/GLM-4.7-Flash weights at 7dd20894a642a0aa287e9827cb1a1f7f91386b67 (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.
Sampling defaults
This package serves with reasoning turned off by default, and caps a single response at 32,768 tokens. The launch line is:
--reasoning_parser glm47
--default-chat-template-kwargs '{"enable_thinking": false}'
--override-generation-config '{"max_new_tokens": 32768}'
The checkpoint's own generation_config.json still applies on top of that, so the served
default is sampled at temperature: 1.0, not greedy. Pass an explicit temperature if
you want something else. Do not serve this model with greedy defaults; see below.
To get reasoning back, ask for it per request. Request-level values override the server default:
{
"model": "zai-org/GLM-4.7-Flash",
"messages": [{"role": "user", "content": "..."}],
"chat_template_kwargs": {"enable_thinking": true}
}
The glm47 reasoning parser is still enabled, so an opted-in response returns its trace in
reasoning and the answer in content.
Known limitations
Reasoning can enter a repeating draft and never finish. On constraint-heavy prompts this
build can loop verbatim inside its own <think> block, run to the token cap without emitting
</think>, and leave the reasoning parser nothing to put in message.content. The client sees
an empty response after a full token budget is spent. Measured on an IFEval prompt ("300+ words,
no commas, three highlighted sections"): under greedy decoding it loops every time from about
token 460; under the served sampled defaults it escapes only by chance.
This is why reasoning is off by default here and why greedy must not be made the default. It is
a defect in this build, not a limit of the checkpoint: a bf16 CPU reference given the same prompt
and the same greedy rule closed its reasoning at token 2,389 and answered correctly, while this
build was still repeating 13,400 tokens later. Teacher forcing puts this build inside the
reference distribution (85.6% argmax agreement, 803 of 804 tokens inside the reference top-p 0.95
nucleus), so no individual operation is wrong; the working hypothesis is trajectory divergence
from near-tie flips introduced by the bf4 expert quantization that lets a 30.6B model fit one
32 GB chip. Investigation, evidence and candidate fixes are in doc/reasoning_loop/ in the
tt-metal branch linked under Provenance.
If you need reasoning on, treat finish_reason == "length" with empty content as a retry
signal. Measured on a paired 10-sample GPQA comparison (same prompts, same seed, only
enable_thinking differs): reasoning enabled hit this exact failure on 2 of 10 requests. It is
not rare when reasoning is on.
The 32,768-token cap applies to every request, not only ones that omit max_tokens. If you
pass an explicit max_tokens above 32,768, it is silently reduced to 32,768 by the server; this
is enforced unconditionally, not as a fallback default. There is no per-request way around it on
this image. If you need longer completions, that is a manifest change on the next republish, not
something a client can request its way past.
A full-scale run (541 ifeval + 198 gpqa, no sampling of the eval sets) confirms the empty-
response failure is gone, and finds a related one that isn't. 0 of 739 responses came back
empty: the reasoning-off default closes that specific failure. But 39 of 739 (5.3%) came back
non-empty and unusable anyway, IE the answer itself gets stuck repeating a short span of text
(a symbol, a phrase, a sentence) for tens of thousands of characters instead of answering. This
includes the exact prompt used throughout this investigation, which fails this way even with
reasoning off. Turning reasoning off changed the failure from silent (empty content) to loud
(garbage content); it did not remove the underlying tendency. Whether this is a consequence of
this build's numerical precision choices or a trait of the checkpoint itself under any precision
is an open question under active investigation; see doc/reasoning_loop/ for the evidence and
the comparison in progress.
Accuracy figures. Full-scale scores: ifeval 70.1% prompt-level strict / 77.3%
instruction-level strict; gpqa_diamond_cot_zeroshot 47.0% (published 75.2, think-mode; not a
fair baseline for this no-think default). An earlier paired 10-doc GPQA comparison (reasoning off
vs on, everything else identical) scored 6/10 vs 5/10 and is superseded by the full-scale numbers
above; both remain in doc/reasoning_loop/ for reference. Full investigation and evidence is in
doc/reasoning_loop/ on the tt-metal branch linked under Provenance.
Provenance
The exact sources the image was built from โ code/ in this repo is byte-identical to the model code inside the image:
| component | built from |
|---|---|
| tt-metal | c57266658df3c58c1328185d38c5158f0e01786b (dirty tree โ the image includes uncommitted changes) |
| vLLM | v0.25.1 |
| vllm-tt-plugin | 6d3bb2854f5f8885acc1b12111d929bffdebc36e |
code/ digest |
37581d8697933da2 (sha256, first 16 hex digits) |
| built | 2026-09-10T03:10:55+00:00 by tt-model 0.1.0 |