glm-4-7-flash-p150

Runs on p150 (mesh P150) โ€” 202,752-token context, up to 32 concurrent sequences.

Packaged and published with tt-model-manager 0.1.0 (manifest schema 5.1).

Quickstart

tt-model pull  stisiTT/glm-4-7-flash-p150 --with-weights
tt-model serve stisiTT/glm-4-7-flash-p150

pull --with-weights downloads the Docker image and the zai-org/GLM-4.7-Flash weights at 7dd20894a642a0aa287e9827cb1a1f7f91386b67 (into your HF cache; they are not in the image). serve starts an OpenAI-compatible server on port 20000 (or the next free port, if that one is busy); the first start compiles kernels for your device, which takes several minutes, and the server is ready when it logs Application startup complete.

Sampling defaults

This package serves with reasoning turned off by default, and caps a single response at 32,768 tokens. The launch line is:

--reasoning_parser glm47
--default-chat-template-kwargs '{"enable_thinking": false}'
--override-generation-config '{"max_new_tokens": 32768}'

The checkpoint's own generation_config.json still applies on top of that, so the served default is sampled at temperature: 1.0, not greedy. Pass an explicit temperature if you want something else. Do not serve this model with greedy defaults; see below.

To get reasoning back, ask for it per request. Request-level values override the server default:

{
  "model": "zai-org/GLM-4.7-Flash",
  "messages": [{"role": "user", "content": "..."}],
  "chat_template_kwargs": {"enable_thinking": true}
}

The glm47 reasoning parser is still enabled, so an opted-in response returns its trace in reasoning and the answer in content.

Known limitations

Reasoning can enter a repeating draft and never finish. On constraint-heavy prompts this build can loop verbatim inside its own <think> block, run to the token cap without emitting </think>, and leave the reasoning parser nothing to put in message.content. The client sees an empty response after a full token budget is spent. Measured on an IFEval prompt ("300+ words, no commas, three highlighted sections"): under greedy decoding it loops every time from about token 460; under the served sampled defaults it escapes only by chance.

This is why reasoning is off by default here and why greedy must not be made the default. It is a defect in this build, not a limit of the checkpoint: a bf16 CPU reference given the same prompt and the same greedy rule closed its reasoning at token 2,389 and answered correctly, while this build was still repeating 13,400 tokens later. Teacher forcing puts this build inside the reference distribution (85.6% argmax agreement, 803 of 804 tokens inside the reference top-p 0.95 nucleus), so no individual operation is wrong; the working hypothesis is trajectory divergence from near-tie flips introduced by the bf4 expert quantization that lets a 30.6B model fit one 32 GB chip. Investigation, evidence and candidate fixes are in doc/reasoning_loop/ in the tt-metal branch linked under Provenance.

If you need reasoning on, treat finish_reason == "length" with empty content as a retry signal. Measured on a paired 10-sample GPQA comparison (same prompts, same seed, only enable_thinking differs): reasoning enabled hit this exact failure on 2 of 10 requests. It is not rare when reasoning is on.

The 32,768-token cap applies to every request, not only ones that omit max_tokens. If you pass an explicit max_tokens above 32,768, it is silently reduced to 32,768 by the server; this is enforced unconditionally, not as a fallback default. There is no per-request way around it on this image. If you need longer completions, that is a manifest change on the next republish, not something a client can request its way past.

A full-scale run (541 ifeval + 198 gpqa, no sampling of the eval sets) confirms the empty- response failure is gone, and finds a related one that isn't. 0 of 739 responses came back empty: the reasoning-off default closes that specific failure. But 39 of 739 (5.3%) came back non-empty and unusable anyway, IE the answer itself gets stuck repeating a short span of text (a symbol, a phrase, a sentence) for tens of thousands of characters instead of answering. This includes the exact prompt used throughout this investigation, which fails this way even with reasoning off. Turning reasoning off changed the failure from silent (empty content) to loud (garbage content); it did not remove the underlying tendency. Whether this is a consequence of this build's numerical precision choices or a trait of the checkpoint itself under any precision is an open question under active investigation; see doc/reasoning_loop/ for the evidence and the comparison in progress.

Accuracy figures. Full-scale scores: ifeval 70.1% prompt-level strict / 77.3% instruction-level strict; gpqa_diamond_cot_zeroshot 47.0% (published 75.2, think-mode; not a fair baseline for this no-think default). An earlier paired 10-doc GPQA comparison (reasoning off vs on, everything else identical) scored 6/10 vs 5/10 and is superseded by the full-scale numbers above; both remain in doc/reasoning_loop/ for reference. Full investigation and evidence is in doc/reasoning_loop/ on the tt-metal branch linked under Provenance.

Provenance

The exact sources the image was built from โ€” code/ in this repo is byte-identical to the model code inside the image:

component built from
tt-metal c57266658df3c58c1328185d38c5158f0e01786b (dirty tree โ€” the image includes uncommitted changes)
vLLM v0.25.1
vllm-tt-plugin 6d3bb2854f5f8885acc1b12111d929bffdebc36e
code/ digest 37581d8697933da2 (sha256, first 16 hex digits)
built 2026-09-10T03:10:55+00:00 by tt-model 0.1.0
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support