gdiamos/amx-moe-eda-p250x4
What this checkpoint actually is
This is not the base model described below --- it is that model fine-tuned on a specific task, and the numbers that matter are the task's.
Task. Given a C program that fails its tests, plus the compiler/test
output, emit a plan in English ("First, include nrl.h. Second, define a
struct type md5_t. ...") and then each function body as a separate turn.
The corpus is Sudnya/classic-eda-c-trajectories, decomposed so that
preamble + sum(lead + body) + tail reconstructs each source file byte for
byte.
Training. Initialised from amx-moe-e256-4day, fine-tuned on 250
programs (757 blocks + 250 plans) for 200 passes per target, no replay,
20,631 steps on one core.
Measured performance
Generated blocks are substituted into the real program and the whole file is
compiled against a reconstructed nrl.h runtime (validated at 97.1% on the
corpus's own corrected programs). A record whose unmodified source does not
compile is excluded rather than scored zero.
| own training programs | held-out programs | |
|---|---|---|
| blocks that compile | 13.9% (n=317) | 10.4% (n=67) |
| compile and attempt the work | 7.6% | 1.5% |
| blocks that are correct | see below | 0 (n=196) |
Loss on own targets 0.5287 nats/token, top-1 0.865.
"Attempt the work" excludes three degenerate strategies that all compile, because C is permissive enough that a model can score well on compile rate without writing anything:
- an empty body;
- a signature with every condition and literal zeroed ---
if (c == 0) return 0;three times where the real block namesfont_H/font_I/font_SP; - a real call, a guard, and a constant ---
uint32_t len = nrl_strlen(s); if (len == 0) return 0; return 0;where the target loops over the string.
The third defeats a naive "does it reference anything outside its own signature" check. It accounts for 9 of this model's 33 apparently-non-trivial seen compiles, and 19 of 31 for its 50-pass sibling --- so any compile-rate number for this task should be published with a guard of this kind stated.
What it cannot do
It does not write correct C for programs it has not seen. Across every arm of the study this checkpoint comes from --- 8 to 301 programs, 50 to 531 passes --- 3,472 held-out blocks were generated and compiled. 27 survived the guards above. Exactly one is correct, and it came from a different arm (30 programs, 500 passes), not from this one:
want: static char *get_domain(const char *email) {
const char *at = nrl_strchr(email, '@');
if (!at) return NULL;
return at + 1; }
got: static char *get_domain(const char *name) { ...otherwise identical... }
The other 26 call real functions and do the wrong thing: a copy_atom that
never copies, a write_u32 that ignores both arguments and prints a usage
message, an err_exit that prints an unrelated hardcoded string and omits the
exit call that is its purpose.
Note that exact string match scores the correct one as wrong, because of the parameter rename. Held-out correctness here should be measured by alpha-equivalence --- identical after renaming parameters positionally.
The measured cause is retrieval, not capacity or data volume: held-out plans keep their structure perfectly and fabricate every identifier, and 11 of 13 identifiers the model needed were verbatim in its own prompt, 71--228 tokens back --- inside the 256-token attention window. Receptive field, data scaling (10^11 programs at the fitted exponent of -0.21) and capacity (a single example memorised to 0.0016 nats, 893 tokens reproduced verbatim) are each ruled out by measurement.
Two axes govern this
Fit --- how well a model reproduces its own training targets --- sets what it can do on material it has seen. Program count sets what transfers. They are close to independent:
| programs | fit | seen | held-out | |
|---|---|---|---|---|
| memoriser | 8 | 0.0150 | 60.0% | 0.0% |
| this model | 250 | 0.5287 | 7.6% | 1.5% |
| 150-program arm | 150 | 1.0757 | 1.5% | 1.3% |
A model fitted almost perfectly on 8 programs transfers nothing; a badly
fitted model on 150 programs transfers more. Held-out loss scales as
programs^-0.213, so every 10x of data buys about 1.4x of held-out output
rate --- and reaching seen-level performance on unseen programs extrapolates to
~2.6e9 programs.
Why this checkpoint and not its sibling
p250 --- same data, 50 passes instead of 200 --- has better held-out loss
(1.53 vs 2.04 nats) and worse compile behaviour (5.5% vs 10.8%
non-trivial). Fitting harder buys valid code and costs generalisation. This
checkpoint is the fitted end of that trade; it is published as a research
artifact for studying what governs compile rate, not as a useful C repairer.
A causal language model trained end to end on one CPU core --- a single
Intel Emerald Rapids core, bf16 through AMX, OMP_NUM_THREADS=1. 3,315,744
active parameters per token, 30,029,426 stored.
The point of the project is not that a small model runs on a CPU. It is that the architecture is derived from a single-core roofline, that training is confined to the same core, and that at this scale the interesting behaviours show up much earlier in the token budget than we expected.
What it does
This is a base model: it predicts the next token and has had no instruction tuning. It is well calibrated under teacher forcing and it cannot generate --- free-running, it enters a repetition basin within about five tokens. That is expected of this checkpoint and is what the instruction-tuned sibling exists to fix. Use it as a starting point for fine-tuning, not as a generator.
Running it
AutoModelForCausalLM.from_pretrained will not work: the architecture is
not one transformers knows --- chunked sliding-window attention interleaved
with log-decay linear attention, and a tied readout and block-routed experts. The model's
own source ships here under m2r/, unmodified from the repository that
trained it.
hf download gdiamos/amx-moe-eda-p250x4 --local-dir amx-moe-eda-p250x4
cd amx-moe-eda-p250x4 && pip install -r requirements.txt && python example.py
import sys, torch
from safetensors.torch import load_file
from tokenizers import Tokenizer
sys.path.insert(0, ".") # the folder you downloaded
from m2r.config import load
from m2r.model.torch_model import Model, swa_mask
cfg = load("training_config.yaml")
model = Model(cfg.model).to(torch.bfloat16)
model.load_state_dict(load_file("model.safetensors"))
model.eval()
tok = Tokenizer.from_file("tokenizer.json")
mask = swa_mask(cfg.model, dtype=torch.bfloat16)
# Pad to a whole number of blocks and read the last REAL position. Right-padding
# is safe -- attention is causal and the MLP is position-wise.
PAD_TO = max(cfg.model.window, cfg.model.route_block or 1, 256)
ids = [1] + tok.encode("The capital of France is").ids # 1 is BOS
for _ in range(60):
n = len(ids)
x = torch.tensor([ids + [0] * ((-n) % PAD_TO)])
with torch.no_grad():
h = model.body(x, mask)[:, n - 1]
logits = (h @ model.emb.t().to(h.dtype)).float()[0] / 0.8
v, i = logits.topk(40)
ids.append(int(i[torch.multinomial(v.softmax(-1), 1)]))
print(tok.decode(ids[1:]))
Sample rather than take the argmax: greedy decoding loops within a few tokens. BOS (id 1) matters --- the model was trained with attention confined to document boundaries keyed on that token, so a prompt without it is unlike anything it saw in training.
What is in this repo
| file | |
|---|---|
model.safetensors |
the weights, bf16 |
generation.json |
decode settings, and the vocabulary mask described below |
config.json |
every architecture field, machine readable |
training_config.yaml |
the run's config, and what example.py loads |
tokenizer.json |
a tokenizers BPE; Tokenizer.from_file loads it alone |
m2r/ |
the model source, imported by example.py |
example.py |
load and generate, correctly |
paper.pdf |
the write-up, when shipped with this export |
LICENSE |
Apache 2.0 |
Architecture
| d_model | 256 |
| layers | 6 |
| layer types | lin, swa, swa, lin, swa, lin |
| mixers | sliding-window attention (window 256), log-decay linear attention (d_state 32) |
| MLP width | 640 |
| vocabulary | 16384 |
| readout | tied to the embedding |
| parameters | 30,029,426 stored, 3,315,744 active per token |
| experts | 64, top-4, d_ff_e 160, route_block 256 |
| MoE layers | [0, 3, 5] |
Attention is confined to document boundaries: a training window packs many documents, and without isolation sliding-window attention reaches into its neighbours while linear attention carries state across the whole window.
The shape is deliberate. One AMX core sustains roughly 2,231 GF/s of bf16 matrix multiply at these dimensions but pays a 1.4--1.5 microsecond floor per GEMM dispatch, so every design choice here is about issuing few large matrix multiplies rather than many small ones.
Training data
| source | tokens | share |
|---|---|---|
blocks.eda_blocks |
84,505,400 | 100.0% |
| total prepared | 84,505,400 | |
| consumed by this run | 84,504,576 | 100.0% |
Every natural-language and code source above is a curated artefact built with the help of large models --- quality classification, rephrasing, model-assisted extraction, and in the case of the reasoning corpus, traces that are themselves generated output. Training a model this small on them is a form of distillation, with no teacher present at training time. This is worth stating plainly, because it means results at this scale depend on corpora that did not exist when models of this size were last studied seriously.
Validation loss
The run's own numbers on its fixed held-out set, as logged. These are a
sampled loss --- a (1 + n_negatives)-way discrimination, not a
full-vocabulary one --- except val_flat, which is full-vocabulary.
{
"steps": 20630,
"tokens": 84504576
}
Limitations
A research artifact, and the honest summary is that the failures are specific rather than diffuse.
It degenerates into repetition. Free-running, it enters an absorbing state within about five tokens, in every domain. No decode-time patch fixes this --- truncation sampling has nothing to reshape in a distribution that concentrated. Instruction tuning does fix it, which is what the instruct sibling of this repo is.
It has had no alignment, safety, or preference training of any kind, and will reproduce the biases and errors of its training corpus.
License and provenance
Apache 2.0, for the weights and for the source in m2r/; full text in
LICENSE.
The training data is a mixture of public code, web, math and instruction corpora, with per-source token counts above. Those corpora carry their own terms, which the Apache licence on this model does not alter and does not extend to them.
Produced by tools/export_hf.py from run torch-20260923T091725-eda-prog-p250x4 at step 20631.
- Downloads last month
- 476