gdiamos/amx-moe-eda-p250x4

What this checkpoint actually is

This is not the base model described below --- it is that model fine-tuned on a specific task, and the numbers that matter are the task's.

Task. Given a C program that fails its tests, plus the compiler/test output, emit a plan in English ("First, include nrl.h. Second, define a struct type md5_t. ...") and then each function body as a separate turn. The corpus is Sudnya/classic-eda-c-trajectories, decomposed so that preamble + sum(lead + body) + tail reconstructs each source file byte for byte.

Training. Initialised from amx-moe-e256-4day, fine-tuned on 250 programs (757 blocks + 250 plans) for 200 passes per target, no replay, 20,631 steps on one core.

Measured performance

Generated blocks are substituted into the real program and the whole file is compiled against a reconstructed nrl.h runtime (validated at 97.1% on the corpus's own corrected programs). A record whose unmodified source does not compile is excluded rather than scored zero.

own training programs held-out programs
blocks that compile 13.9% (n=317) 10.4% (n=67)
compile and attempt the work 7.6% 1.5%
blocks that are correct see below 0 (n=196)

Loss on own targets 0.5287 nats/token, top-1 0.865.

"Attempt the work" excludes three degenerate strategies that all compile, because C is permissive enough that a model can score well on compile rate without writing anything:

  1. an empty body;
  2. a signature with every condition and literal zeroed --- if (c == 0) return 0; three times where the real block names font_H/font_I/font_SP;
  3. a real call, a guard, and a constant --- uint32_t len = nrl_strlen(s); if (len == 0) return 0; return 0; where the target loops over the string.

The third defeats a naive "does it reference anything outside its own signature" check. It accounts for 9 of this model's 33 apparently-non-trivial seen compiles, and 19 of 31 for its 50-pass sibling --- so any compile-rate number for this task should be published with a guard of this kind stated.

What it cannot do

It does not write correct C for programs it has not seen. Across every arm of the study this checkpoint comes from --- 8 to 301 programs, 50 to 531 passes --- 3,472 held-out blocks were generated and compiled. 27 survived the guards above. Exactly one is correct, and it came from a different arm (30 programs, 500 passes), not from this one:

want:  static char *get_domain(const char *email) {
           const char *at = nrl_strchr(email, '@');
           if (!at) return NULL;
           return at + 1; }
got:   static char *get_domain(const char *name) { ...otherwise identical... }

The other 26 call real functions and do the wrong thing: a copy_atom that never copies, a write_u32 that ignores both arguments and prints a usage message, an err_exit that prints an unrelated hardcoded string and omits the exit call that is its purpose.

Note that exact string match scores the correct one as wrong, because of the parameter rename. Held-out correctness here should be measured by alpha-equivalence --- identical after renaming parameters positionally.

The measured cause is retrieval, not capacity or data volume: held-out plans keep their structure perfectly and fabricate every identifier, and 11 of 13 identifiers the model needed were verbatim in its own prompt, 71--228 tokens back --- inside the 256-token attention window. Receptive field, data scaling (10^11 programs at the fitted exponent of -0.21) and capacity (a single example memorised to 0.0016 nats, 893 tokens reproduced verbatim) are each ruled out by measurement.

Two axes govern this

Fit --- how well a model reproduces its own training targets --- sets what it can do on material it has seen. Program count sets what transfers. They are close to independent:

programs fit seen held-out
memoriser 8 0.0150 60.0% 0.0%
this model 250 0.5287 7.6% 1.5%
150-program arm 150 1.0757 1.5% 1.3%

A model fitted almost perfectly on 8 programs transfers nothing; a badly fitted model on 150 programs transfers more. Held-out loss scales as programs^-0.213, so every 10x of data buys about 1.4x of held-out output rate --- and reaching seen-level performance on unseen programs extrapolates to ~2.6e9 programs.

Why this checkpoint and not its sibling

p250 --- same data, 50 passes instead of 200 --- has better held-out loss (1.53 vs 2.04 nats) and worse compile behaviour (5.5% vs 10.8% non-trivial). Fitting harder buys valid code and costs generalisation. This checkpoint is the fitted end of that trade; it is published as a research artifact for studying what governs compile rate, not as a useful C repairer.

A causal language model trained end to end on one CPU core --- a single Intel Emerald Rapids core, bf16 through AMX, OMP_NUM_THREADS=1. 3,315,744 active parameters per token, 30,029,426 stored.

The point of the project is not that a small model runs on a CPU. It is that the architecture is derived from a single-core roofline, that training is confined to the same core, and that at this scale the interesting behaviours show up much earlier in the token budget than we expected.

What it does

This is a base model: it predicts the next token and has had no instruction tuning. It is well calibrated under teacher forcing and it cannot generate --- free-running, it enters a repetition basin within about five tokens. That is expected of this checkpoint and is what the instruction-tuned sibling exists to fix. Use it as a starting point for fine-tuning, not as a generator.

Running it

AutoModelForCausalLM.from_pretrained will not work: the architecture is not one transformers knows --- chunked sliding-window attention interleaved with log-decay linear attention, and a tied readout and block-routed experts. The model's own source ships here under m2r/, unmodified from the repository that trained it.

hf download gdiamos/amx-moe-eda-p250x4 --local-dir amx-moe-eda-p250x4
cd amx-moe-eda-p250x4 && pip install -r requirements.txt && python example.py
import sys, torch
from safetensors.torch import load_file
from tokenizers import Tokenizer
sys.path.insert(0, ".")                 # the folder you downloaded

from m2r.config import load
from m2r.model.torch_model import Model, swa_mask

cfg = load("training_config.yaml")
model = Model(cfg.model).to(torch.bfloat16)
model.load_state_dict(load_file("model.safetensors"))
model.eval()
tok = Tokenizer.from_file("tokenizer.json")
mask = swa_mask(cfg.model, dtype=torch.bfloat16)

# Pad to a whole number of blocks and read the last REAL position. Right-padding
# is safe -- attention is causal and the MLP is position-wise.
PAD_TO = max(cfg.model.window, cfg.model.route_block or 1, 256)
ids = [1] + tok.encode("The capital of France is").ids    # 1 is BOS
for _ in range(60):
    n = len(ids)
    x = torch.tensor([ids + [0] * ((-n) % PAD_TO)])
    with torch.no_grad():
        h = model.body(x, mask)[:, n - 1]
    logits = (h @ model.emb.t().to(h.dtype)).float()[0] / 0.8
    v, i = logits.topk(40)
    ids.append(int(i[torch.multinomial(v.softmax(-1), 1)]))
print(tok.decode(ids[1:]))

Sample rather than take the argmax: greedy decoding loops within a few tokens. BOS (id 1) matters --- the model was trained with attention confined to document boundaries keyed on that token, so a prompt without it is unlike anything it saw in training.

What is in this repo

file
model.safetensors the weights, bf16
generation.json decode settings, and the vocabulary mask described below
config.json every architecture field, machine readable
training_config.yaml the run's config, and what example.py loads
tokenizer.json a tokenizers BPE; Tokenizer.from_file loads it alone
m2r/ the model source, imported by example.py
example.py load and generate, correctly
paper.pdf the write-up, when shipped with this export
LICENSE Apache 2.0

Architecture

d_model 256
layers 6
layer types lin, swa, swa, lin, swa, lin
mixers sliding-window attention (window 256), log-decay linear attention (d_state 32)
MLP width 640
vocabulary 16384
readout tied to the embedding
parameters 30,029,426 stored, 3,315,744 active per token
experts 64, top-4, d_ff_e 160, route_block 256
MoE layers [0, 3, 5]

Attention is confined to document boundaries: a training window packs many documents, and without isolation sliding-window attention reaches into its neighbours while linear attention carries state across the whole window.

The shape is deliberate. One AMX core sustains roughly 2,231 GF/s of bf16 matrix multiply at these dimensions but pays a 1.4--1.5 microsecond floor per GEMM dispatch, so every design choice here is about issuing few large matrix multiplies rather than many small ones.

Training data

source tokens share
blocks.eda_blocks 84,505,400 100.0%
total prepared 84,505,400
consumed by this run 84,504,576 100.0%

Every natural-language and code source above is a curated artefact built with the help of large models --- quality classification, rephrasing, model-assisted extraction, and in the case of the reasoning corpus, traces that are themselves generated output. Training a model this small on them is a form of distillation, with no teacher present at training time. This is worth stating plainly, because it means results at this scale depend on corpora that did not exist when models of this size were last studied seriously.

Validation loss

The run's own numbers on its fixed held-out set, as logged. These are a sampled loss --- a (1 + n_negatives)-way discrimination, not a full-vocabulary one --- except val_flat, which is full-vocabulary.

{
  "steps": 20630,
  "tokens": 84504576
}

Limitations

A research artifact, and the honest summary is that the failures are specific rather than diffuse.

It degenerates into repetition. Free-running, it enters an absorbing state within about five tokens, in every domain. No decode-time patch fixes this --- truncation sampling has nothing to reshape in a distribution that concentrated. Instruction tuning does fix it, which is what the instruct sibling of this repo is.

It has had no alignment, safety, or preference training of any kind, and will reproduce the biases and errors of its training corpus.

License and provenance

Apache 2.0, for the weights and for the source in m2r/; full text in LICENSE.

The training data is a mixture of public code, web, math and instruction corpora, with per-source token counts above. Those corpora carry their own terms, which the Apache licence on this model does not alter and does not extend to them.

Produced by tools/export_hf.py from run torch-20260923T091725-eda-prog-p250x4 at step 20631.

Downloads last month
476
Safetensors
Model size
30M params
Tensor type
I64
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support