PyCraft-1: A 55M Python Code LLM Trained From Scratch on Consumer Hardware
Model Description
PyCraft-1 is a 55.3M parameter decoder-only transformer trained entirely from scratch on Python code using a single NVIDIA RTX 3050 laptop GPU (4GB VRAM). It demonstrates that a domain-specific code LLM can be trained without cloud compute, following a quality-first data curriculum inspired by the Phi-1 "Textbooks Are All You Need" approach.
It runs comfortably on CPU — no GPU required for inference.
Architecture
| Component | Choice |
|---|---|
| Architecture | Decoder-only transformer |
| Parameters | 55.3M |
| Attention | Grouped Query Attention (8Q / 2KV heads) |
| Positional encoding | RoPE |
| QK-Norm | RMSNorm on Q and K (OLMo 2 / Qwen 3, 2025) |
| FFN | SwiGLU |
| Normalisation | RMSNorm pre-norm |
| Training objective | Causal LM + Fill-in-the-Middle (FIM, 50%) |
| Context window | 1024 tokens |
| Vocabulary | 32,000 (custom BPE, Python-tuned) |
Training
Pretraining:
- 309,221 curated Python examples from 6 open sources
- Quality curriculum: scored on 5 heuristics, ordered best-first
- 4,000 steps | 1.05B tokens | Loss 1.16 | PPL 3.2
Supervised Fine-Tuning:
- Magicoder-OSS-Instruct-75K (Python subset, 40k examples)
- 400 steps | Loss 1.15 | PPL 3.15
Hardware: NVIDIA RTX 3050 Laptop GPU 4GB, Ryzen 7 6000, 16GB RAM Total training time: ~22 hours
Novel Contributions
Quality-first data curriculum — 309k Python examples scored and ordered by educational value (docstrings, type hints, comments, naming conventions, length). Validates Phi-1 hypothesis at the resource-constrained regime.
QK-Norm — RMSNorm applied to Q and K before RoPE, adopted from OLMo 2 and Qwen 3 (2025), improving training stability in small models.
FIM pretraining on 4GB VRAM — Fill-in-the-Middle objective using PSM format trained on a consumer GPU via gradient checkpointing and BF16 mixed precision.
Full reproducibility — complete open-source pipeline runnable on consumer hardware in under one week.
Evaluation
| Metric | Value |
|---|---|
| Pretraining loss | 1.16 |
| Pretraining PPL | 3.20 |
| SFT loss | 1.15 |
| SFT PPL | 3.15 |
| Held-out PPL (binary search) | 1.4 |
| Held-out PPL (Stack class) | 1.7 |
| Held-out PPL (average, 5 samples) | 2.16 |
HumanEval
| Model | Size | Training tokens | Hardware | Pass@1 |
|---|---|---|---|---|
| GPT-Neo | 125M | 300B | multi-GPU | 0.83% |
| PyCraft-1 | 55M | 1.05B | 1× RTX 3050 | 3.66% |
| CodeParrot | 110M | 50B | multi-GPU | 3.80% |
3.66% Pass@1 (6/164) with greedy decoding, the conventional and
reproducible setting for single-sample pass@1. Passing problems:
greatest_common_divisor, strlen, get_positive, is_prime,
remove_vowels, add.
That is 4.4× GPT-Neo 125M on 0.35% of its training tokens, and close to CodeParrot 110M on 2% of its tokens, from a model roughly half their size.
Usage
pip install "torch>=2.1" safetensors tokenizers
Download this repository's model.safetensors, the model/ package, and
tokenizer/tokenizer.json, then:
import torch
from safetensors.torch import load_file
from model.config import get_config_120m
from model.pycraft_model import PyCraftModel
from tokenizer.tokenizer_utils import PyCraftTokenizer
device = "cuda" if torch.cuda.is_available() else "cpu"
tokenizer = PyCraftTokenizer("tokenizer/tokenizer.json")
cfg = get_config_120m()
cfg.vocab_size = 32000
cfg.dropout = 0.0
model = PyCraftModel(cfg).to(device)
model.load_state_dict(load_file("model.safetensors", device=device))
model.eval()
prompt = "def is_palindrome(s: str) -> bool:\n "
ids = tokenizer.encode(prompt)
inp = torch.tensor(ids, dtype=torch.long).unsqueeze(0).to(device)
with torch.no_grad():
out = model.generate(
inp,
max_new_tokens=80,
temperature=0.2, # 0.0 selects greedy decoding
top_k=20,
repetition_penalty=1.1,
)
print(tokenizer.decode(out[0, len(ids):].tolist(), skip_special_tokens=True))
generate() uses a KV cache, so cost per token stays flat instead of
growing with context. On CPU (8 threads) that is roughly 49–62 tokens/s
depending on context length, against 7–31 without it.
Generation stops at <|endoftext|> by default. stop_strings=[...]
(with tokenizer=) and top_p are also supported.
Fill in the Middle
Half of pretraining used the FIM objective in PSM order, so infilling is trained behaviour rather than a prompting trick:
tok = tokenizer
prefix = "def factorial(n):\n if n <= 1:\n return 1\n "
suffix = "\n\nprint(factorial(5))\n"
ids = ([tok.prefix_id] + tok.encode(prefix)
+ [tok.suffix_id] + tok.encode(suffix)
+ [tok.middle_id])
inp = torch.tensor(ids, dtype=torch.long).unsqueeze(0).to(device)
with torch.no_grad():
out = model.generate(inp, max_new_tokens=40, temperature=0.2, top_k=20,
eos_token_id=[tok.eot_id, tok.middle_id])
print(tokenizer.decode(out[0, len(ids):].tolist(), skip_special_tokens=True))
# -> return n * factorial(n-1)
CLI and local REST API
The GitHub repository ships an installable package with a CLI and a
FastAPI server (completions with SSE streaming, plus a /v1/fim endpoint):
pip install -e ".[serve]"
pycraft generate "def is_prime(n):"
pycraft serve # http://127.0.0.1:8000/docs
Limitations
- 55M parameters: generates plausible Python but may produce logical errors
- Context window: 1024 tokens, and it is a hard stop rather than a sliding window (the KV cache stores post-RoPE keys)
- Best on standard algorithms, data structures, and common library patterns
- The SFT data contained markdown code fences, so raw output often includes them; strip them before executing
- Not a conversational assistant
Citation
@misc{inamdar2026pycraft,
title={PyCraft-1: Training a Python Code LLM From Scratch on Consumer Hardware},
author={Inamdar, Rohan},
year={2026},
institution={University of Manchester, MSc Artificial Intelligence},
}
Links
- Downloads last month
- 28