A newer version of this model is available: RuVI-AI/pragya-test-3

Pragya Β· test2

A 21.4M-parameter language model trained from scratch β€” for testing, not for production.

Params Context Vocab Val PPL License


⚠️ Not a product

This is a scratch checkpoint used to shake out bugs in the training stack and sanity-check how the architecture scales. It is published because it might be useful as a small, from-scratch baseline β€” not because it is finished, aligned, or safe.

If you're looking for a capable small language model, this is not it.


Why this exists

A test artifact for the training pipeline. It was trained to answer questions like:

  • Does the training loop converge cleanly at this architecture and scale?
  • Do the fused Triton kernels produce numerically correct output vs. an eager reference?
  • Does loss scale as expected with batch size, context length, and depth?
  • Do the optimizer, LR schedule, and mixed-precision paths behave as designed?

The checkpoint itself is disposable. What it demonstrates is that the stack works end-to-end on a from-scratch 21M model.


Model card

Parameters 21.4M
Architecture Decoder-only transformer
Layers 16
Attention 12 query heads Β· 6 key/value heads (GQA)
Hidden size 384
Head dimension 32
FFN SwiGLU
Context length 1024 tokens
Vocabulary 16,384 (byte-level BPE, tiktoken / RustBPE)
Positional RoPE, base 10,000
Attention mask Dense causal
Embeddings Tied (input embedding = LM head)
Training precision BF16

On the training config: many flags in the internal config file β€” MoE, Mixture-of-Depths, MTP, quantization β€” are off or unused in this run. Treat the model card above as the spec, not the config file.


Training

Unique tokens ~2.9B
Epochs 1
Steps 12,000
Batch 320 sequences Γ— 1024 tokens
Optimizer Muon (2D weights) + AdamW (embeddings, norms, biases)
LR schedule 500-step warmup β†’ WSD with 2,500-step decay
Final val perplexity 16.0

Data mix

Source Role
roneneldan/TinyStories Simple narrative English
littlelearner/LittleCurriculum Structured educational text
HuggingFaceTB/cosmopedia Synthetic textbook-style content
HuggingFaceFW/fineweb-edu Filtered educational web text
HuggingFaceTB/smollm-corpus Mixed high-quality web text

English only. No additional filtering for toxicity, bias, or factual accuracy beyond each dataset's upstream filtering.


Capabilities

βœ… It can

  • Continue English prompts for a sentence or two
  • Produce grammatically plausible text in a few registers (narrative, encyclopedic, instructional)
  • Serve as a fine-tuning starting point
  • Act as a control baseline for training experiments

❌ It can't

  • Answer questions factually
  • Follow instructions
  • Stay coherent past the first paragraph
  • Handle non-English input
  • Deal with context longer than 1024 tokens

At 21M parameters, output drifts. That's the point of publishing it β€” you can see exactly where a from-scratch model at this size falls apart, and use it as a control when testing a new training change.


Sampling

For usable output, sample tightly:

temperature        0.6 – 0.9
top_k              20 – 50
top_p              0.9        (if using nucleus instead of top-k)
repetition_penalty 1.2 – 1.4
no_repeat_ngram    3

At this size, tight sampling matters more than usual β€” the model has very little capacity to recover from an off-distribution token. Loose settings turn grammatical drift into word salad within two sentences.


Files

pragya.pt                 ← checkpoint (state_dict + config)
tokenizer/tokenizer.pkl   ← pickled tiktoken Encoding (vocab 16,384)
model.py                  ← GPT architecture (GPT / GPTConfig)
inference.py              ← generate(), load_model(), sampling helpers

The tokenizer is a pickled tiktoken.Encoding, not a HuggingFace tokenizer. Loading requires the model code from the same repo. It will not work with AutoTokenizer.from_pretrained(...).

Quick start

python inference.py --checkpoint pragya.pt --prompt "The Internet is" \
    --top-k 20 --top-p 0.9 --temperature 0.6

inference.py handles checkpoint loading, KV-cached generation, sampling (top-k / top-p / repetition penalty / no-repeat n-grams), and streaming output. model.py contains the full architecture β€” no external dependencies beyond PyTorch and the tokenizer.


License

Apache 2.0. Each training dataset carries its own license β€” check upstream terms before redistributing derivatives.

Pragya Β· test2 Β· by Arush Kumar

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train RuVI-AI/pragya-test2