Qwen3.5-4B TikZ LoRA

Stage 1 LoRA supervised fine-tuning for instruction-to-TikZ generation. This flat repository contains the exact pinned BF16 base checkpoint and the unmerged LoRA adapter at repository root, plus tokenizer, metrics, provenance, and checksums.

Provenance

  • Base model: Qwen/Qwen3.5-4B
  • Base revision: 851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a
  • Training method: BF16 LoRA, not QLoRA
  • Trainer: Unsloth native loading with TRL SFTTrainer
  • Hardware: NVIDIA RTX PRO 6000 Blackwell Server Edition, one GPU
  • Maximum sequence length: 8192
  • Trainable parameters observed: 84,934,656
  • Total parameters observed: 4,624,200,192
  • Trainable fraction: 1.8376%

Dataset and tokens

  • Training examples: 94674
  • Eligible selected examples: 94674
  • Quarantined examples excluded: 1841
  • Prompt tokens: 20010926
  • Supervised assistant tokens: 61945327
  • Unpadded dataset tokens: 81956253
  • Trainer-reported input tokens processed: 116653068

Only assistant/TikZ tokens contributed to loss. Prompt tokens were masked and overlength examples were quarantined rather than truncated.

Optimization

  • Epochs completed: 1.000000
  • Effective batch size: 16
  • Per-device batch size: 2
  • Gradient accumulation: 8
  • Learning rate: 0.0001
  • Scheduler: cosine
  • LoRA rank: 64
  • LoRA alpha: 64
  • LoRA dropout: 0.0

Results

  • Overall gate passed: True
  • Training completed: True
  • Full epoch completed: True
  • Artifact complete: True
  • Validation loss: 0.320700
  • Trainer-reported train loss: 0.070897
  • First logged loss: 0.824556
  • Final logged loss: 0.293612
  • Minimum logged loss: 0.241120
  • Mean logged loss: 0.355989
  • Trainer-reported input tokens: 116653068
  • Total FLOPs: 2.78997394992674e+18
  • Accumulated Slurm allocation time: 6h 24m 36s
  • Training allocations: 4

Loss is token-level cross-entropy against one reference TikZ implementation; it does not directly measure compilation or rendered-image similarity.

Reporting note: The Trainer-reported aggregate training loss belongs to the final resumed allocation and is not the mean over the complete run. Validation loss and the individual logged losses are the reliable loss records for this resumed training.

Slurm allocations

Job ID Name State Elapsed Node
4327259 qwen35-native-full TIMEOUT 01:00:02 a2841
4327405 qwen35-native-full-r2 TIMEOUT 02:00:28 a2841
4327662 qwen35-native-full-r3 TIMEOUT 02:00:06 a2041
4327913 qwen35-native-full-final COMPLETED 01:24:00 a2841

Loading

The standard Transformers PEFT integration can load the adapter using its recorded base-model reference:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "Praha-Labs/Qwen3.5-4B-TikZ-LoRA"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo, torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True
)

Explicit PEFT loading against the pinned upstream base is also supported:

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

repo = "Praha-Labs/Qwen3.5-4B-TikZ-LoRA"
base = AutoModelForCausalLM.from_pretrained(
    "Qwen/Qwen3.5-4B", revision="851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a",
    torch_dtype=torch.bfloat16, device_map="auto", trust_remote_code=True,
)
tokenizer = AutoTokenizer.from_pretrained(repo)
model = PeftModel.from_pretrained(base, repo)

Included evidence

training-metrics.json, resolved-config.json, training-config.yaml, native-run.json, token-report.json, training-jobs.tsv, environment.lock, source-commit.txt, BASE_MODEL_CARD.md, and SHA256SUMS document the run and its provenance.

Limitations

Compilation rate, rendered-image similarity, and human evaluation should be measured before deployment because multiple valid TikZ programs can render the same diagram.

Downloads last month
16
Safetensors
Model size
5B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Praha-Labs/Qwen3.5-4B-TikZ-LoRA

Finetuned
Qwen/Qwen3.5-4B
Adapter
(674)
this model