metadata
license: apache-2.0
language:
- multilingual
tags:
- content-classification
- byte-level
- onnx
- matryoshka
- lightweight
- classifier
pipeline_tag: text-classification
library_name: pico-type
pico-type
A tiny byte-level multi-head content classifier (~1.5M parameters) that classifies any content into 7 categories simultaneously from raw bytes — no tokenizer, no pretrained embeddings.
Architecture
ByteEmbed → Conv1D×3 → BiAttention×2 → Pool → Matryoshka Heads
- Byte-level: operates directly on UTF-8 bytes, supports any language
- Matryoshka heads: 7 independent classification heads with 4 tiers (tiny/small/base/pro)
- 1.5M params: fits in ~9MB single-file ONNX (FP32), runs in ~18ms on CPU
- No tokenizer: zero vocabulary dependencies
Classification Heads
| Head | Classes | Description |
|---|---|---|
| coarse | 12 | text, code, link, image, file, config, markup, data, error, secret, archive, binary |
| modality | 8 | textual, binary_image, binary_archive, binary_executable, binary_document, binary_audio, binary_video, binary_other |
| subtype | 24 | json, yaml, toml, ini, csv, html, xml, markdown, sql, log, diff, dockerfile, etc. |
| code_lang | 62 | python, javascript, typescript, java, c, cpp, go, rust, swift, bash, sql, etc. |
| text_lang | 30 | en, es, fr, de, it, pt, ru, zh, ja, ko, ar, hi, etc. |
| file_mime | 90 | text/html, application/json, application/pdf, image/png, video/mp4, etc. |
| risk | 6 | api_key, jwt, password, email, phone, ssh_key (probabilities) |
Performance
Benchmarked on synthetic data (500 samples, 1024 bytes max, base tier, 1700 training steps):
| Head | Accuracy | Support |
|---|---|---|
| coarse | 100.0% | 500 |
| modality | 100.0% | 500 |
| subtype | 93.8% | 128 |
| code_lang | 41.7% | 48 |
| text_lang | 94.3% | 35 |
| file_mime | 100.0% | 131 |
| risk (mAP) | 100.0% | — |
- Inference: ~18ms per sample on CPU (ONNX Runtime, M2)
- Model size: ~9.25MB single-file FP32 (base tier; 206 KB graph + 9.05 MB weights)
- Loss: 1.97 eval_loss (best, step 1700)
code_lang accuracy (54.2%) reflects 62-class coverage; improves with longer sequences (>256 bytes). v0.2 will target better code language discrimination.
Usage
CLI
# Pipe content
echo "def hello(): pass" | picotype --pretty
# File
picotype --file document.txt
# Clipboard (macOS)
picotype --clip
Python
from model.pico_type.labels import decode_output
# Run with ONNX session
result = {"coarse": "code", "modality": "textual", ...}
decoded = decode_output(result, tier="base")
MCP Server
PICOTYPE_MODEL_DIR=./checkpoints python -m model.pico_type.mcp_server
Model Tiers
| Tier | Head Dim | Params | ONNX Size |
|---|---|---|---|
| tiny | 16 | 1.43M | 9.09 MB |
| small | 64 | 1.45M | 9.13 MB |
| base | 192 | 1.48M | 9.25 MB |
| pro | 576 | 1.56M | 9.61 MB |
All tiers share the same trunk; only the final linear layer differs per tier.
Deployment
HuggingFace Space
The Gradio Space provides:
- Text input and file upload
- Real-time 7-head classification
- Tier selection (tiny/small/base/pro)
ONNX Runtime
import onnxruntime
session = ort.InferenceSession("picotype_base.onnx")
Training
Trained on synthetic data (11 content buckets, 62 code languages, 30 text languages, 90 MIME types) using multi-task loss with 500 optimization steps.
- Loss: weighted cross-entropy (coarse) + binary cross-entropy (risk)
- Optimizer: AdamW (lr=1e-3, weight_decay=0.01)
- GPU: ~100ms/step on MPS, ~3.5s/step on CPU
License
Apache 2.0