safa-ckb-sub-onnx-ensemble
Ensemble classifier for online gender-based violence (OGBV) in Sorani Kurdish social
media comments. Four fine-tuned encoders read each comment; their logits are averaged,
a fixed per-class bias is added, and the argmax over 13 sub-categories is the
prediction. The 6 main categories and the binary hate/no-hate view derive from the
predicted sub-category through a fixed sub_to_main map. Built by the Jordan Open
Source Association (JOSA) with INSM and DRI under the SAFA project.
Architecture
| Member | Base model | Params | Precision | Size |
|---|---|---|---|---|
members/s42_int8 |
FacebookAI/xlm-roberta-base, seed 42 | 279M | INT8 dynamic | 279 MB |
members/s43_int8 |
FacebookAI/xlm-roberta-base, seed 43 | 279M | INT8 dynamic | 279 MB |
members/glot500_int8 |
cis-lmu/glot500-base | 395M | INT8 dynamic | 395 MB |
members/bernice_int8 |
jhu-clsp/bernice | 279M | INT8 dynamic | 279 MB |
1.23B parameters total, 1.23 GB. All members are RoBERTa-family, which survives dynamic INT8 quantization near-lossless; BERT-family encoders do not.
Aggregation: equal-weight mean of the four members' raw logits, plus the bias
vector in bias.json (13 values, tuned on training-side logits under an escape-rate
guardrail), then argmax. MANIFEST.json is the machine-readable serving contract.
Head: 13-way sequence classification, max_length 128. The Glot500 and Bernice
members carry Sorani orthographic canonicalization inside their saved tokenizers; it
applies at tokenize time and needs no extra preprocessing step.
Preprocessing
Inputs must pass the SAFA serving preprocessor before tokenization: strip links,
mentions, and photo tags; demojize; fold alef variants; strip diacritics; keep
Arabic-script text; cap at 50 words. No leetspeak decoding and no alef-maqsura fold
for Kurdish. Each member's training_config.json records its flags.
Inference
Per comment: preprocess once, then for each member tokenize with that member's own
tokenizer and run its session (RoBERTa graphs take no token_type_ids), mean the four
logit vectors, add the bias vector, argmax, and map sub to main. CPU is the intended
runtime; the full ensemble serves within a 2 GB budget with no GPU.
Evaluation
Frozen held-out test set, 6,518 comments, never used during development. Scores are on the INT8 ONNX serving path with the bias applied.
| Level | Accuracy | Weighted-F1 | Macro-F1 |
|---|---|---|---|
| Sub (13 classes) | 0.722 | 0.719 | 0.548 |
| Main (6, derived) | 0.726 | 0.723 | 0.712 |
| Binary (derived) | 0.836 | 0.835 | 0.834 |
Harmful-class macro-F1 (the selection metric): 0.5149. Escape rate, the share of gold-harmful comments predicted benign: 0.137, down from 0.206 for the previous single-model deployment.
Training data
The SAFA Sorani Kurdish corpus: 65,774 labeled comments (48,333 general monitoring, 17,441 election period) from monitored public pages on Facebook, TikTok, and X. Two annotators labeled each comment; a lead annotator adjudicated disagreements. The dataset is gated: thejosango/safa-ckb-dataset, access on request.
Intended use and limitations
Built for monitoring, research, and human-review triage. Outputs must feed human decisions; do not use the model for automated enforcement against individuals. Per-class performance drops hard on the rarest classes (death threats: 2 test examples; political harassment: 14), and quality degrades outside public dialectal comments. Confidence scores support triage, and reviewers should stay closest to the rare classes. This is, to our knowledge, the first OGBV sub-category classifier for Sorani Kurdish; treat it as a strong baseline, not a solved problem.
Model tree for thejosango/safa-ckb-sub-onnx-ensemble
Base model
FacebookAI/xlm-roberta-baseDataset used to train thejosango/safa-ckb-sub-onnx-ensemble
Evaluation results
- Binary macro-F1 (derived) on SAFA Sorani Kurdish Datasettest set self-reported0.834
- Harmful-class macro-F1 on SAFA Sorani Kurdish Datasettest set self-reported0.515
- Accuracy (13-class) on SAFA Sorani Kurdish Datasettest set self-reported0.722