dfine-det-large-baseline-stage0
Stage-0 multiscript pretrained weights for dfine-det — a D-FINE Large (HGNetv2-B4) detector adapted to predict text-line baselines as polylines (B-spline control points), not boxes.
Hub repo:
johnlockejrr/dfine-det-large-baseline-stage0
Primary file:best_cbad_f1.safetensors(~116 MB)
This checkpoint is intended as a generic baseline pretrain for downstream Stage-1 fine-tunes (e.g. Hebrew/Samaritan, ICDAR cBAD, other PAGE/ALTO corpora).
Model summary
| Architecture | PolylineDFINE — D-FINE detection core + polyline head |
| Size class | Large (D-FINE-L / backbone B4) |
| Parameters | ~30.1 M trainable |
| Queries | 300 |
| Geometry | (K=8) cubic B-spline control points per line + height |
| Canvas | (1280\times1280) letterbox |
| Init | Official Peterande/D-FINE dfine_l_obj2coco_e25.pth (strict=False) |
| Task | Document baseline / text-line detection → PAGE or ALTO export |
Detection is set prediction of scored polylines. Polygonal line environments (BLLA) are applied only at PAGE/ALTO serialize time, not in the training objective.
Intended use
Use for
- Starting point for Stage-1 fine-tuning on a target script/corpus
- Zero-shot / few-shot probing of baselines on heterogeneous manuscripts (expect domain gap)
- Research on polyline baseline detectors built on D-FINE
Not for
- Claiming official ICDAR cBAD 2019 test numbers from these Stage-0 val scores (see below)
- Production OCR transcription (this model detects lines; it does not recognize text)
- Replacing layout region detectors (paragraphs, tables, etc.)
Training data (Stage-0)
Multiscript PAGE-XML / ALTO baselines under baseline_pretrain_raw (HTR-United–style open corpora + private packs), compiled to Arrow with adaptive simplify → B-spline (K=8).
| Split | Pages |
|---|---|
| Train | 46 513 |
| Val (holdout) | 1 923 (~5%, seed 42) |
| Total XMLs inventoried | 48 436 across 29 corpora |
Corpora (train / val pages)
| Corpus | Train | Val | Scripts / notes |
|---|---|---|---|
| de-dh-unibe | 14309 | 80 | German manuscripts |
| ar-Bulac | 7461 | 394 | Arabic |
| he-Sofer_Mahir | 5474 | 289 | Hebrew |
| se-Riksarkivet_page | 3861 | 217 | Swedish archival |
| cu-Generic_Church_Slavonic_Handwriting | 2638 | 80 | Church Slavonic |
| smp-private | 2143 | 113 | Samaritan (private) |
| la-cmmhwr26 | 2052 | 275 | Latin |
| eu-cBAD-ICDAR2017 | 1859 | 98 | cBAD 2017 |
| eu-cBAD-ICDAR2019 | 1398 | 74 | cBAD 2019 (subset in pretrain mix) |
| syr-MIDRASH | 1055 | 56 | Syriac |
| ar-Muharaf | 907 | 48 | Arabic |
| ar-OpenITI-arabic_print_data | 819 | 45 | Arabic print (ALTO) |
| fi-FCR_500 | 475 | 25 | Finnish |
| de-ICFHR-2016 | 427 | 21 | German (list entry; image availability varies) |
| sa-…Pracalit…Sanskrit_and_Newar… | 417 | 22 | Sanskrit / Newar |
| syr-HTR_Winter_School_2024 | 259 | 14 | Syriac |
| la-HTR_Winter_School_Vienna | 255 | 16 | Latin |
| he-BiblIA | 125 | 7 | Hebrew |
| la-CREMMA | 100 | 21 | Latin |
| + smaller packs | … | … | Greek, Czech, Polish, Judeo-Spanish, Icelandic, Estonian, papyri, etc. |
Important: ICDAR cBAD pages appear inside this pretrain mixture. Stage-0 validation F1 is an internal multiscript holdout, not a fair official cBAD 2019 test score. For apples-to-apples cBAD bake-offs, use a separate Stage-1-cBAD fine-tune and the official test split.
Training recipe
| Hyperparameter | Value |
|---|---|
| Init weights | dfine-l-obj2coco (Peterande/D-FINE) |
| Optimizer | AdamW; base LR (1\times10^{-4}) (backbone (0.1\times)) |
| Schedule | Linear warmup + cosine |
| Precision | bf16-mixed |
| Effective batch | 32 (droplet: micro-batch 32 × accum 1 on MI300X-class) |
| Epochs (cap) | 80; early-stop on cbad_f1_max (patience 12, min 15) |
| Augment | Mild photometric + ±2° rotation (OOB lines dropped) |
| Match distance | 20 px on the 1280 canvas |
| Monitor | cbad_f1_max over conf ∈ {0.1…0.6} |
| Seed | 42 |
Config references: configs/baseline_pretrain.yaml, configs/baseline_pretrain_droplet.yaml.
Training was stopped around epoch 39 (credits / wall-clock), after the monitor had plateaued near 0.89.
Evaluation (Stage-0 val)
Metric: cBAD-style line F1 — Hungarian matching of densified polylines with mean bidirectional Chamfer cost; TP if cost ≤ dist_thresh (20 px).
Released checkpoint
| Metric | Value | Notes |
|---|---|---|
cbad_f1_max |
≈ 0.893 | Best monitor; operating conf typically 0.5 |
| Weights file | best_cbad_f1.safetensors |
Also tagged best_0.8930.safetensors |
| Source | checkpoints/best-epoch=39.ckpt |
Converted with dfine-det convert |
Curves / logs (bundled)
history.csv— TensorBoard export for epochs 1–38 (peak in CSV: 0.8896 @ ep38, conf 0.5)stage0_metrics.png/.pdf— training curves from that CSV- At ep38 @ conf 0.5: precision 0.943, recall 0.842, mean Chamfer ≈ 4.32 px
Fixed conf 0.1 F1 is substantially lower than cbad_f1_max — always sweep (or use ~0.5) before comparing systems.
How to use
Install
pip install -e ".[dev]" # from the dfine-det repo; see QUICKSTART for torch/CUDA notes
Inference (PAGE XML)
dfine-det -d cuda:0 infer \
--weights best_cbad_f1.safetensors \
--image page.jpg \
-o page.xml \
--format page \
--conf 0.5
Fine-tune (Stage-1)
train:
load: best_cbad_f1.safetensors # this Stage-0 Large pretrain
backbone: B4
lr: 5.0e-5
loss_height: 3.0
cost_y: 3.0
monitor: cbad_f1_max
dfine-det --config your_stage1.yaml -d cuda:0 train --augment
Files in this release
| File | Description |
|---|---|
best_cbad_f1.safetensors |
Recommended Stage-0 Large weights |
best_0.8930.safetensors |
Same snapshot, score-tagged filename |
history.csv |
Per-epoch val metrics (epochs 1–38 export) |
stage0_metrics.png |
Summary figure |
README.md |
This model card |
Limitations
- Val F1 is not official cBAD 2019 test performance.
- Pretrain includes cBAD pages → do not treat Stage-0 zero-shot cBAD scores as an external bake-off.
- (K=8) control points underfit strongly curved / torn lines; consider higher (K) in future Stage-0 runs.
- EMA is used for export in training runs; released safetensors are the monitored best snapshot.
- Some listed corpora historically had missing sidecars (e.g. de-ICFHR-2016); compile filtered what resolved on disk.
Citation & credits
Please credit D-FINE (architecture & COCO/Objects365 pretrain)
This work builds on official D-FINE by Yansong Peng et al. (ICLR 2025 Spotlight). Backbone, hybrid encoder, deformable attention, and CDN utilities are adapted from their Apache-2.0 codebase; detection init uses their D-FINE-L Objects365→COCO checkpoint.
@inproceedings{peng2025dfine,
title = {D-FINE: Redefine Regression Task in DETRs as Fine-grained Distribution Refinement},
author = {Peng, Yansong and Li, Hebei and Wu, Peixi and Zhang, Yueyi and Sun, Xiaoyan and Wu, Feng},
booktitle = {The Thirteenth International Conference on Learning Representations},
year = {2025},
url = {https://arxiv.org/abs/2410.13842}
}
- Code: github.com/Peterande/D-FINE
- Weights: huggingface.co/Peterande/D-FINE (
dfine_l_obj2coco_e25.pth)
This Stage-0 polyline adaptation (dfine-det)
@software{dfine_det_stage0_large,
title = {dfine-det Large Stage-0: Multiscript Polyline Baseline Pretrain},
author = {John Locke Jrr, johnlockejrr},
year = {2026},
note = {Initialized from Peterande D-FINE-L; polyline baseline head and PAGE/ALTO training pipeline},
url = {https://huggingface.co/johnlockejrr/dfine-det-large-baseline-stage0}
}
Acknowledgements
We thank AMD for complimentary GPU compute on the AMD Developer Cloud (AMD Instinct™ MI300X), which supported Stage-0 pretraining of this model.
Additional notices
- Vendored D-FINE modules: see
THIRD_PARTY_NOTICES.md(Apache-2.0, © 2024 The D-FINE Authors). - PAGE polygonalization helpers derive from Kraken/BLLA (Apache-2.0, © Benjamin Kiessling) — used at export, not for this pretrain objective.
- Training corpora: respect each dataset’s original license (HTR-United / Zenodo / institutional terms). Private packs (
smp-private,et-private, …) are not redistributed with this card; only model weights trained with them are released if your license allows.
License
Apache License 2.0 for the dfine-det code and these weights, consistent with D-FINE’s Apache-2.0 release. Downstream users must also comply with licenses of any datasets used in further fine-tuning.
Model tree for johnlockejrr/dfine-det-large-baseline-stage0
Paper for johnlockejrr/dfine-det-large-baseline-stage0
Evaluation results
- cbad_f1_max (conf sweep) on Stage-0 multiscript holdout (5% val split)validation set self-reported0.893
- cbad_f1 @ conf=0.5 on Stage-0 multiscript holdout (5% val split)validation set self-reported0.893
- Precision @ conf=0.5 (ep38 history) on Stage-0 multiscript holdout (5% val split)validation set self-reported0.943
- Recall @ conf=0.5 (ep38 history) on Stage-0 multiscript holdout (5% val split)validation set self-reported0.842