Hongyang-Du commited on
Commit
56ea4be
·
verified ·
1 Parent(s): 3f6f7c5

Document available checkpoints, provenance checks, and paper-reported metrics

Browse files
README.md CHANGED
@@ -14,14 +14,14 @@ datasets:
14
  - ILSVRC/imagenet-1k
15
  ---
16
 
17
- # FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
18
 
19
  Official checkpoints for **FuseReg**. FuseReg replaces heuristic encoder-layer
20
  fusion in representation autoencoders (RAEs) with training over random subsets of
21
  encoder layers, so a single pixel decoder reconstructs reliably under any layer-subset
22
  fusion and pairs better with the diffusion generator.
23
 
24
- **Paper:** coming soon · **Code:** coming soon
25
 
26
  ## Repository layout
27
 
@@ -30,25 +30,47 @@ dinov3-vitl/ # DINOv3 ViT-L encoder, ImageNet 256x256
30
  decoder_k23/
31
  p0.00.safetensors # reproduced RAEv2 decoder (no layer drop)
32
  p0.05.safetensors ... p0.95.safetensors
 
 
 
33
  ditxl_k23/
34
  raev2_pdit0.0_ep040.safetensors
35
  raev2_pdit0.0_ep080.safetensors
36
  fusereg_pdit0.3_ep040.safetensors ... fusereg_pdit0.9_ep040.safetensors
37
- eupe-vitb/ # coming soon
38
- siglip/ # coming soon
39
  ```
40
 
41
- All files contain **EMA weights only** in `safetensors` format (no optimizer,
42
- discriminator, or encoder weights). Each file's `safetensors` metadata records
43
- `encoder`, `p`, `epoch`, `step`, and the encoder `layers` the model was trained on.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
44
 
45
  ## DINOv3 ViT-L
46
 
47
  ### Pixel decoders (`dinov3-vitl/decoder_k23/`)
48
 
49
  ViT decoder (hidden size 1152, 415.6M params), trained for 16 epochs on ImageNet-256
50
- with pixel, LPIPS, adversarial, and classifier-surrogate losses, over DINOv3-L layers
51
- 1-23 with random layer-drop rate `p`. Reconstruction on ImageNet val (50k) under three
52
  inference-time fusions:
53
 
54
  | File | p | feed k=7 PSNR / SSIM / rFID | feed k=23 PSNR / SSIM / rFID | feed l11 PSNR / SSIM / rFID |
@@ -62,6 +84,14 @@ inference-time fusions:
62
  | `p0.90` | 0.9 | 23.62 / 0.671 / 0.610 | 27.60 / 0.827 / 0.415 | 24.82 / 0.722 / 0.458 |
63
  | **`p0.95`** (recommended) | 0.95 | **23.77 / 0.678 / 0.604** | 27.52 / 0.826 / 0.421 | **25.13 / 0.735 / 0.455** |
64
 
 
 
 
 
 
 
 
 
65
  ### DiT-XL generators (`dinov3-vitl/ditxl_k23/`)
66
 
67
  Class-conditional ImageNet-256 DiT-XL (hidden size 1440, 875.3M params) generating in
@@ -79,7 +109,7 @@ the k=23 DINOv3-L latent.
79
  FuseReg DiT-XL generators are trained with random layer-drop rate `p_dit` over the same
80
  k=23 latent; the full `p_dit` x `p_dec` grid is in the paper appendix (DiT-XL rate sweep).
81
 
82
- Pairing a fixed RAEv2 k=23 generator with the FuseReg `p0.95` decoder reduces
83
  unguided gFID from 3.01 to 2.21 while keeping guided gFID at 1.25 (internal guidance
84
  1.78, 50k samples, 50 Euler steps).
85
 
@@ -94,14 +124,22 @@ state_dict = load_file(path)
94
  decoder.load_state_dict(state_dict) # decoder built from the FuseReg codebase
95
  ```
96
 
97
- The encoder is not included; obtain DINOv3 ViT-L from
 
 
 
 
 
 
 
 
98
  [Meta](https://github.com/facebookresearch/dinov3) under the DINOv3 License.
99
 
100
  ## Citation
101
 
102
  ```bibtex
103
  @article{du2026fusereg,
104
- title = {FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders},
105
  author = {Du, Hongyang and Ye, Junjie and Yang, Jiawei and Xie, Yunfei and Cong, Xiaoyan and Zhang, Haodong and Huang, Yongchao and Li, Zongxia and Kadav, Asim and Wei, Chen and Balestriero, Randall and Wang, Yue},
106
  year = {2026}
107
  }
 
14
  - ILSVRC/imagenet-1k
15
  ---
16
 
17
+ # Regularizing Layer Fusion Closes the Reconstruction-Generation Gap in Representation Autoencoders
18
 
19
  Official checkpoints for **FuseReg**. FuseReg replaces heuristic encoder-layer
20
  fusion in representation autoencoders (RAEs) with training over random subsets of
21
  encoder layers, so a single pixel decoder reconstructs reliably under any layer-subset
22
  fusion and pairs better with the diffusion generator.
23
 
24
+ **Paper:** coming soon · **Code:** [Hongyang-Du/FuseReg](https://github.com/Hongyang-Du/FuseReg)
25
 
26
  ## Repository layout
27
 
 
30
  decoder_k23/
31
  p0.00.safetensors # reproduced RAEv2 decoder (no layer drop)
32
  p0.05.safetensors ... p0.95.safetensors
33
+ decoder_k7/
34
+ p0.60.safetensors
35
+ p0.90.safetensors
36
  ditxl_k23/
37
  raev2_pdit0.0_ep040.safetensors
38
  raev2_pdit0.0_ep080.safetensors
39
  fusereg_pdit0.3_ep040.safetensors ... fusereg_pdit0.9_ep040.safetensors
40
+ eupe-vitb/ # no checkpoints published yet
41
+ siglip/ # no checkpoints published yet
42
  ```
43
 
44
+ The repository currently contains **16 inference checkpoints**: eight DINOv3-L
45
+ k=23 decoders, two DINOv3-L k=7 decoders, and six DiT-XL generators. The k=7
46
+ `p=0.3`, four EUPE, and four SigLIP variants have not been published. A 25-file
47
+ candidate collection is tracked in the checkpoint manifest; only entries in its
48
+ `files` array are available for download.
49
+
50
+ [Checkpoint manifest](checkpoint-manifest.json) records file size, SHA-256,
51
+ selected EMA state, and the scope of validation. Fourteen pre-existing files have
52
+ been checked against source checkpoint epoch/step and EMA tensor names/shapes.
53
+ The k=23 `p=0.95` source archive also matches the paper provenance SHA-256, and all
54
+ 456 EMA tensors equal the published file. Both newly published k=7 decoders were
55
+ checked against Drive CRC32C, round-trip tensor equality, and Hub SHA-256.
56
+
57
+ All checkpoint files contain **EMA weights only** in `safetensors` format (no optimizer,
58
+ discriminator, or encoder weights). The `safetensors` metadata records available training provenance such as
59
+ `encoder`, drop rate, `epoch`, and `step`; some older files do not include `layers`.
60
+ Use the matching evaluation configuration and checkpoint manifest for fusion layers.
61
+
62
+ All numerical results below are **paper-reported**, not independently reproduced
63
+ during checkpoint packaging. File integrity and tensor equality checks do not
64
+ validate PSNR, SSIM, rFID, or gFID. Full metric reproduction requires the exact
65
+ ImageNet evaluation split, encoder checkpoints, preprocessing, and latent statistics.
66
 
67
  ## DINOv3 ViT-L
68
 
69
  ### Pixel decoders (`dinov3-vitl/decoder_k23/`)
70
 
71
  ViT decoder (hidden size 1152, 415.6M params), trained for 16 epochs on ImageNet-256
72
+ with pixel, LPIPS, and adversarial losses, over DINOv3-L layers
73
+ 1-23 with random layer-drop rate `p`. Paper-reported reconstruction on ImageNet val (50k) under three
74
  inference-time fusions:
75
 
76
  | File | p | feed k=7 PSNR / SSIM / rFID | feed k=23 PSNR / SSIM / rFID | feed l11 PSNR / SSIM / rFID |
 
84
  | `p0.90` | 0.9 | 23.62 / 0.671 / 0.610 | 27.60 / 0.827 / 0.415 | 24.82 / 0.722 / 0.458 |
85
  | **`p0.95`** (recommended) | 0.95 | **23.77 / 0.678 / 0.604** | 27.52 / 0.826 / 0.421 | **25.13 / 0.735 / 0.455** |
86
 
87
+ ### Pixel decoders trained on k=7 (`dinov3-vitl/decoder_k7/`)
88
+
89
+ Published variants: `p0.60.safetensors` and `p0.90.safetensors`. Their source
90
+ checkpoint metadata identifies the seven encoder layers as
91
+ `[11, 13, 15, 17, 19, 21, 23]`; both were saved at epoch 16. These are distinct
92
+ from feeding a k=23-trained decoder with the same seven-layer subset. No new
93
+ reconstruction metrics are claimed for the packaged files.
94
+
95
  ### DiT-XL generators (`dinov3-vitl/ditxl_k23/`)
96
 
97
  Class-conditional ImageNet-256 DiT-XL (hidden size 1440, 875.3M params) generating in
 
109
  FuseReg DiT-XL generators are trained with random layer-drop rate `p_dit` over the same
110
  k=23 latent; the full `p_dit` x `p_dec` grid is in the paper appendix (DiT-XL rate sweep).
111
 
112
+ The paper reports that pairing a fixed RAEv2 k=23 generator with the FuseReg `p0.95` decoder reduces
113
  unguided gFID from 3.01 to 2.21 while keeping guided gFID at 1.25 (internal guidance
114
  1.78, 50k samples, 50 Euler steps).
115
 
 
124
  decoder.load_state_dict(state_dict) # decoder built from the FuseReg codebase
125
  ```
126
 
127
+ The full RAE checkpoint used for k=23 `p=0.00` stores decoder weights under
128
+ `ema` with a `decoder.` prefix. The published file contains that decoder subtree.
129
+ Its source archive does not record output normalization settings; confirm the
130
+ original RAE evaluation path before using it for a numerical baseline comparison.
131
+ The trained drop decoders and the original RAE baseline must each use the
132
+ appropriate output normalization. The optional CLS-surrogate operation in the
133
+ code is an additive last-layer token mean, not a classifier-surrogate loss.
134
+
135
+ The encoder and latent normalization statistics are not included. Obtain DINOv3 ViT-L from
136
  [Meta](https://github.com/facebookresearch/dinov3) under the DINOv3 License.
137
 
138
  ## Citation
139
 
140
  ```bibtex
141
  @article{du2026fusereg,
142
+ title = {Regularizing Layer Fusion Closes the Reconstruction-Generation Gap in Representation Autoencoders},
143
  author = {Du, Hongyang and Ye, Junjie and Yang, Jiawei and Xie, Yunfei and Cong, Xiaoyan and Zhang, Haodong and Huang, Yongchao and Li, Zongxia and Kadav, Asim and Wei, Chen and Balestriero, Randall and Wang, Yue},
144
  year = {2026}
145
  }
checkpoint-manifest.json ADDED
@@ -0,0 +1,609 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "repo_id": "Hongyang-Du/FuseReg",
3
+ "revision": "3f6f7c58434256c835bad6e79df1b598db5d59a1",
4
+ "format": "safetensors",
5
+ "weight_selection": "EMA inference weights only",
6
+ "note": "Checkpoint integrity checks are not an independent reproduction of paper metrics. Encoder weights and latent normalization statistics are separate dependencies.",
7
+ "present_count": 16,
8
+ "planned_count": 25,
9
+ "blocked_files": [
10
+ {
11
+ "path": "dinov3-vitl/decoder_k7/p0.30.safetensors",
12
+ "status": "blocked",
13
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
14
+ },
15
+ {
16
+ "path": "eupe-vitb/decoder_k11/p0.90.safetensors",
17
+ "status": "blocked",
18
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
19
+ },
20
+ {
21
+ "path": "eupe-vitb/decoder_k11/p0.60.safetensors",
22
+ "status": "blocked",
23
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
24
+ },
25
+ {
26
+ "path": "eupe-vitb/decoder_k11/p0.30.safetensors",
27
+ "status": "blocked",
28
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
29
+ },
30
+ {
31
+ "path": "eupe-vitb/decoder_k11/p0.00.safetensors",
32
+ "status": "blocked",
33
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
34
+ },
35
+ {
36
+ "path": "siglip/decoder_k23/p0.90.safetensors",
37
+ "status": "blocked",
38
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
39
+ },
40
+ {
41
+ "path": "siglip/decoder_k23/p0.60.safetensors",
42
+ "status": "blocked",
43
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
44
+ },
45
+ {
46
+ "path": "siglip/decoder_k23/p0.30.safetensors",
47
+ "status": "blocked",
48
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
49
+ },
50
+ {
51
+ "path": "siglip/decoder_k23/p0.00.safetensors",
52
+ "status": "blocked",
53
+ "reason": "Google Drive download quota exceeded; connected Drive raw-file tool has a 256 MiB file limit."
54
+ }
55
+ ],
56
+ "files": [
57
+ {
58
+ "path": "dinov3-vitl/decoder_k23/p0.00.safetensors",
59
+ "size_bytes": 1662643616,
60
+ "sha256": "6b8cd680a15cfa82e62936bf723dca3d57cbfdb4d7cb5d2b5b7686b901086cf4",
61
+ "encoder": "dinov3-vitl",
62
+ "model": "decoder",
63
+ "p": 0.0,
64
+ "epoch": 16,
65
+ "layers": [
66
+ 1,
67
+ 2,
68
+ 3,
69
+ 4,
70
+ 5,
71
+ 6,
72
+ 7,
73
+ 8,
74
+ 9,
75
+ 10,
76
+ 11,
77
+ 12,
78
+ 13,
79
+ 14,
80
+ 15,
81
+ 16,
82
+ 17,
83
+ 18,
84
+ 19,
85
+ 20,
86
+ 21,
87
+ 22,
88
+ 23
89
+ ],
90
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
91
+ "source_key": "ema.decoder.* (prefix removed)",
92
+ "layers_verification": "Inferred from source filename and existing repository documentation; absent in safetensors metadata.",
93
+ "output_normalization": "unresolved: source checkpoint lacks output normalization configuration; confirm original RAE evaluation path before reporting metrics"
94
+ },
95
+ {
96
+ "path": "dinov3-vitl/decoder_k23/p0.05.safetensors",
97
+ "size_bytes": 1662643712,
98
+ "sha256": "0d43b5477cefa739160103e371594e279733fadc8fc7c41daa5cbda73161ded2",
99
+ "encoder": "dinov3-vitl",
100
+ "model": "decoder",
101
+ "p": 0.05,
102
+ "epoch": 16,
103
+ "layers": [
104
+ 1,
105
+ 2,
106
+ 3,
107
+ 4,
108
+ 5,
109
+ 6,
110
+ 7,
111
+ 8,
112
+ 9,
113
+ 10,
114
+ 11,
115
+ 12,
116
+ 13,
117
+ 14,
118
+ 15,
119
+ 16,
120
+ 17,
121
+ 18,
122
+ 19,
123
+ 20,
124
+ 21,
125
+ 22,
126
+ 23
127
+ ],
128
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
129
+ "source_key": "ema_dec"
130
+ },
131
+ {
132
+ "path": "dinov3-vitl/decoder_k23/p0.10.safetensors",
133
+ "size_bytes": 1662643712,
134
+ "sha256": "8b5f611f6f69b1527ef061214044a93ed5f1a9e7a5e30d07a4b7a38c26c6b00b",
135
+ "encoder": "dinov3-vitl",
136
+ "model": "decoder",
137
+ "p": 0.1,
138
+ "epoch": 16,
139
+ "layers": [
140
+ 1,
141
+ 2,
142
+ 3,
143
+ 4,
144
+ 5,
145
+ 6,
146
+ 7,
147
+ 8,
148
+ 9,
149
+ 10,
150
+ 11,
151
+ 12,
152
+ 13,
153
+ 14,
154
+ 15,
155
+ 16,
156
+ 17,
157
+ 18,
158
+ 19,
159
+ 20,
160
+ 21,
161
+ 22,
162
+ 23
163
+ ],
164
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
165
+ "source_key": "ema_dec"
166
+ },
167
+ {
168
+ "path": "dinov3-vitl/decoder_k23/p0.30.safetensors",
169
+ "size_bytes": 1662643712,
170
+ "sha256": "47ea1d8ab595098c7ef2a1639b65907ace9506e0e3e30165de8107f378685fbc",
171
+ "encoder": "dinov3-vitl",
172
+ "model": "decoder",
173
+ "p": 0.3,
174
+ "epoch": 16,
175
+ "layers": [
176
+ 1,
177
+ 2,
178
+ 3,
179
+ 4,
180
+ 5,
181
+ 6,
182
+ 7,
183
+ 8,
184
+ 9,
185
+ 10,
186
+ 11,
187
+ 12,
188
+ 13,
189
+ 14,
190
+ 15,
191
+ 16,
192
+ 17,
193
+ 18,
194
+ 19,
195
+ 20,
196
+ 21,
197
+ 22,
198
+ 23
199
+ ],
200
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
201
+ "source_key": "ema_dec"
202
+ },
203
+ {
204
+ "path": "dinov3-vitl/decoder_k23/p0.50.safetensors",
205
+ "size_bytes": 1662643712,
206
+ "sha256": "4fba21d3a8ae23df7d5ea7bd143de00f885d562678b50521a11b132321f3191a",
207
+ "encoder": "dinov3-vitl",
208
+ "model": "decoder",
209
+ "p": 0.5,
210
+ "epoch": 16,
211
+ "layers": [
212
+ 1,
213
+ 2,
214
+ 3,
215
+ 4,
216
+ 5,
217
+ 6,
218
+ 7,
219
+ 8,
220
+ 9,
221
+ 10,
222
+ 11,
223
+ 12,
224
+ 13,
225
+ 14,
226
+ 15,
227
+ 16,
228
+ 17,
229
+ 18,
230
+ 19,
231
+ 20,
232
+ 21,
233
+ 22,
234
+ 23
235
+ ],
236
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
237
+ "source_key": "ema_dec"
238
+ },
239
+ {
240
+ "path": "dinov3-vitl/decoder_k23/p0.70.safetensors",
241
+ "size_bytes": 1662643712,
242
+ "sha256": "330f76d9775065adee902a5657656af23be13496f6ef97cdf5abba240ba0a253",
243
+ "encoder": "dinov3-vitl",
244
+ "model": "decoder",
245
+ "p": 0.7,
246
+ "epoch": 16,
247
+ "layers": [
248
+ 1,
249
+ 2,
250
+ 3,
251
+ 4,
252
+ 5,
253
+ 6,
254
+ 7,
255
+ 8,
256
+ 9,
257
+ 10,
258
+ 11,
259
+ 12,
260
+ 13,
261
+ 14,
262
+ 15,
263
+ 16,
264
+ 17,
265
+ 18,
266
+ 19,
267
+ 20,
268
+ 21,
269
+ 22,
270
+ 23
271
+ ],
272
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
273
+ "source_key": "ema_dec"
274
+ },
275
+ {
276
+ "path": "dinov3-vitl/decoder_k23/p0.90.safetensors",
277
+ "size_bytes": 1662643712,
278
+ "sha256": "042190ff72f3606dc8c1431878f71fc7aeb73275e698bc32b8708066521fbfa7",
279
+ "encoder": "dinov3-vitl",
280
+ "model": "decoder",
281
+ "p": 0.9,
282
+ "epoch": 16,
283
+ "layers": [
284
+ 1,
285
+ 2,
286
+ 3,
287
+ 4,
288
+ 5,
289
+ 6,
290
+ 7,
291
+ 8,
292
+ 9,
293
+ 10,
294
+ 11,
295
+ 12,
296
+ 13,
297
+ 14,
298
+ 15,
299
+ 16,
300
+ 17,
301
+ 18,
302
+ 19,
303
+ 20,
304
+ 21,
305
+ 22,
306
+ 23
307
+ ],
308
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
309
+ "source_key": "ema_dec"
310
+ },
311
+ {
312
+ "path": "dinov3-vitl/decoder_k23/p0.95.safetensors",
313
+ "size_bytes": 1662643664,
314
+ "sha256": "f82adb05ca36c5bc22630a86c7439a630769bc2dd7c8f2a5e4043f0e38120476",
315
+ "encoder": "dinov3-vitl",
316
+ "model": "decoder",
317
+ "p": 0.95,
318
+ "epoch": 16,
319
+ "layers": [
320
+ 1,
321
+ 2,
322
+ 3,
323
+ 4,
324
+ 5,
325
+ 6,
326
+ 7,
327
+ 8,
328
+ 9,
329
+ 10,
330
+ 11,
331
+ 12,
332
+ 13,
333
+ 14,
334
+ 15,
335
+ 16,
336
+ 17,
337
+ 18,
338
+ 19,
339
+ 20,
340
+ 21,
341
+ 22,
342
+ 23
343
+ ],
344
+ "verification": "Source archive SHA-256 matches paper provenance; all 456 EMA tensors equal existing safetensors; local safetensors SHA-256 equals Hub LFS SHA-256.",
345
+ "source_key": "ema_dec",
346
+ "source_sha256": "49cf7f50310dd836f9d8282966c4ea5c188fbf03cbfe5504a09efe5fcec79a24"
347
+ },
348
+ {
349
+ "path": "dinov3-vitl/decoder_k7/p0.60.safetensors",
350
+ "size_bytes": 1662643816,
351
+ "sha256": "1b627589651d9755ad974f62fe0c54b159016e1a6803a931c1fb0b93d17566d1",
352
+ "encoder": "dinov3-vitl",
353
+ "model": "decoder",
354
+ "p": 0.6,
355
+ "epoch": 16,
356
+ "layers": [
357
+ 11,
358
+ 13,
359
+ 15,
360
+ 17,
361
+ 19,
362
+ 21,
363
+ 23
364
+ ],
365
+ "source_sha256": "adfcae18e0059d98274d60d017b286f09e3681b582cefc190a15edebd58d875b",
366
+ "verification": "Drive source size and CRC32C verified; EMA tensors round-trip equal; uploaded size and SHA-256 match Hub LFS metadata."
367
+ },
368
+ {
369
+ "path": "dinov3-vitl/decoder_k7/p0.90.safetensors",
370
+ "size_bytes": 1662643816,
371
+ "sha256": "25e0ceed420efdb957246e406e2a623974b55f2b8e1bccfdc9ac69d6eeeb6071",
372
+ "encoder": "dinov3-vitl",
373
+ "model": "decoder",
374
+ "p": 0.9,
375
+ "epoch": 16,
376
+ "layers": [
377
+ 11,
378
+ 13,
379
+ 15,
380
+ 17,
381
+ 19,
382
+ 21,
383
+ 23
384
+ ],
385
+ "source_sha256": "6110c7e689703215c19feefeac1dd40db64949fb3c863b9ca1c7f7ccb510160f",
386
+ "verification": "Drive source size and CRC32C verified; EMA tensors round-trip equal; uploaded size and SHA-256 match Hub LFS metadata."
387
+ },
388
+ {
389
+ "path": "dinov3-vitl/ditxl_k23/fusereg_pdit0.3_ep040.safetensors",
390
+ "size_bytes": 3501338104,
391
+ "sha256": "303c28286c6d97aea769cbf88fc767828dd52b6da9ea77f50f40b392cb70cd35",
392
+ "encoder": "dinov3-vitl",
393
+ "model": "dit-xl",
394
+ "p": 0.3,
395
+ "epoch": 40,
396
+ "layers": [
397
+ 1,
398
+ 2,
399
+ 3,
400
+ 4,
401
+ 5,
402
+ 6,
403
+ 7,
404
+ 8,
405
+ 9,
406
+ 10,
407
+ 11,
408
+ 12,
409
+ 13,
410
+ 14,
411
+ 15,
412
+ 16,
413
+ 17,
414
+ 18,
415
+ 19,
416
+ 20,
417
+ 21,
418
+ 22,
419
+ 23
420
+ ],
421
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
422
+ "source_key": "ema",
423
+ "layers_verification": "Inferred from source filename and existing repository documentation; absent in safetensors metadata."
424
+ },
425
+ {
426
+ "path": "dinov3-vitl/ditxl_k23/fusereg_pdit0.5_ep040.safetensors",
427
+ "size_bytes": 3501338104,
428
+ "sha256": "e92e22a1f3fc1b00687d0172fabe123b197c714cbfdb39647cf3cdfaa5a4f94f",
429
+ "encoder": "dinov3-vitl",
430
+ "model": "dit-xl",
431
+ "p": 0.5,
432
+ "epoch": 40,
433
+ "layers": [
434
+ 1,
435
+ 2,
436
+ 3,
437
+ 4,
438
+ 5,
439
+ 6,
440
+ 7,
441
+ 8,
442
+ 9,
443
+ 10,
444
+ 11,
445
+ 12,
446
+ 13,
447
+ 14,
448
+ 15,
449
+ 16,
450
+ 17,
451
+ 18,
452
+ 19,
453
+ 20,
454
+ 21,
455
+ 22,
456
+ 23
457
+ ],
458
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
459
+ "source_key": "ema",
460
+ "layers_verification": "Inferred from source filename and existing repository documentation; absent in safetensors metadata."
461
+ },
462
+ {
463
+ "path": "dinov3-vitl/ditxl_k23/fusereg_pdit0.7_ep040.safetensors",
464
+ "size_bytes": 3501338104,
465
+ "sha256": "48124b905b5040f3e78e6eeb5bfe7aa233bf07518392b0cfe9344fa65a937797",
466
+ "encoder": "dinov3-vitl",
467
+ "model": "dit-xl",
468
+ "p": 0.7,
469
+ "epoch": 40,
470
+ "layers": [
471
+ 1,
472
+ 2,
473
+ 3,
474
+ 4,
475
+ 5,
476
+ 6,
477
+ 7,
478
+ 8,
479
+ 9,
480
+ 10,
481
+ 11,
482
+ 12,
483
+ 13,
484
+ 14,
485
+ 15,
486
+ 16,
487
+ 17,
488
+ 18,
489
+ 19,
490
+ 20,
491
+ 21,
492
+ 22,
493
+ 23
494
+ ],
495
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
496
+ "source_key": "ema",
497
+ "layers_verification": "Inferred from source filename and existing repository documentation; absent in safetensors metadata."
498
+ },
499
+ {
500
+ "path": "dinov3-vitl/ditxl_k23/fusereg_pdit0.9_ep040.safetensors",
501
+ "size_bytes": 3501338104,
502
+ "sha256": "10d1b10d343273de043f43ed1f7bb93ebc58c462ef20c5bbf610f86d6e252ce8",
503
+ "encoder": "dinov3-vitl",
504
+ "model": "dit-xl",
505
+ "p": 0.9,
506
+ "epoch": 40,
507
+ "layers": [
508
+ 1,
509
+ 2,
510
+ 3,
511
+ 4,
512
+ 5,
513
+ 6,
514
+ 7,
515
+ 8,
516
+ 9,
517
+ 10,
518
+ 11,
519
+ 12,
520
+ 13,
521
+ 14,
522
+ 15,
523
+ 16,
524
+ 17,
525
+ 18,
526
+ 19,
527
+ 20,
528
+ 21,
529
+ 22,
530
+ 23
531
+ ],
532
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
533
+ "source_key": "ema",
534
+ "layers_verification": "Inferred from source filename and existing repository documentation; absent in safetensors metadata."
535
+ },
536
+ {
537
+ "path": "dinov3-vitl/ditxl_k23/raev2_pdit0.0_ep040.safetensors",
538
+ "size_bytes": 3501338168,
539
+ "sha256": "47090a0e892af0aa60fb9961f9fbf38873c2dd5b940eb08a5028ad8450e4fd62",
540
+ "encoder": "dinov3-vitl",
541
+ "model": "dit-xl",
542
+ "p": 0.0,
543
+ "epoch": 40,
544
+ "layers": [
545
+ 1,
546
+ 2,
547
+ 3,
548
+ 4,
549
+ 5,
550
+ 6,
551
+ 7,
552
+ 8,
553
+ 9,
554
+ 10,
555
+ 11,
556
+ 12,
557
+ 13,
558
+ 14,
559
+ 15,
560
+ 16,
561
+ 17,
562
+ 18,
563
+ 19,
564
+ 20,
565
+ 21,
566
+ 22,
567
+ 23
568
+ ],
569
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
570
+ "source_key": "ema"
571
+ },
572
+ {
573
+ "path": "dinov3-vitl/ditxl_k23/raev2_pdit0.0_ep080.safetensors",
574
+ "size_bytes": 3501338168,
575
+ "sha256": "e07d5ab7e1479a6f0dc9b3ed1ced797970b36a9d9c1a12c36b0c62719b03d074",
576
+ "encoder": "dinov3-vitl",
577
+ "model": "dit-xl",
578
+ "p": 0.0,
579
+ "epoch": 80,
580
+ "layers": [
581
+ 1,
582
+ 2,
583
+ 3,
584
+ 4,
585
+ 5,
586
+ 6,
587
+ 7,
588
+ 8,
589
+ 9,
590
+ 10,
591
+ 11,
592
+ 12,
593
+ 13,
594
+ 14,
595
+ 15,
596
+ 16,
597
+ 17,
598
+ 18,
599
+ 19,
600
+ 20,
601
+ 21,
602
+ 22,
603
+ 23
604
+ ],
605
+ "verification": "Drive archive and Hub headers agree on epoch, step, and all selected EMA tensor names and shapes; weight values not independently compared.",
606
+ "source_key": "ema"
607
+ }
608
+ ]
609
+ }
dinov3-vitl/README.md CHANGED
@@ -1,6 +1,7 @@
1
  # DINOv3 ViT-L
2
 
3
- - `decoder_k23/`: FuseReg pixel decoders over DINOv3-L layers 1-23, layer-drop rate `p` in {0, 0.05, 0.1, 0.3, 0.5, 0.7, 0.9, 0.95}. `p0.00` is the reproduced RAEv2 decoder.
4
- - `ditxl_k23/`: class-conditional ImageNet-256 DiT-XL generators in the k=23 latent (RAEv2 `p_dit=0` at 40/80 epochs; FuseReg `p_dit` in {0.3, 0.5, 0.7, 0.9} at 40 epochs).
 
5
 
6
- See the top-level README for metrics and usage.
 
1
  # DINOv3 ViT-L
2
 
3
+ - `decoder_k23/`: eight pixel decoders with `p` in {0, 0.05, 0.1, 0.3, 0.5, 0.7, 0.9, 0.95}.
4
+ - `decoder_k7/`: pixel decoders with `p` in {0.6, 0.9}, trained on layers [11, 13, 15, 17, 19, 21, 23]. The p=0.3 variant is not published.
5
+ - `ditxl_k23/`: six class-conditional DiT-XL generators (RAEv2 p=0 at epochs 40/80; FuseReg p in {0.3, 0.5, 0.7, 0.9} at epoch 40).
6
 
7
+ All files contain EMA inference weights only. See the top-level README and checkpoint manifest for validation scope, baseline normalization caveats, and paper-reported results. Latent statistics and encoder weights are separate dependencies.
eupe-vitb/README.md CHANGED
@@ -1,3 +1,3 @@
1
  # EUPE ViT-B
2
 
3
- Coming soon: FuseReg pixel decoders over EUPE ViT-B layers 1-11 (`decoder_k11/`).
 
1
  # EUPE ViT-B
2
 
3
+ No EUPE checkpoints have been published in this repository yet.
siglip/README.md CHANGED
@@ -1,3 +1,3 @@
1
  # SigLIP
2
 
3
- Coming soon: FuseReg pixel decoders over SigLIP layers 1-23 (`decoder_k23/`).
 
1
  # SigLIP
2
 
3
+ No SigLIP checkpoints have been published in this repository yet.