ChenmingWu commited on
Commit
ebd5829
·
verified ·
1 Parent(s): 195056b

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +76 -0
README.md ADDED
@@ -0,0 +1,76 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: other
3
+ license_name: fair-noncommercial-research-derivative
4
+ license_link: https://github.com/facebookresearch/vggt-omega/blob/main/LICENSE
5
+ tags: [3d-gaussian-splatting, novel-view-synthesis, autonomous-driving, vggt, feed-forward]
6
+ ---
7
+
8
+ # MapVGGT — Map-grounded feed-forward 3DGS on a VGGT-Omega backbone (PRIVATE)
9
+
10
+ > **PRIVATE research artifact.** Non-commercial, research-only. See **License** below before any use.
11
+
12
+ MapVGGT is a feed-forward novel-view-synthesis model for driving scenes:
13
+ **VGGT-Omega (1B)** predicts per-pixel **metric depth**; each input pixel is lifted to a
14
+ world-space 3D Gaussian (positions from depth + **known** camera poses); a per-pixel head
15
+ predicts opacity/scale/rotation; the union is rendered with `gsplat`; a small 2D **UNet
16
+ refines** the rendered image. MapGS components (HD-map–anchored tokens, scene-graph
17
+ dynamics, map-depth / free-space losses) are included but — see results — found neutral.
18
+
19
+ ## Honest results (held-out-SCENE, segment-disjoint Waymo, 40 distinct scenes, 256×448, n_in=8)
20
+
21
+ | model | PSNR | SSIM | notes |
22
+ |---|---|---|---|
23
+ | VGGT-Omega + gentle finetune (backbone) | 21.7 | 0.66 | `abl_base_best` |
24
+ | + MAGT map tokens + scene-graph dynamics | 21.7 | 0.66 | `abl_full_best` — **neutral** (ablation) |
25
+ | **+ UNet render-refine** | **22.67** | **0.689** | `mapvggt_refine_best` — **headline** |
26
+
27
+ **Be candid about scope.** This is a **research/system artifact, not SOTA**: ~22.7 dB is
28
+ **~5 dB below** published feed-forward driving NVS (DGGT 27.4, PointForward 28.5, on
29
+ different protocols). Established by clean ablation: the entire gain over a generic
30
+ backbone is **VGGT-Omega + gentle backbone finetuning**; the single extra lever that
31
+ moved the metric is the **UNet refine (+0.85 dB)**. HD-map tokens, scene-graph dynamics,
32
+ higher resolution, multi-view color fusion, uncertainty-shaped covariance, and a skybox
33
+ were all **measured neutral** on this metric (the image-space UNet subsumes them). The
34
+ binding constraint is data scale (1/3 Waymo, ~1157 clips; overfits ~step 1000). Per-clip
35
+ PSNR anti-correlates with view-extrapolation distance (r=-0.57): the model is strong on
36
+ slow/overlapping scenes, weak on fast ego-motion / disocclusion.
37
+
38
+ ## Contents
39
+ - `mapvggt/` — model (`model.py`), heads (`heads.py`: MAGT map tokens, scene-graph dynamics),
40
+ `refine.py` (RefineUNet). `crosscolor.py` / `uncertainty.py` are **experimental, validated
41
+ negative** (kept for the record; not used in training).
42
+ - `mapgs/` — data pipeline (unified clip format, Waymo/AV2 converters), HD-map, losses, metrics.
43
+ - `scripts/` — `train_mapvggt_refine.py` (main trainer), `train_mapvggt_full.py` (map+dyn),
44
+ `eval_mapvggt.py` (canonical loader + held-out eval), data-restore utilities.
45
+ - `checkpoints/` — `mapvggt_refine_best.safetensors` (headline 22.67), `abl_base_best`,
46
+ `abl_full_best`. **Each ~4.6 GB and embeds the finetuned VGGT-Omega 1B backbone** (keys
47
+ `model.vggt.*`, `model.head.*`, `unet.*` for the refine ckpt).
48
+
49
+ ## NOT included (by design)
50
+ - **Base VGGT-Omega weights** (`vggt_omega_1b_512.pt`) — obtain from its FAIR-licensed source;
51
+ set `MAPVGGT_VGGT_CKPT`. (Our refine ckpt already contains a finetuned copy of these weights.)
52
+ - **Training data** — Waymo Open clips (its license **forbids redistribution**) and AV2 clips
53
+ (regenerate with `mapgs/data/convert/*` from your own licensed copies).
54
+ - Vendored clones (`_vggt_omega_repo`, `_tokengs_repo`); clone yourself and set `VGGT_OMEGA_REPO`.
55
+
56
+ ## Usage
57
+ ```bash
58
+ export VGGT_OMEGA_REPO=/path/to/vggt-omega # facebookresearch/vggt-omega clone
59
+ export MAPVGGT_VGGT_CKPT=/path/to/vggt_omega_1b_512.pt # base weights (FAIR-licensed)
60
+ # eval the released checkpoint on a segment-disjoint Waymo val split:
61
+ python -m scripts.eval_mapvggt --ckpt checkpoints/mapvggt_refine_best.safetensors \
62
+ --roots /path/to/data/unified/waymo
63
+ ```
64
+ The refine checkpoint round-trips to 22.67±3.76 / 0.689 via `scripts/eval_mapvggt.py`.
65
+
66
+ ## License & provenance (read before use)
67
+ - **Derivative of VGGT-Omega (Meta FAIR), under the FAIR Noncommercial Research License.**
68
+ The checkpoints contain finetuned VGGT-Omega weights → they inherit FAIR terms:
69
+ **non-commercial, research-only; do not redistribute.** This repo is **PRIVATE** for that reason.
70
+ ⚠️ Commercial use (incl. by a commercial org) is **not permitted** under FAIR terms.
71
+ - Lineage: **TokenGS** (NVIDIA, research-only) — earlier backbone, code under Apache-2.0;
72
+ **Depth-Anything-V2** (Apache-2.0); **PointForward** (scene-graph dynamics formulation).
73
+ - Training data: **Waymo Open Dataset** (subject to Waymo terms, no redistribution) and
74
+ **Argoverse 2** (CC BY-NC-SA 4.0). MapGS code is the authors' own.
75
+
76
+ *Reproducibility note:* gsplat + bf16 make runs reproducible at the seed/config level, not bit-exact.