Upload README.md with huggingface_hub
Browse files
README.md
ADDED
|
@@ -0,0 +1,76 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: fair-noncommercial-research-derivative
|
| 4 |
+
license_link: https://github.com/facebookresearch/vggt-omega/blob/main/LICENSE
|
| 5 |
+
tags: [3d-gaussian-splatting, novel-view-synthesis, autonomous-driving, vggt, feed-forward]
|
| 6 |
+
---
|
| 7 |
+
|
| 8 |
+
# MapVGGT — Map-grounded feed-forward 3DGS on a VGGT-Omega backbone (PRIVATE)
|
| 9 |
+
|
| 10 |
+
> **PRIVATE research artifact.** Non-commercial, research-only. See **License** below before any use.
|
| 11 |
+
|
| 12 |
+
MapVGGT is a feed-forward novel-view-synthesis model for driving scenes:
|
| 13 |
+
**VGGT-Omega (1B)** predicts per-pixel **metric depth**; each input pixel is lifted to a
|
| 14 |
+
world-space 3D Gaussian (positions from depth + **known** camera poses); a per-pixel head
|
| 15 |
+
predicts opacity/scale/rotation; the union is rendered with `gsplat`; a small 2D **UNet
|
| 16 |
+
refines** the rendered image. MapGS components (HD-map–anchored tokens, scene-graph
|
| 17 |
+
dynamics, map-depth / free-space losses) are included but — see results — found neutral.
|
| 18 |
+
|
| 19 |
+
## Honest results (held-out-SCENE, segment-disjoint Waymo, 40 distinct scenes, 256×448, n_in=8)
|
| 20 |
+
|
| 21 |
+
| model | PSNR | SSIM | notes |
|
| 22 |
+
|---|---|---|---|
|
| 23 |
+
| VGGT-Omega + gentle finetune (backbone) | 21.7 | 0.66 | `abl_base_best` |
|
| 24 |
+
| + MAGT map tokens + scene-graph dynamics | 21.7 | 0.66 | `abl_full_best` — **neutral** (ablation) |
|
| 25 |
+
| **+ UNet render-refine** | **22.67** | **0.689** | `mapvggt_refine_best` — **headline** |
|
| 26 |
+
|
| 27 |
+
**Be candid about scope.** This is a **research/system artifact, not SOTA**: ~22.7 dB is
|
| 28 |
+
**~5 dB below** published feed-forward driving NVS (DGGT 27.4, PointForward 28.5, on
|
| 29 |
+
different protocols). Established by clean ablation: the entire gain over a generic
|
| 30 |
+
backbone is **VGGT-Omega + gentle backbone finetuning**; the single extra lever that
|
| 31 |
+
moved the metric is the **UNet refine (+0.85 dB)**. HD-map tokens, scene-graph dynamics,
|
| 32 |
+
higher resolution, multi-view color fusion, uncertainty-shaped covariance, and a skybox
|
| 33 |
+
were all **measured neutral** on this metric (the image-space UNet subsumes them). The
|
| 34 |
+
binding constraint is data scale (1/3 Waymo, ~1157 clips; overfits ~step 1000). Per-clip
|
| 35 |
+
PSNR anti-correlates with view-extrapolation distance (r=-0.57): the model is strong on
|
| 36 |
+
slow/overlapping scenes, weak on fast ego-motion / disocclusion.
|
| 37 |
+
|
| 38 |
+
## Contents
|
| 39 |
+
- `mapvggt/` — model (`model.py`), heads (`heads.py`: MAGT map tokens, scene-graph dynamics),
|
| 40 |
+
`refine.py` (RefineUNet). `crosscolor.py` / `uncertainty.py` are **experimental, validated
|
| 41 |
+
negative** (kept for the record; not used in training).
|
| 42 |
+
- `mapgs/` — data pipeline (unified clip format, Waymo/AV2 converters), HD-map, losses, metrics.
|
| 43 |
+
- `scripts/` — `train_mapvggt_refine.py` (main trainer), `train_mapvggt_full.py` (map+dyn),
|
| 44 |
+
`eval_mapvggt.py` (canonical loader + held-out eval), data-restore utilities.
|
| 45 |
+
- `checkpoints/` — `mapvggt_refine_best.safetensors` (headline 22.67), `abl_base_best`,
|
| 46 |
+
`abl_full_best`. **Each ~4.6 GB and embeds the finetuned VGGT-Omega 1B backbone** (keys
|
| 47 |
+
`model.vggt.*`, `model.head.*`, `unet.*` for the refine ckpt).
|
| 48 |
+
|
| 49 |
+
## NOT included (by design)
|
| 50 |
+
- **Base VGGT-Omega weights** (`vggt_omega_1b_512.pt`) — obtain from its FAIR-licensed source;
|
| 51 |
+
set `MAPVGGT_VGGT_CKPT`. (Our refine ckpt already contains a finetuned copy of these weights.)
|
| 52 |
+
- **Training data** — Waymo Open clips (its license **forbids redistribution**) and AV2 clips
|
| 53 |
+
(regenerate with `mapgs/data/convert/*` from your own licensed copies).
|
| 54 |
+
- Vendored clones (`_vggt_omega_repo`, `_tokengs_repo`); clone yourself and set `VGGT_OMEGA_REPO`.
|
| 55 |
+
|
| 56 |
+
## Usage
|
| 57 |
+
```bash
|
| 58 |
+
export VGGT_OMEGA_REPO=/path/to/vggt-omega # facebookresearch/vggt-omega clone
|
| 59 |
+
export MAPVGGT_VGGT_CKPT=/path/to/vggt_omega_1b_512.pt # base weights (FAIR-licensed)
|
| 60 |
+
# eval the released checkpoint on a segment-disjoint Waymo val split:
|
| 61 |
+
python -m scripts.eval_mapvggt --ckpt checkpoints/mapvggt_refine_best.safetensors \
|
| 62 |
+
--roots /path/to/data/unified/waymo
|
| 63 |
+
```
|
| 64 |
+
The refine checkpoint round-trips to 22.67±3.76 / 0.689 via `scripts/eval_mapvggt.py`.
|
| 65 |
+
|
| 66 |
+
## License & provenance (read before use)
|
| 67 |
+
- **Derivative of VGGT-Omega (Meta FAIR), under the FAIR Noncommercial Research License.**
|
| 68 |
+
The checkpoints contain finetuned VGGT-Omega weights → they inherit FAIR terms:
|
| 69 |
+
**non-commercial, research-only; do not redistribute.** This repo is **PRIVATE** for that reason.
|
| 70 |
+
⚠️ Commercial use (incl. by a commercial org) is **not permitted** under FAIR terms.
|
| 71 |
+
- Lineage: **TokenGS** (NVIDIA, research-only) — earlier backbone, code under Apache-2.0;
|
| 72 |
+
**Depth-Anything-V2** (Apache-2.0); **PointForward** (scene-graph dynamics formulation).
|
| 73 |
+
- Training data: **Waymo Open Dataset** (subject to Waymo terms, no redistribution) and
|
| 74 |
+
**Argoverse 2** (CC BY-NC-SA 4.0). MapGS code is the authors' own.
|
| 75 |
+
|
| 76 |
+
*Reproducibility note:* gsplat + bf16 make runs reproducible at the seed/config level, not bit-exact.
|