Latent-Foresight (Cityscapes + nuScenes + CoVLA): End-to-End Learning Predictable Representations for Latent World Models
Latent-Foresight learns a latent world model end to end: an autoencoder compresses frozen DINOv2 features into a compact latent space, and a diffusion transformer (DiT), trained with a flow-matching (rectified flow) objective, predicts the latent of future frames. The autoencoder and the predictor are trained jointly, so the latent space is shaped to be predictable.
This repository contains the scaled-up model, trained on the combination of Cityscapes, nuScenes and CoVLA instead of Cityscapes alone. Scaling the training data improves performance over the Cityscapes-only model.
Paper
Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models
Code
The official implementation can be found on GitHub: https://github.com/Sta8is/Latent-Foresight
Checkpoint
| File | Description |
|---|---|
latent_foresight_csnscovla.ckpt |
Latent-Foresight trained on Cityscapes + nuScenes + CoVLA, high resolution (448x896) |
Sample Usage
The feature statistics (dinov2_stats.pth) and the DPT heads (head_segm.pth, head_depth.pth, head_normals.pth) are not duplicated here. Download them from Sta8is/Latent-Foresight_cs:
hf download Sta8is/Latent-Foresight_csnscovla --local-dir checkpoints
hf download Sta8is/Latent-Foresight_cs dinov2_stats.pth head_segm.pth head_depth.pth head_normals.pth --local-dir checkpoints
For installation, training and evaluation (use --ckpt checkpoints/latent_foresight_csnscovla.ckpt), refer to the GitHub repository.
Citation
If you found Latent-Foresight useful in your research, please consider starring ⭐ us on GitHub and citing 📚 us in your research!
@article{karypidis2026latent,
title={Latent-Foresight: End-to-End Learning Predictable Representations for Latent World Models},
author={Karypidis, Efstathios and Gidaris, Spyros and Komodakis, Nikos},
journal={arXiv preprint arXiv:2610.01942},
year={2026}
}
Acknowledgements
Our code is partially based on Dino-Foresight. We also thank the authors of DINOv2 and DPT for their work and open-source code.
