VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 12 days ago • 269
Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published Jul 17 • 44
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 77
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 199
Has GPT-5 Achieved Spatial Intelligence? An Empirical Study Paper • 2508.13142 • Published Aug 18, 2025 • 35
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation Paper • 2510.26794 • Published Oct 30, 2025 • 27
Scaling Spatial Intelligence with Multimodal Foundation Models Paper • 2511.13719 • Published Nov 17, 2025 • 50
Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer Paper • 2603.19227 • Published Mar 19 • 42
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation Paper • 2510.26794 • Published Oct 30, 2025 • 27
ConsistCompose: Unified Multimodal Layout Control for Image Composition Paper • 2511.18333 • Published Nov 23, 2025 • 5
Scaling Spatial Intelligence with Multimodal Foundation Models Paper • 2511.13719 • Published Nov 17, 2025 • 50
Large Motion Model for Unified Multi-Modal Motion Generation Paper • 2404.01284 • Published Apr 1, 2024
SMPLest-X: Ultimate Scaling for Expressive Human Pose and Shape Estimation Paper • 2501.09782 • Published Jan 16, 2025
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos Paper • 2501.13335 • Published Jan 23, 2025
Controllable Human-centric Keyframe Interpolation with Generative Prior Paper • 2506.03119 • Published Jun 3, 2025 • 3
DPoser-X: Diffusion Model as Robust 3D Whole-body Human Pose Prior Paper • 2508.00599 • Published Aug 1, 2025 • 7