Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory Paper • 2607.24368 • Published 23 days ago • 33
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 23 days ago • 37
Beacon: Knowing When and How to Perform Agentic Visual Reasoning Paper • 2607.28595 • Published 20 days ago • 55
HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 21 days ago • 76
Running on Zero Agents Featured 103 Mage-VL 🎞 103 Codec-native video & image understanding with Mage-VL 4B
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published about 1 month ago • 167
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published 24 days ago • 125
GNM Head: A Generative aNthropometric Model of the human head Paper • 2607.23687 • Published 24 days ago • 5