FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow Paper • 2609.03563 • Published 4 days ago • 19
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning Paper • 2608.27549 • Published 11 days ago • 50
GigaBrain-0.7: Scaling Embodied Foundation Models to Emergent Capabilities with a Three-System Architecture Paper • 2608.15875 • Published 22 days ago • 103
Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs Paper • 2608.12781 • Published 21 days ago • 35
Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 14 days ago • 11
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 18 days ago • 274
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published 19 days ago • 52
Embodied-Navigator: Point, Think, Memorize, and Align for Efficient Navigation Paper • 2608.17512 • Published 20 days ago • 51
Demystifying Agent Skills: Why They Work-Until They Don't Paper • 2608.14036 • Published 24 days ago • 169
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published 21 days ago • 340
PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment Paper • 2608.14284 • Published 24 days ago • 14
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published 25 days ago • 58
Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence Paper • 2608.12743 • Published 25 days ago • 44
Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures Paper • 2607.28802 • Published Jul 30 • 10
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation Paper • 2608.02287 • Published Aug 3 • 31
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published Aug 3 • 183