WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces Paper • 2606.09426 • Published Jun 8 • 50
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents Paper • 2606.02031 • Published Jun 1 • 21
OpenForgeRL: Train Harness-native Agents in Any Environment Paper • 2607.21557 • Published Jul 23 • 13
AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Paper • 2608.00155 • Published Jul 31 • 21
XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding Paper • 2608.00036 • Published Jul 21 • 3
ResearchStudio-Reel: Automate the Last Mile of Research from Paper to Poster, Video, and Blog Paper • 2607.04438 • Published Jul 5 • 62
ResearchStudio-Idea: An Evidence-Grounded Research-Ideation Skill Suite from ML Conference Outcomes Paper • 2607.04439 • Published Jul 5 • 64
HealthAgentBench: A Unified Benchmark Suite of Realistic Agentic Healthcare Environments for Challenging Frontier AI Agents Paper • 2606.31179 • Published Jun 30 • 7
AutoSaddler: Automatic Harness Optimization with Durable Updates from Agent Execution Traces Paper • 2608.23041 • Published Aug 24 • 66
CUAWright: A Minimal Unified Interface for Digital Agents Paper • 2610.04116 • Published 6 days ago • 5
Can Computation from Earlier Problems Help LLMs Solve New Ones? Paper • 2609.39394 • Published 8 days ago • 9
Efficient Reasoning Training Does Not Always Harm CoT Faithfulness and Monitorability Paper • 2610.03509 • Published 6 days ago • 15
Equal Ranking Quality, Different Decisions: Measuring and Reducing Order Dependence in LLM Scorers Paper • 2608.26762 • Published 12 days ago • 22
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes Paper • 2610.03715 • Published 6 days ago • 29
TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization Paper • 2605.21318 • Published 3 days ago • 8
What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document Paper • 2610.04128 • Published 6 days ago • 9
MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining Paper • 2609.35652 • Published 10 days ago • 6
DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents Paper • 2610.02320 • Published 7 days ago • 11