SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 5 days ago • 155
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 4 days ago • 91
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 9 days ago • 277
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Paper • 2608.11752 • Published 10 days ago • 22
How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review Paper • 2608.08975 • Published 13 days ago • 48
SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill Libraries Paper • 2608.05604 • Published 17 days ago • 78
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 22 days ago • 261
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? Paper • 2608.10366 • Published 12 days ago • 10
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Paper • 2608.06374 • Published 17 days ago • 23
MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 23 days ago • 97
WCM: A World Critic Model for Vision-Language-Action Reinforcement Learning Paper • 2607.29613 • Published 23 days ago • 28
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 20 days ago • 156
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 24 days ago • 307