SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation Paper • 2608.02287 • Published 10 days ago • 31
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published 7 days ago • 92
CAST: Game Solvers as Turn-Level Teachers for LLM Agents Paper • 2607.25308 • Published 16 days ago • 41
SkillRise: Agentic Reinforcement Learning for Cross-Task Skill Evolution Paper • 2607.26784 • Published 15 days ago • 28
MPIE-Bench: Benchmarking Anatomically Plausible Multi-Person Interaction Editing Paper • 2607.27616 • Published 14 days ago • 39
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published 14 days ago • 34
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published 14 days ago • 50
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 24 days ago • 61
Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation Paper • 2607.05382 • Published Jul 9 • 88
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 85
SynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding Paper • 2607.10400 • Published Jul 11 • 71
EnvScaler: Scaling Tool-Interactive Environments for LLM Agent via Programmatic Synthesis Paper • 2601.05808 • Published Jan 9 • 38
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published Jul 9 • 77
AgenticSTS: A Bounded-Memory Testbed for Long-Horizon LLM Agents Paper • 2607.02255 • Published Jul 2 • 68
OpenThinker-Agent2 Collection OpenThinker-Agent2: agentic SFT/RL datasets and 8B/32B models (cold-start SFT, RL, and the OpenThinkerAgent-32B release). • 11 items • Updated Jun 11 • 10
DomainShuttle: Freeform Open Domain Subject-driven Text-to-video Generation Paper • 2606.26058 • Published Jun 24 • 67
FastContext: Training Efficient Repository Explorer for Coding Agents Paper • 2606.14066 • Published Jun 12 • 95
Rethinking Continual Experience Internalization for Self-Evolving LLM Agents Paper • 2606.04703 • Published Jun 3 • 26