FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents Paper • 2606.12087 • Published Jun 10 • 79
EvoTrainer: Co-Evolving LLM Policies and Training Harnesses for Autonomous Agentic Reinforcement Learning Paper • 2606.03108 • Published Jun 2 • 12
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research Paper • 2606.07591 • Published May 28 • 102
GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks Paper • 2510.04374 • Published Oct 5, 2025 • 1
PatRe: A Full-Stage Office Action and Rebuttal Generation Benchmark for Patent Examination Paper • 2605.03571 • Published May 5 • 7
InteractWeb-Bench: Can Multimodal Agent Escape Blind Execution in Interactive Website Generation? Paper • 2604.27419 • Published Apr 30 • 13
Beyond Quantity: Trajectory Diversity Scaling for Code Agents Paper • 2602.03219 • Published Feb 3 • 2
Benchmarks and Datasets Collection IP Intelligence Team Works about Benchmarks and Datasets. • 4 items • Updated May 18 • 1
AI for Patents Collection IP Intelligence Team Works about AI for Patents. • 1 item • Updated Feb 2 • 1
FlowPIE Collection Resources of Our Proposed Scientific Idea Generation Algorithm, FlowPIE. • 1 item • Updated Apr 1 • 2
FlowPIE: Test-Time Scientific Idea Evolution with Flow-Guided Literature Exploration Paper • 2603.29557 • Published Mar 31 • 17
WithAnyone: Towards Controllable and ID Consistent Image Generation Paper • 2510.14975 • Published Oct 16, 2025 • 86
Lingma SWE-GPT: An Open Development-Process-Centric Language Model for Automated Software Improvement Paper • 2411.00622 • Published Nov 1, 2024 • 3
Implicit Actor Critic Coupling via a Supervised Learning Framework for RLVR Paper • 2509.02522 • Published Sep 2, 2025 • 26
SWE-Flow: Synthesizing Software Engineering Data in a Test-Driven Manner Paper • 2506.09003 • Published Jun 10, 2025 • 17
Model Merging in Pre-training of Large Language Models Paper • 2505.12082 • Published May 17, 2025 • 40