Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 5 days ago • 56
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? Paper • 2609.00787 • Published 19 days ago • 19
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? Paper • 2609.00787 • Published 19 days ago • 19
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 17 days ago • 98
OS-MAP: How Far Can Computer-Using Agents Go in Breadth and Depth? Paper • 2507.19132 • Published Jul 25, 2025
CPMobius: Iterative Coach-Player Reasoning for Data-Free Reinforcement Learning Paper • 2602.02979 • Published Feb 3 • 1
StudyBench: Can Self-Evolution Squeeze Textbooks for Olympiad Capability? Paper • 2609.00787 • Published 19 days ago • 19
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 144
Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe Paper • 2604.13016 • Published Apr 14 • 116