Running 229 The ultimate guide to multi-harness RL 🔀 229 Train open models with RL inside real agent harnesses
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 24 days ago • 138
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning Paper • 2608.27549 • Published Aug 27 • 48
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 144
view article Article Welcome Inkling by Thinking Machines +3 burtenshaw, merve, pcuenq, ariG23498, andito • Jul 15 • 167