Jais 2: A Family of Arabic-Centric Open Large Language Models Paper • 2608.13580 • Published Jul 7 • 1
view article Article 🪢 Langfuse and 🤗 Hugging Face: 5 Ways to use them Together MJannik • Mar 14, 2025 • 14
view article Article Welcome Gemma 4: Frontier multimodal intelligence on device +5 merve, pcuenq, sergiopaniego, burtenshaw, Steveeeeeeen, alvarobartt, SaylorTwift • Apr 2 • 922
ResearchGym: Evaluating Language Model Agents on Real-World AI Research Paper • 2602.15112 • Published Feb 16 • 21
view article Article OpenEnv in Practice: Evaluating Tool-Using Agents in Real-World Environments +3 christian-washington, ajasuja, santosh-iima, lewtun, burtenshaw • Feb 12 • 36
view article Article Custom Kernels for All from Codex and Claude +2 burtenshaw, sayakpaul, ariG23498, evalstate • Feb 13 • 81
view article Article We Got Claude to Build CUDA Kernels and teach open models! +2 burtenshaw, evalstate, merve, pcuenq • Jan 28 • 159
view article Article Community Evals: Because we're done trusting black-box leaderboards over the community +5 burtenshaw, SaylorTwift, kramp, merve, davanstrien, nielsr, julien-c • Feb 4 • 90
view article Article Alyah ⭐️: Toward Robust Evaluation of Emirati Dialect Capabilities in Arabic LLMs tiiuae • Jan 27 • 26
view article Article AssetOpsBench: Bridging the Gap Between AI Agent Benchmarks and Industrial Reality ibm-research • Jan 21 • 33
view article Article Open Responses: What you need to know +2 evalstate, burtenshaw, merve, pcuenq • Jan 15 • 113
view article Article NVIDIA Cosmos Reason 2 Brings Advanced Reasoning To Physical AI nvidia • Jan 5 • 64
AgriLLM Collection A collection of the artifacts for the AgriLLM initiative. • 5 items • Updated Dec 15, 2025 • 6
view article Article The Open Evaluation Standard: Benchmarking NVIDIA Nemotron 3 Nano with NeMo Evaluator nvidia • Dec 17, 2025 • 50