Add community evaluation results for GPQA, HLE, DEEP-SWE, APEX-AGENTS

#37
by nielsr HF Staff - opened

This PR adds community-provided evaluation results for the following benchmarks:

These results were extracted from the model card. This is based on the new evaluation results feature.

Note: This is an automated PR. Please review the evaluation results before merging.

@bigeagle this PR is required to enable eval results on the right hand side of the model repo, it also makes a model show up on our leaderboards
On Qwen3.6 repo for instance right hand side you can see it, the model also shows up on associated leaderboard. we use your results directly
Screenshot 2026-07-27 at 18.32.28

Screenshot 2026-07-27 at 18.31.51

bigeagle changed pull request status to merged

Sign up or log in to comment