What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks Paper • 2504.07825 • Published Apr 10, 2025 • 1
Faithful or Fabricated? A Causal Framework for Rationalization Bias in LLM Judges Paper • 2605.23970 • Published May 13 • 2
Don't Judge an LLM Only by Its Activations: Discovering Suppressed Safety Features via Counterfactual Activation Potential Paper • 2610.05541 • Published 3 days ago • 2
Heretic Waitlist Collection Models that may or may not eventually be re-ablated with SOTA Heretic methods like SOMPOA • 8 items • Updated 2 days ago