metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language Models.
Alexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric Schulz
Browse the full ICLR paper archive.
Alexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric Schulz
Browse the full ICLR paper archive.