Skip to content

MazeEval: A Benchmark for Testing Sequential Decision-Making in Language Models.

Hafsteinn Einarsson

VenueBLREC
Year2026
ProceedingsLREC

Browse the full LREC paper archive.