Chatbot Arena Estimate: towards a generalized performance benchmark for LLM capabilities.
Lucas Spangher, Tianle Li, William F. Arnold, Nick Masiewicki, Xerxes Dotiwalla, Rama Kumar Pasumarthi, Peter Grabowski, Eugene Ie, Daniel Gruhl
Browse the full NAACL paper archive.