A live prompt becomes a pairwise vote

Chatbot Arena asks a user to submit a prompt, then presents two anonymous chatbot replies for comparison. In Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference, Chiang and colleagues analyze more than 240,000 votes collected during the platform's first months. Their ranking method uses those preferences to estimate model performance; they also compare crowd votes with expert ratings and report good agreement.

The voters shape the measure

Arena's questions come from participating users, rather than a fixed evaluation set. Conversations in the analyzed data span more than 100 languages, with 77% in English. The vote stream therefore describes choices made by this platform's users and the prompts they supplied. It does not answer whether a preferred response is factually correct or safe; each vote records a participant's choice for a particular exchange. That 77% share is a useful boundary to retain when interpreting Arena's multilingual coverage.

Bibliography

  • Chiang, W., Zheng, L., Sheng, Y., Angelopoulos, A. N., Li, T., Li, D., Zhang, H., Zhu, B., Jordan, M., Gonzalez, J. E., & Stoica, I. (2024). Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. arXiv:2403.04132v1 [cs.AI]. https://arxiv.org/pdf/2403.04132