A battery of strategic choices
Yutong Xie and colleagues benchmark five chatbot families on behavioral-economics games that probe trust, fairness, cooperation, and risk in How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games. The battery includes prisoner’s dilemma, trust, ultimatum, dictator, public-goods, and bomb-risk tasks. The authors compare choice distributions with human-player behavior, then estimate payoff preferences and consistency across games.
Concentrated choices, different payoff patterns
The tested chatbots often produced concentrated choice distributions, while human choices were more varied. In the payoff analysis, models put more weight on fairness than the human comparison data. Inferred preferences also varied by game: a pattern in one setting did not always carry through another. These are measured response patterns on the benchmark, rather than observations of private motives.
The authors report behavior shifts across model versions and sizes, so version-specific measurement matters when interpreting a game profile. The suite’s clearest use is comparative: showing where each tested chatbot’s choices align with or depart from the human distributions.
Bibliography
- Xie, Y., Liu, Y., Ma, Z., Shi, L., Wang, X., Yuan, W., Jackson, M. O., & Mei, Q. (2024). How Different AI Chatbots Behave? Benchmarking Large Language Models in Behavioral Economics Games. arXiv:2412.12362v1 [cs.CL]. Full paper (PDF).
