HELM maps model behavior across scenarios, then reports how scores shift when the metric changes.
Where the comparison widens
Liang and colleagues' Holistic Evaluation of Language Models (HELM) organizes possible use cases and desiderata into a taxonomy before choosing what to measure. The team evaluated 30 prominent models on 42 scenarios; 21 had not been part of mainstream LM evaluation at the time. Across 16 core scenarios they measured up to seven metrics, including accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency. Seven targeted evaluations covered 26 additional scenarios, including reasoning and disinformation.
A model's profile depends on the metric
The scorecard shows trade-offs across measures: a strong accuracy result can coexist with weaker calibration or higher toxicity. HELM's contribution is to put these outcomes in view together, so comparisons reveal more than one dimension of model behavior.
Missing cells stay visible
The authors identify underrepresented areas, including question answering for neglected English dialects and trustworthiness measures. Metrics were available in 87.5% of the cases across the core scenarios, and the paper describes HELM as a living benchmark. Its scenario selection balances coverage with feasibility; the resulting scorecard represents those choices, not every use or value an application might care about. For a particular deployment, the relevant next step is to select the scenarios and metrics that match its users and costs.
Bibliography
- Percy Liang; Rishi Bommasani; Tony Lee; Dimitris Tsipras; Dilara Soylu; Michihiro Yasunaga; Yian Zhang; Deepak Narayanan; Yuhuai Wu; Ananya Kumar; Benjamin Newman; Binhang Yuan; Bobby Yan; Ce Zhang; Christian Cosgrove; Christopher D. Manning; Christopher Ré; Diana Acosta-Navas; Drew A. Hudson; Eric Zelikman; Esin Durmus; Faisal Ladhak; Frieda Rong; Hongyu Ren; Huaxiu Yao; Jue Wang; Keshav Santhanam; Laurel Orr; Lucia Zheng; Mert Yuksekgonul; Mirac Suzgun; Nathan Kim; Neel Guha; Niladri Chatterji; Omar Khattab; Peter Henderson; Qian Huang; Ryan Chi; Sang Michael Xie; Shibani Santurkar; Surya Ganguli; Tatsunori Hashimoto; Thomas Icard; Tianyi Zhang; Vishrav Chaudhary; William Wang; Xuechen Li; Yifan Mai; Yuhui Zhang; Yuta Koreeda. (2022). Holistic Evaluation of Language Models. arXiv:2211.09110v2. https://arxiv.org/abs/2211.09110v2
