A leaderboard needs more than one wording
A benchmark score usually reflects a model answering questions under one prompt template. Felipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin instead ask what the model’s performance distribution looks like across possible templates in Efficient Multi-Prompt Evaluation of LLMs.
Their PromptEval method estimates that distribution by borrowing information across prompts and examples, so an evaluator need not run every example against every possible wording. They assess it on MMLU, BIG-bench Hard, and LMentry, including an analysis of 100 MMLU prompt templates. The authors report that the method estimated performance quantiles accurately with a budget equivalent to two single-prompt evaluations, compared with evaluating the full prompt pool. Quantiles make it possible to report a median or a high-performing percentile alongside the result for any one chosen prompt.
The prompt pool still needs a reason
PromptEval estimates behavior over a predefined set; it does not decide which templates belong in that set or solve prompt engineering. The authors identify template selection as a remaining challenge. The estimate is also tied to the model, benchmark, prompt pool, and budget used in the evaluation.
A benchmark that uses this approach can show whether a reported score is typical or an outlier within its selected prompt family. The paper does not establish that one percentile is right for every application; it makes the distribution visible so that choice can be made explicitly.
Bibliography
- Felipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva, Onkar Bhardwaj, Leshem Choshen, Allysson Flavio Melo de Oliveira, Yuekai Sun, and Mikhail Yurochkin (2024). Efficient Multi-Prompt Evaluation of LLMs. arXiv:2405.17202v3. Full paper (PDF).
