A template is more than decoration

A colon, line break, or answer label can leave a task’s meaning intact for a person while changing a model’s result. In Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting, Melanie Sclar and colleagues build a grammar of equivalent prompt formats and measure how model accuracy moves across them.

Their FORMATSPREAD method varies features such as separators, whitespace, casing, and the representation of answer choices. Across the tested settings, LLaMA-2-13B showed a spread of up to 76 accuracy points between formats. The authors report that adding few-shot examples, scaling model size, and instruction tuning did not eliminate the effect. In a separate evaluation, GPT-3.5 was tested across 320 formats on 53 Super-NaturalInstructions tasks; the authors describe the approach as costing less than $10 per task on average.

A range instead of one number

The practical concern is comparison. A format that helps one model need not help another, so a leaderboard built from one arbitrary template can make a model ranking depend on the template choice. Reporting the range across plausible formats would expose that dependency rather than hiding it inside a single score.

The format grammar is manually designed, and some generated combinations are less natural than ordinary prompts. The authors also focus on tasks with relatively short instructions and input fields, leaving longer prompt effects open. Their result is evidence about the formats and tasks they sampled, not a claim that every cosmetic edit changes every answer.

Bibliography

  • Melanie Sclar, Yejin Choi, Yulia Tsvetkov, and Alane Suhr (2023). Quantifying Language Models’ Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formatting. arXiv:2310.11324v2. Full paper (PDF).