Benchmark contamination has a simple worst case: test examples are included in training, then used to measure the same model.
A taxonomy of exposure
Sainz and colleagues' position paper on contamination in NLP evaluation distinguishes levels of benchmark exposure. The clearest case is training on test-split examples and later evaluating on that split. The authors argue this can overestimate performance compared with a model that has not seen the examples, potentially changing which scientific conclusions appear supported.
An unknown rate, not a measured prevalence
The paper offers a taxonomy and a call for community methods; it does not report a new benchmark or estimate how often contamination occurs. Its authors say the extent of the problem is difficult to measure, and propose automatic or semi-automatic exposure detection plus flags for conclusions that could be affected. Detecting text overlap alone would not establish the model's training history or learning process.
Their position paper leaves the central measurement task open: benchmark-specific exposure checks that work when training corpora are incomplete. Until then, published scores can state which exposure information and detection procedure were available for that evaluation.
Bibliography
- Oscar Sainz; Jon Ander Campos; Iker García-Ferrero; Julen Etxaniz; Oier Lopez de Lacalle; Eneko Agirre. (2023). NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark. arXiv:2310.18018v1. https://arxiv.org/abs/2310.18018v1
