Sampling before submission
AlphaCode did not submit its first program for a problem and stop. Li and colleagues describe a pipeline that generates many candidate programs, filters them using example tests, and clusters the survivors before selecting a small set. In Competition-Level Code Generation with AlphaCode, this search-and-filter process is central to the evaluation.
The model was pretrained on selected GitHub code and fine-tuned on CodeContests, a curated set of competitive-programming problems.
Two contest evaluations
On held-out CodeContests problems, the best system solved 34.2% using at most ten submissions per problem. In simulated participation across ten recent Codeforces contests, each with more than 5,000 participants, AlphaCode ranked within the top 54.3% on average.
What those contests leave untested
These are competitive-programming tasks with formal problem statements and testable solutions. The paper studies sensitivity to problem descriptions and metadata and discusses possible copying from training data. It uses temporal splits and adds tests to reduce test false positives and training overlap, but the contest results cover only ten contests. Performance on changing requirements and multi-file maintenance remains outside these evaluations.
Bibliography
- Li, Y., Choi, D., Chung, J., Kushman, N., Schrittwieser, J., Leblond, R., Eccles, T., Keeling, J., Gimeno, F., Lago, A. D., Hubert, T., Choy, P., d'Autume, C. d. M., Babuschkin, I., Chen, X., Huang, P., Welbl, J., Gowal, S., Cherepanov, A., Molloy, J., Mankowitz, D. J., Robson, E. S., Kohli, P., Freitas, N. d., Kavukcuoglu, K., & Vinyals, O. (2022). Competition-Level Code Generation with AlphaCode. arXiv:2203.07814v1 [cs.NE]. https://arxiv.org/pdf/2203.07814
