A function name can mislead a code generator

Give a model a function name that conflicts with the written specification and it may follow the name. Jones and Steinhardt use this kind of intervention in Capturing Failures of Large Language Models via Human Cognitive Biases. They borrow ideas from human cognitive-bias research to design tests for recurring errors in generated code.

Their main experiments use OpenAI Codex and CodeGen on HumanEval programming problems, with a smaller set of experiments involving GPT-3. The researchers change the framing around a task, insert anchors resembling flawed solutions, alter example order and set function names against descriptions. Each manipulation gives them a hypothesis about the errors to look for.

From anchors to file deletion

An unrelated function placed before a target task can sharply lower functional correctness. An example resembling a bad solution can draw a model's output toward that anchor. The authors also test how the ordering of unary and binary examples affects generated answers. Together these comparisons show why a single benchmark prompt can miss systematic sensitivities.

One targeted experiment has a more immediate consequence. With sufficiently many package imports in the prompt, Codex can generate code that incorrectly deletes files. The result identifies a failure that an evaluator can try to elicit and inspect. It does not measure the rate of accidental file deletion in deployed coding assistants.

What this test does and does not explain

The experiments use older model versions, a limited programming benchmark and deliberately designed prompts. Smaller GPT-3 tests add less evidence than the main code-generation comparisons. A label such as 'anchoring' describes the observed response to a manipulation here; it makes no claim about human-like psychology inside the model.

For safety testing, the distinction between elicitation and prevalence matters. The paper shows how to build reproducible probes for consequential errors. Estimating how often such errors occur in ordinary use would require a different study.

Bibliography

  • Jones, E., & Steinhardt, J. (2022). Capturing Failures of Large Language Models via Human Cognitive Biases. In Advances in Neural Information Processing Systems 35 (NeurIPS 2022), pp. 11785–11799. https://doi.org/10.52202/068431-0856. arXiv:2202.12299v2 [cs.CL]. Full paper (PDF).