GPT-4 cleared average test-taker scores on several exams in AGIEval, while remaining below top human performance.
The human reference
Zhong and colleagues compare foundation models with scores from public admissions and qualification exams in AGIEval. Under zero-shot chain-of-thought prompting, GPT-4 exceeded average human performance on the SAT, LSAT, and math competitions. The paper reports 95% accuracy on SAT Math and 92.5% on the English test of the Chinese college entrance exam; a gap with top test-takers remained.
Twenty tasks, several prompting setups
The benchmark includes 20 language-only tasks from tests such as the SAT, Gaokao, law-school admissions, math competitions, lawyer qualification, and civil-service exams. The authors evaluate GPT-4, ChatGPT, Text-Davinci-003, and Vicuna using few-shot, zero-shot, and chain-of-thought settings. They also find weaker performance on complex reasoning and domain-specific items, including analytical reasoning, physics, law, and chemistry.
Reading the score in context
AGIEval's public exams include Chinese and English materials, but results remain tied to those exams and prompt conditions. The benchmark's language-only version cannot represent non-text assessment tasks. In the paper's results, GPT-4's relative standing changes with the comparison group and exam: above-average results coexist with a gap to top performers and weaker scores in specialized areas.
Bibliography
- Wanjun Zhong; Ruixiang Cui; Yiduo Guo; Yaobo Liang; Shuai Lu; Yanlin Wang; Amin Saied; Weizhu Chen; Nan Duan. (2023). AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models. arXiv:2304.06364v2. https://arxiv.org/abs/2304.06364v2
