AISCHOLAR / BLOG
Research on the
fascinating minds of LLMs.
Notes on evidence, experimental design, and the changing craft of studying intelligent systems.
Share
LATEST WRITING01 / BLOG

Chatbot Arena ranks models through anonymous head-to-head votes
The live evaluation platform collected more than 240,000 votes and compared crowd preferences with expert ratings.
Read the article ↗02October 1, 2026
HELM scores models across scenarios, not a single leaderboard
A 30-model evaluation applies multiple metrics across 42 language-model scenarios and makes gaps in benchmark coverage explicit.03September 30, 2026Benchmark contamination can invalidate a clean-looking score
A position paper distinguishes forms of benchmark exposure and argues for explicit contamination checks and warnings.04September 29, 2026AlphaCode searched thousands of programs for contest solutions
The system sampled, filtered, and clustered code; in ten Codeforces contests it ranked within the top 54.3% on average.05September 28, 2026AGIEval compares foundation models with public human exams
Across 20 standardized-exam tasks, GPT-4 beat average human scores in several settings but lagged top human performance.GOOD RESEARCH STARTS WITH BETTER QUESTIONS.See research in practice ↗
About AIScholar
Explore AIScholar ↗Research LLMs with confidence.
AIScholar is a research platform for studying LLMs as research subjects. Design experiments, collect responses, and analyze results in one place—so it’s easier to conduct rigorous, transparent studies.