The LLM-as-research-subject platform

Study the behavior of large language models.

AIScholar brings design, repeated responses, coding and analysis together for supported single-turn text studies. Explore bounded questions relevant to AI safety and alignment, benchmarking and other model behavior research; observed responses do not establish deployment safety.

For single-turn text studies. Running trials requires your OpenRouter key.

The research pipeline

09 STEPS

Plan

  1. 01Hypotheses
  2. 02Design
  3. 03Prompts

Run

  1. 04Execute
  2. 05Collect
  3. 06Code

Analyze

  1. 07Visualize
  2. 08Analyze
  3. 09Write-up
Why researchers choose AIScholar

A simpler way to run a careful study.

  • Ask what changes behavior

    Compare how models respond under different conditions. See the spread of responses, not just an average, and keep the study's limits in view.

  • AI help when you need it

    Use AI wizards to draft designs, prompts, coding rules and report sections. You set the question, check the method and interpret the results. Availability depends on your plan.

  • Check your method as you go

    Warnings and analysis tools can flag design and measurement problems. You still need to validate your coding rules and review the findings.

  • Keep the method inspectable

    Save prompt and coding versions alongside trial settings and results. Others can inspect what you ran, though model and provider changes may affect a rerun.

  • Bring your own data

    Upload responses you collected elsewhere and use AIScholar's coding, charts and analysis tools.

  • Take your work with you

    Export charts, data and a draft report. Check the citations and claims before you share it.

How it works

Your study, step by step.

AIScholar keeps your design, prompts, responses and analysis together. You choose the factors, check the coding rules, review responses and interpret the results.

STAGE 01

Hypotheses

Frame testable predictions about LLM behavior, with AI as a sounding board.

Learn more about Hypotheses

STAGE 02

Design

Build factorial, fractional, matched-pair, or custom designs with discrete factors and levels.

Learn more about Design

STAGE 03

Prompts

Construct prompt templates from a snippet library bound to your factor levels.

Learn more about Prompts

STAGE 04

Execute

Pilot then run single-turn text trials with OpenRouter models; review estimated cost first.

Learn more about Execute

STAGE 05

Collect

Capture every response with full trial metadata for reproducibility.

Learn more about Collect

STAGE 06

Code

Score responses with LLM-as-a-judge, multi-judge agreement, and human review.

Learn more about Code

STAGE 07

Visualize

Explore results through charts, descriptive statistics, and cell coverage.

Learn more about Visualize

STAGE 08

Analyze

Run inferential models, effect sizes, and LLM-specific methodological checks.

Learn more about Analyze

STAGE 09

Write-up

Draft a report for expert checking and export to Word, Markdown, or LaTeX.

Learn more about Write-up

Explore all features in depth →

What's inside

Tools for each part of your study.

See how each tool works, from study design and prompts to coding, analysis and reporting.

  • Methodology knowledge base

    Give the AI wizards curated methods entries as context when you need them.

  • Citation traceability

    A curated citation library tracked per project, with inline references that survive into the write-up.

  • AI Analyst chat

    Ask your coded data questions in plain English, on top of the formal statistical tests.

  • BYOD dataset mode

    Analyze response datasets you collected elsewhere with the same coding and statistics stages.

Why AIScholar is different

How AIScholar fits alongside other tools.

Eval platforms, survey tools and notebooks can all study model behavior. AIScholar is for researchers who want a guided way to design and run single-turn text studies with OpenRouter models, then code and analyze the responses.

vs. LLM eval & observability tools

Eval and observability tools can run datasets, custom scorers and comparable experiments, including behavioral evaluations.

AIScholar offers a guided route through factors, replicated trials, coding, analysis and a write-up draft for supported studies.

vs. generic design-of-experiments software

DOE packages offer deeper design optimization and diagnostics; survey tools support human respondents.

AIScholar connects supported discrete conditions to OpenRouter text trials in one workflow.

vs. academic research code

Custom notebooks offer maximum control, local or private models, and open methods code.

The integrated workflow reduces setup for supported studies, while constraining providers, modalities and turn count.

vs. R&D tax & compliance tooling

Documents experimentation after the fact for reporting.

Runs the experiment and documents it, with an optional R&D module that captures the work as you do it.

Compare approaches and their boundaries →

AI assistance

AI help that uses your study context.

The AI wizards can use your hypotheses, design, prompts and results to help draft a design, write prompts, suggest coding rules or discuss an analysis. Check their work before using it. Wizard availability depends on your plan; subject and judge calls on OpenRouter use your key.

  • Hypotheses

    Brainstorm and sharpen testable predictions.

  • Design

    Propose factorial conditions and levels.

  • Prompts

    Generate and validate snippet-bound prompt templates.

  • Coding

    Draft coding schemes and judge prompts.

  • Analysis

    Talk through results with the AI Analyst.

  • Write-up

    Draft prose with citations and references.

Bring your own OpenRouter key
For running trials against subject models.
Wizard options
Platform-provided AI wizard runs follow your plan's current allowance. Choosing OpenRouter for a wizard call uses your key and provider billing.
Subject models are BYOK
Studying LLM responses requires your own OpenRouter API key. Those OpenRouter usage costs are not included in AIScholar's pricing.
Estimated before you run
AIScholar estimates planned subject-model cost before execution. Actual charges vary, and response-coding judge calls can add separately billed OpenRouter usage.
Frequently asked questions

What to know before you start.

What exactly can I study on AIScholar?
Study text-derivable outcomes from single-turn OpenRouter model responses. Vary discrete prompt, model or supported parameter factors; replicate trials; and code observable response characteristics. You still need a valid question and measurement plan.
What is out of scope?
Subjects are OpenRouter-served models only: no multi-turn dialogues, tools or agents, multimodal inputs, fine-tuning or model internals. Subject-call reasoning effort is not reliably applied. Model-only trials do not recruit people, but human data or review may require institutional assessment.
Do I need my own API key?
Subject-model trials and response-coding judge calls use your OpenRouter key and may incur separate charges. Run estimates are not fixed prices; model rates, output length and retries affect final charges. Platform-provided AI wizard allowance depends on your plan; choosing OpenRouter for wizard calls uses your key.
Which models can be research subjects?
Available OpenRouter-served text models can be included as subject conditions. Check availability, exact model string and pricing when you run; provider or version changes may affect comparisons.
What statistics are built in?
Linear mixed models, factorial and repeated-measures ANOVA, logistic GLMM, chi-square, Bayesian tests with Bayes factors, non-parametric tests, equivalence testing (TOST), multiple-comparison corrections, inter-rater reliability, and effect sizes with bootstrap confidence intervals. You can also examine variance decomposition and prompt sensitivity.
What do I get on the Free plan?
The current platform-provided AI wizard allowance is temporarily unavailable. The Free plan includes access to the manual study workflow. Check the current plan details for project and feature limits; OpenRouter subject, response-coding judge and optional OpenRouter wizard usage is billed separately.
Does a model-only study require ethics review?
Model-inference trials alone do not recruit people. Human annotation or human-derived data can involve separate consent, privacy, or institutional requirements, which vary by context and are outside AIScholar's scope.
How should I interpret automated coding?
Automated codes are measurements, not ground truth. Define observable outcomes, inspect disagreements and validate codes for your research question before drawing conclusions.

Read the full FAQ for details on models, costs, data ownership and reproducibility.

Pricing

Start free. Check the plans when you need more.

The current platform-provided AI wizard allowance is temporarily unavailable. Check current plan terms for project limits and AI feature availability. Optional OpenRouter wizard calls use your own key and are billed separately.

Running trials against subject models uses your own OpenRouter API key. Those OpenRouter usage costs are billed by OpenRouter and are not included in any AIScholar plan. Automated coding/judge calls also use your OpenRouter key and can incur separate charges. Run estimates are not final charges; rates, tokens and retries vary. Review API-key costs.

Free

Run the full research workflow and try AI on the house.

$0

Free forever

  • Up to 3 projects
  • Platform-provided AI wizard allowance (current count unavailable)
  • Full pipeline through Visualization
  • Experiment execution & response collection
  • Built-in statistical analyses & charts
  • Data export (CSV / XLSX)
  • AI wizards capped — upgrade for unlimited
  • No AI Analyst chat
  • No Write-Up assistance
Most popular

Pro

Unlimited AI across every stage — go from idea to manuscript in a day.

$39.99/mo

Checking billing availability…

  • Up to 50 projects
  • Everything in Free
  • Unlimited AI hypothesis, design, prompt & coding-scheme wizards
  • AI Analyst — ask your data questions in plain English
  • AI Write-Up — citation-backed draft from Intro to Discussion

Enterprise

Everything in Pro, plus compliance-grade R&D documentation for grants & tax credits.

$99.99/mo

Checking billing availability…

  • Up to 500 projects
  • Everything in Pro
  • Reproducibility package export
  • R&D documentation module (OECD Frascati-aligned)
  • Technological uncertainty tracking
  • Effort & expenditure timesheets
  • R&D tax-incentive export package
  • Documentation supports SR&ED tax credits in Canada

Have a discount code? Enter it at checkout to apply your savings.

Prices in USD. Cancel anytime from the billing portal. Need a custom plan? Contact us.

Start with a question about model behavior.

Design your first single-turn text study; review subject-model costs before running.