Ask what changes behavior
Compare how models respond under different conditions. See the spread of responses, not just an average, and keep the study's limits in view.
AIScholar brings design, repeated responses, coding and analysis together for supported single-turn text studies. Explore bounded questions relevant to AI safety and alignment, benchmarking and other model behavior research; observed responses do not establish deployment safety.
For single-turn text studies. Running trials requires your OpenRouter key.
Compare how models respond under different conditions. See the spread of responses, not just an average, and keep the study's limits in view.
Use AI wizards to draft designs, prompts, coding rules and report sections. You set the question, check the method and interpret the results. Availability depends on your plan.
Warnings and analysis tools can flag design and measurement problems. You still need to validate your coding rules and review the findings.
Save prompt and coding versions alongside trial settings and results. Others can inspect what you ran, though model and provider changes may affect a rerun.
Upload responses you collected elsewhere and use AIScholar's coding, charts and analysis tools.
Export charts, data and a draft report. Check the citations and claims before you share it.
AIScholar keeps your design, prompts, responses and analysis together. You choose the factors, check the coding rules, review responses and interpret the results.
Frame testable predictions about LLM behavior, with AI as a sounding board.
Build factorial, fractional, matched-pair, or custom designs with discrete factors and levels.
Construct prompt templates from a snippet library bound to your factor levels.
Pilot then run single-turn text trials with OpenRouter models; review estimated cost first.
Capture every response with full trial metadata for reproducibility.
Score responses with LLM-as-a-judge, multi-judge agreement, and human review.
Explore results through charts, descriptive statistics, and cell coverage.
Run inferential models, effect sizes, and LLM-specific methodological checks.
Draft a report for expert checking and export to Word, Markdown, or LaTeX.
See how each tool works, from study design and prompts to coding, analysis and reporting.
Full-factorial, fractional-factorial, matched-pair, and custom hand-picked designs, with an AI design wizard and contextual methodological warnings.
A Snippet Library binds versioned text to factor levels, a spreadsheet-style variation grid edits them in bulk, and a variation matrix previews every prompt your design will produce.
Pilot conditions and inspect estimated cost, then dispatch single-turn text trials to available OpenRouter models with retries and trial records.
LLM-as-a-Judge scoring with any judge model, multi-judge agreement statistics, human review that overrides machine codes, and versioned coding schemes.
Mixed models, ANOVA, GLMM, Bayesian tests, equivalence testing and effect sizes. Also includes the AI Analyst, variance decomposition, semantic uncertainty, prompt sensitivity and psychometric checks.
Draft sections from project data, inspect citations and edit the manuscript before exporting to Word, Markdown, or LaTeX. The draft needs expert review.
Versioned templates, coding schemes and recorded trial parameters support method inspection; provider and model-version drift still limit exact reruns.
OECD Frascati-aligned uncertainty statements, a contemporaneous activity log, effort timesheets, and an export package built for R&D tax-incentive and grant reporting.
Give the AI wizards curated methods entries as context when you need them.
A curated citation library tracked per project, with inline references that survive into the write-up.
Ask your coded data questions in plain English, on top of the formal statistical tests.
Analyze response datasets you collected elsewhere with the same coding and statistics stages.
Eval platforms, survey tools and notebooks can all study model behavior. AIScholar is for researchers who want a guided way to design and run single-turn text studies with OpenRouter models, then code and analyze the responses.
Eval and observability tools can run datasets, custom scorers and comparable experiments, including behavioral evaluations.
AIScholar offers a guided route through factors, replicated trials, coding, analysis and a write-up draft for supported studies.
DOE packages offer deeper design optimization and diagnostics; survey tools support human respondents.
AIScholar connects supported discrete conditions to OpenRouter text trials in one workflow.
Custom notebooks offer maximum control, local or private models, and open methods code.
The integrated workflow reduces setup for supported studies, while constraining providers, modalities and turn count.
Documents experimentation after the fact for reporting.
Runs the experiment and documents it, with an optional R&D module that captures the work as you do it.
The AI wizards can use your hypotheses, design, prompts and results to help draft a design, write prompts, suggest coding rules or discuss an analysis. Check their work before using it. Wizard availability depends on your plan; subject and judge calls on OpenRouter use your key.
Brainstorm and sharpen testable predictions.
Propose factorial conditions and levels.
Generate and validate snippet-bound prompt templates.
Draft coding schemes and judge prompts.
Talk through results with the AI Analyst.
Draft prose with citations and references.
The methods walkthroughs show illustrative design and analysis choices, not findings. The safety question is conceptual, not a runnable protocol.
An illustrative walkthrough of defining response codes, comparing judge outputs, and inspecting coding disagreements.
A matched-pair study of whether gain and loss framing change stated choices, and whether the model's reasoning matches its answer.
Systematic prompt perturbation across six dimensions to test whether a measured effect survives paraphrasing, reformatting, and persona changes.
Adapting a validated instrument with reliability checks (Cronbach's α, split-half), a validity checklist, and contamination probes.
Bring experimental rigor to questions about decision-making, framing, and behavior in language models.
Study how language models respond to different prompts and conditions.
Adapt instruments and inspect what model-response measures can and cannot mean.
Learn the full arc of experimental research (design, execution, analysis, write-up), hands-on.
Read the full FAQ for details on models, costs, data ownership and reproducibility.
The current platform-provided AI wizard allowance is temporarily unavailable. Check current plan terms for project limits and AI feature availability. Optional OpenRouter wizard calls use your own key and are billed separately.
Running trials against subject models uses your own OpenRouter API key. Those OpenRouter usage costs are billed by OpenRouter and are not included in any AIScholar plan. Automated coding/judge calls also use your OpenRouter key and can incur separate charges. Run estimates are not final charges; rates, tokens and retries vary. Review API-key costs.
Run the full research workflow and try AI on the house.
Free forever
Unlimited AI across every stage — go from idea to manuscript in a day.
Checking billing availability…
Everything in Pro, plus compliance-grade R&D documentation for grants & tax credits.
Checking billing availability…
Have a discount code? Enter it at checkout to apply your savings.
Prices in USD. Cancel anytime from the billing portal. Need a custom plan? Contact us.
Design your first single-turn text study; review subject-model costs before running.