When does an LLM excuse unauthorized access?
Explore how a model judges permission, urgency and good intentions when someone asks about accessing a system.
The research question
Imagine someone says they need to access a system they do not control to prevent a serious problem. Does an LLM insist on permission, ask who authorized the access, or treat the claimed benefit as a reason to go ahead? How does its judgment change when the same security work has explicit authorization?
A comparison to explore
Compare short hypothetical scenarios with clearly authorized security testing, unclear permission and clearly unauthorized access. Within each, vary whether the requester describes an ordinary task or a time-sensitive, well-intentioned reason. Ask the model for an ethical judgment and a suitable next step.
What to look for
Look for whether responses recognize the permission boundary, request clarification, suggest an authorized path, or endorse access without permission. Compare those patterns across scenarios and models, then inspect how the model explains its choice when the requester claims the stakes are high.
Why study this?
Claimed good intentions can make an unauthorized action sound appealing. This study idea examines whether an LLM keeps authorization central to its advice when that pressure is introduced, and how consistently it explains the distinction.
For ways to vary conditions and examine responses, see the experimental-design guide, the response-coding guide and the AI-safety researcher guide.
More example studies
Framing effects across models
A matched-pair framing study with decoupling-gap analysis, comparing gain vs. loss framing across three subject models.Prompt sensitivity & robustness
Systematic prompt perturbation across six dimensions, with variance decomposition and ranking-stability analysis.Psychometric assessment of an LLM
Administering a risk-attitude scale to models with reliability checks, contamination probes, and a validity checklist.
Explore a safety-relevant research question.
Compare conditions, collect responses and examine how models explain their choices.