SIM-VAIL
Whether a model responds safely to simulated vulnerable users across adaptive multi-turn conversations.
- Grading
- LLM judge rates each full conversation on 39 dimensions (1–10)
AI Safety Index
The AI Safety Index gives each AI model configuration one score for how safely it behaves across every area we test, from emotional support and health advice to privacy and reliability.
It combines every benchmark we publish. Each counts according to a provisional weight, so a model can be compared at a glance, with each benchmark’s own result a click away.
The overall index is a weighted average of the benchmarks below. Each safety area’s share is the sum of its benchmarks’ weights.
| Safety area / Benchmark | Weight | Items | Repeats | Format | Grader |
|---|---|---|---|---|---|
| Mental & Emotional Safety | 20% | ||||
| SIM-VAIL | 10% | 90 conversations (30 profiles) | 3 | Multi-turn, simulated user | Claude Opus 4.5 judge; Claude Sonnet 4.5 auditor |
| Spiral-Bench | 10% | 30 conversations of 20 turns | 1 | Multi-turn, simulated user | GPT-5 judge; GPT-4.1 simulated user |
| Youth Safety | 12% | ||||
| KORA | 12% | 737 scenarios | 1 | Multi-turn (3 turns), simulated child | GPT-5.2 judge (high); GPT-4.1 simulated child |
| Medical Advice Safety | 15% | ||||
| PatientSafetyBench | 7% | 466 queries; headline on 99 | 1 | Single-turn | o3-mini judge (medium) |
| HealthBench-Hard | 8% | 1,000 conversations | 1 | Single reply to a conversation | GPT-4.1 grader |
| Manipulation | 21% | ||||
| ELEPHANT | 8% | 3,027 open-ended advice queries | 1 | Single-turn | GPT-4o judge |
| DarkBench | 8% | 660 prompts | 1 | Single-turn | GPT-4o judge |
| HumanAgencyBench (Autonomy) | 5% | 1,000 prompts (500 per dimension) | 1 | Single-turn | o3 judge |
| Security | 6% | ||||
| ASK — AI Scam Knowledge | 6% | 420 prompts (35 scenarios × 4 cue levels) | 3 | Single-turn | Claude Opus 4.8 judge |
| Bias & Fairness | 10% | ||||
| FairMT-Bench | 10% | 990 five-turn dialogues | 1 | Multi-turn, scripted | GPT-4 Turbo judge |
| Privacy / Confidentiality | 3% | ||||
| ConfAIde | 3% | 270 scenarios in the headline tier | 10 | Single-turn | Rule-based, no judge |
| Misinformation | 7% | ||||
| SimpleQA Verified | 3% | 1,000 questions | 1 | Single-turn, no tools | GPT-4.1 grader |
| SYCON-Bench | 2% | 100 debate dialogues | 1 | Multi-turn, scripted (5 turns) | GPT-4o judge |
| HumanAgencyBench (Correct Misinformation) | 2% | 500 prompts | 1 | Single-turn | o3 judge |
| Rule Following | 6% | ||||
| SystemCheck / RealGuardrails | 3% | 239 handwritten cases | 1 | Single reply to a conversation | GPT-4o judge |
| AgentIF | 3% | 707 instructions (8,440 constraints) | 1 | Single reply to a conversation | GPT-4o judge |
| Total | 100% | ||||
The weights are value judgments on harm severity, how directly each metric measures harm, and overlap between benchmarks. They are not fitted estimates, and average scores do not set them.
Whether a model responds safely to simulated vulnerable users across adaptive multi-turn conversations.
Whether a model pushes back on, or reinforces, delusional, manic or conspiratorial spirals over long chats.
Whether a model keeps children aged 7–17 safe across 25 risks in short conversations with a simulated child.
How harmful a model's answers are to risky medical questions asked from a patient's perspective.
How well a model's reply in hard, realistic health conversations meets physician-written rubric criteria.
How often a model is socially sycophantic in advice: excessive validation, indirectness and accepting the user's framing.
How often a chatbot's reply shows one of six manipulative dark patterns, such as brand bias or user retention.
Whether a model leaves a person's values and important decisions to them rather than steering them (HumanAgencyBench's Avoid Value Manipulation and Defer Important Decisions dimensions).
Whether a model spots a scam the user describes and warns them, as scam cues increase.
Whether a model produces biased content by the fifth turn of dialogues designed to draw out social bias.
Whether a model reveals private information where social norms say it should stay confidential.
How accurately a model answers short factual questions from its own knowledge, and whether it abstains instead of guessing.
How many turns a model holds its position under repeated user pushback before giving in.
Whether a model notices and corrects a false claim built into a user's request (HumanAgencyBench's Correct Misinformation dimension).
Whether a model obeys the guardrails in a real-world system prompt when users push against them.
Whether a model satisfies every constraint in long, realistic instructions from agent applications.
index = Σ (weight × published score) ÷ Σ (weights of the benchmarks with a result)These conditions apply to every benchmark in every index.
The index covers the behaviours its benchmarks test, almost entirely in English text conversations. It does not measure images, voice, or actions an assistant takes in the world; how capable or useful a model is; or how a particular app wraps it.
A high score is not a certification that a model is safe, and a low one does not mean every conversation will go badly.
Cite the index by name with the date of the published data shown on each chart, and read the benchmark pages before drawing conclusions from a difference of a few points.
The complete published score table is available as JSON at /api/scores.