AI Safety Index

Rule Following Index methodology

Index. Rule Following.

What this is

The Rule Following index gives each AI model configuration one score for this area of safety (how reliably models follow instructions and respect boundaries they have been given).

It averages the benchmarks listed below, each with an equal share, so no single test decides the result.

Index composition

The Rule Following index gives each benchmark below an equal share.

Index composition
BenchmarkWeightItemsRepeatsFormatGrader
SystemCheck / RealGuardrails50%239 handwritten cases1Single reply to a conversationGPT-4o judge
AgentIF50%707 instructions (8,440 constraints)1Single reply to a conversationGPT-4o judge
Total100%

What each benchmark measures

AgentIF

Whether a model satisfies every constraint in long, realistic instructions from agent applications.

Grading
Per-constraint checks by LLM judge and code

How the index is calculated

index = Σ published scores ÷ number of benchmarks with a result
  • Every input is a benchmark’s published 0–100 score, used as it is. No score is normalized, before or after averaging.
  • A missing result is left out, never counted as zero; the remaining weights are rescaled for that model.
  • The index summarizes evaluations. It is not a percentage of safety or a guarantee about any product.

Testing parameters

These conditions apply to every benchmark in every index.

Published scores
Each benchmark contributes one headline metric, converted to a 0–100 scale where higher is better. Values past either end are clipped.
Model configurations
The reasoning setting is part of each configuration. Configurations are published separately and never averaged together.
Sampling
Benchmarks set no temperature for the tested model; each provider's default applies.
Output limit
The lower of the model's own maximum and the benchmark's cap, which is 65,536 tokens where one is set.
Retries
Failed API calls are retried up to 3 times with backoff; a stopped run is resumed up to twice.
Provider refusals
An explicit provider safety block is recorded, not retried. DarkBench, FairMT-Bench, ELEPHANT and KORA score it as the safe outcome; elsewhere it is left unscored.
Judge failures
A refused or failed judge, simulator or auditor call is an execution failure and is never scored.
Which run is published
The newest complete full run on the pinned protocol. Partial runs never replace complete ones.
Scope
Scores describe API model configurations. Consumer apps may add their own prompts, tools and safeguards.

What this does not measure

The index covers the behaviours its benchmarks test, almost entirely in English text conversations. It does not measure images, voice, or actions an assistant takes in the world; how capable or useful a model is; or how a particular app wraps it.

A high score is not a certification that a model is safe, and a low one does not mean every conversation will go badly.

Using this data

Cite the index by name with the date of the published data shown on each chart, and read the benchmark pages before drawing conclusions from a difference of a few points.

The complete published score table is available as JSON at /api/scores.