HumanAgencyBench (Correct Misinformation)

Misinformation. Headline metric: Correct Misinformation dimension score, 0–1 (higher is better).

Data published

Model comparison

Scores use the published 0–100 transformations; higher is better on the selected metric. Indexes average these published scores; none is normalized. Raw scores below retain their published scale. A missing result is not a zero.

Showing 14 of 14 model configurations.

Each bar is a toggle button. Activate a bar to pin or unpin that model. The data table contains exact scores and sources.
  1. AnthropicClaude Opus 5.5
  2. OpenAIGPT-6 Astra
  3. PerplexityPerplexity Agent
  4. OpenAIGPT-6 Sol
  5. OpenAIGPT-5.6 Terra
  6. Zhipu AIGLM 5.3 Flash
  7. AnthropicClaude Sonnet 5
  8. xAIGrok 4.6
  9. xAIGrok 4.5
  10. OpenAIGPT-5.6 Luna
  11. DeepSeekDeepSeek V4 Flash
  12. InklingInkling
  13. GoogleGemini 3.6 Flash
  14. Mistral AIMistral Medium 3.5
HumanAgencyBench (Correct Misinformation): every model’s score, interval and source
ModelProviderModel versionReasoning settingScore out of 100IntervalSourceMeasuredSample
Claude Opus 5.5 · MediumAnthropicClaude Opus 5.5Medium95.80Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
GPT-6 Astra · MediumOpenAIGPT-6 AstraMedium89.66Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
Perplexity Agent · medium presetPerplexityPerplexity Agent · medium presetNot recorded88.66Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
GPT-6 Sol · MediumOpenAIGPT-6 SolMedium85.28Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
GPT-5.6 Terra · MediumOpenAIGPT-5.6 TerraMedium82.42Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
GLM 5.3 Flash · HighZhipu AIGLM 5.3 FlashHigh82.08Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
Claude Sonnet 5 · MediumAnthropicClaude Sonnet 5Medium78.52Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
Grok 4.6 · MediumxAIGrok 4.6Medium72.36Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
Grok 4.5 · MediumxAIGrok 4.5Medium67.90Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
GPT-5.6 Luna · MediumOpenAIGPT-5.6 LunaMedium62.50Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
DeepSeek V4 Flash · MediumDeepSeekDeepSeek V4 FlashMedium58.42Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
Inkling · MediumInklingInklingMedium57.74Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
Gemini 3.6 Flash · MediumGoogleGemini 3.6 FlashMedium43.44Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied
Mistral Medium 3.5Mistral AIMistral Medium 3.5Not recorded15.16Not suppliedHumanAgencyBench (Correct Misinformation)Not suppliedNot supplied

How it’s measured

Whether a model notices and corrects a false claim built into a user's request (HumanAgencyBench's Correct Misinformation dimension).

Items
500 prompts
Repeats
1
Format
Single-turn
Grading
LLM judge lists rubric issues; each prompt scores 10 minus deductions
Grader
o3 judge
Headline metric
Correct Misinformation dimension score, 0–1 (higher is better)
On the 0–100 scale
Score × 100

Scores and coverage

All configurations, including missing results
Model configurationPublished score / 100Native score (0–1)
Claude Opus 5.5 · Medium95.800.95
GPT-6 Astra · Medium89.660.89
Perplexity Agent · medium preset88.660.88
GPT-6 Sol · Medium85.280.85
GPT-5.6 Terra · Medium82.420.82
GLM 5.3 Flash · High82.080.82
Claude Sonnet 5 · Medium78.520.78
Grok 4.6 · Medium72.360.72
Grok 4.5 · Medium67.900.67
GPT-5.6 Luna · Medium62.500.62
DeepSeek V4 Flash · Medium58.420.58
Inkling · Medium57.740.57
Gemini 3.6 Flash · Medium43.440.43
Mistral Medium 3.515.160.15

About these results

These are evaluations of API model configurations. They do not establish how a consumer app behaves with its own prompts, tools, or safeguards.

The date identifies the published dataset. Measurement dates, sample sizes, and confidence intervals are not included in this score-only feed.

Published September 25, 2026. Read the benchmark methodology.

HumanAgencyBench (Correct Misinformation) Leaderboard — AI Safety Benchmark — Prosaic Intelligence