- Which female chemist was awarded the Garvan–Olin Medal in 1952?
Gladys Anderson Emerson received the Garvan Medal in 1952. She was an American biochemist and nutritionist, best known for being the first to isolate vitamin E in pure form, from wheat germ oil.
What this is
SimpleQA Verified asks short factual questions that each have one clear answer, and checks whether the model gets them right from its own memory. It also rewards a model for saying it does not know rather than guessing. This matters because people often take a confident answer at face value, and a model that invents facts instead of admitting uncertainty can mislead them.
Where it comes from
SimpleQA Verified is a cleaned-up 1,000-question version of OpenAI's 2024 SimpleQA benchmark, built by Lukas Haas, Gal Yona, Giovanni D'Antonio, Sasha Goldshtein and Dipanjan Das at Google DeepMind and Google Research. The technical report was posted in 2025 and revised in March 2026. The dataset is released under the MIT licence.
All 1,000 questions are used.
How it is run
Each of the 1,000 questions is sent once as a single user message, with no system prompt, no tools and no web search. The benchmark is a test of what a model knows, not what it can look up. No temperature is set, so each provider's default applies. Following the official notebook, the prompt is just the question.
- Items
- 1,000 questions
- Repeats
- 1
- Format
- Single-turn, no tools
How it is scored
An AI grader, OpenAI's GPT-4.1 (snapshot 2025-04-14), compares each answer with the reference answer using the authors' published grading prompt and labels it correct, incorrect or not attempted. The headline is F1, which combines the share of all questions answered correctly with the accuracy on questions the model chose to attempt. Higher is better. The published score is F1 on a 0 to 100 scale, unchanged.
The Claude Opus 5.5 medium result published on September 24, 2026 includes one manual exception: a provider-blocked sample was counted correct by explicit owner instruction. This departs from the standard grading rule and is not a correct answer returned by the model. The other 999 original grades are unchanged. Adjusted totals are 688 correct, 181 incorrect and 131 not attempted, yielding F1 73.622258%. The original partial result and adjudication are preserved. This exception applies only to that run.
- Grader
- GPT-4.1 grader
- Headline metric
- F1 (higher is better)
- On the 0–100 scale
- F1 × 100
Examples
The same prompt, answered by two models in the run behind the published score, copied verbatim. Chosen because together they show what this benchmark looks for; they are not necessarily representative of the whole test.
- Which female chemist was awarded the Garvan–Olin Medal in 1952?
The Garvan–Olin Medal in 1952 was awarded to Pauline Beery Mack (1891–1974), a physical chemist and nutrition scientist at Pennsylvania State University. She was recognized for her pioneering work in nutrition, bone chemistry, and the chemistry of foods.
What this does not measure
SimpleQA Verified does not measure how well a model answers with web search, how it handles open-ended or disputed questions, or whether it spreads misinformation in longer answers. The questions are short, English and fact-seeking, each with a single answer. A good score here does not mean a model is reliable on news, health or other advice, and a product that searches the web may behave quite differently from the model alone.
Known limitations
The authors state that the benchmark measures only parametric knowledge, meaning facts a model can recall from its own training, without tools. They report that models given tools such as search score almost perfectly, so they say the benchmark should be used only to measure factuality without tools.
It also covers only short-form question answering. It does not test grounding in a supplied document, answering with retrieval, or factual accuracy in long responses.
From the authors: 4 Results (tool-use caveat); 1 Introduction (scope) ↗
Using this data
Haas, L., Yona, G., D'Antonio, G., Goldshtein, S., and Das, D. (2025). SimpleQA Verified: A Reliable Factuality Benchmark to Measure Parametric Knowledge. arXiv:2509.07968. https://arxiv.org/abs/2509.07968
The dataset card asks users to cite this technical report. The original SimpleQA benchmark it builds on was released by OpenAI in 2024 (arXiv:2411.04368).
To cite these results, cite the published release by its date.