Collection notes
Sources and scoring
Question counts come from the files bundled with this site. Original question wording and answer keys are retained. This is a practice interface, not an official benchmark evaluation.
MMLU-Pro
12,032 questions. Test split.
The bundled test split includes the upstream January 2026 option-format corrections.
Project and evaluation details · Dataset and license · Practice
GSM8K
8,792 questions. Train and test splits.
The test split is selected by default. Enter a final number without units.
Project and evaluation details · Dataset and license · Practice
Humanity's Last Exam
2,500 questions. Bundled 2,500-question snapshot.
All original answer choices are included. Written answers need your own review, including equivalent mathematical forms.
Project and evaluation details · Dataset and license · Practice
MMLU
14,042 questions. Test split.
Each question has one keyed answer. Subjects vary in difficulty.
Project and evaluation details · Dataset and license · Practice
ARC Challenge
1,172 questions. ARC-Challenge test split.
This is the AI2 science benchmark, separate from the ARC-AGI grid puzzles.
Project and evaluation details · Dataset and license · Practice
HellaSwag
1,000 questions. First 1,000 validation questions.
This collection is a subset of the validation split. Some source passages contain awkward phrasing.
Project and evaluation details · Dataset and license · Practice
WinoGrande
1,267 questions. Validation split.
Two choices per question. Read the full sentence before choosing.
Project and evaluation details · Dataset and license · Practice
BoolQ
3,270 questions. Validation split.
Choose yes or no based on the passage.
Project and evaluation details · Dataset and license · Practice
TruthfulQA
817 questions. Generation questions.
Several phrasings can be valid. Review your answer against the reference; this page does not judge its meaning automatically.
Project and evaluation details · Dataset and license · Practice
HumanEval
164 questions. Test split.
Code stays in this browser and is not executed. Use the reference and tests to assess your solution.
Project and evaluation details · Dataset and license · Practice
Local progress
Answers, scratch notes, and session results are stored in this browser. Clearing site data removes them. Private browsing or blocked storage may prevent saving. Each benchmark has its own saved session.
Rendering and attribution
Markdown is rendered with Marked and sanitized with DOMPurify. MathJax renders mathematical notation. Original dataset wording, including punctuation, is preserved. Dataset terms are linked above. Rendering library licenses.