Benchmark lab / kaileh.dev

Collection notes

Sources and scoring

Question counts come from the files bundled with this site. Original question wording and answer keys are retained. This is a practice interface, not an official benchmark evaluation.

MMLU-Pro

12,032 questions. Test split.

The bundled test split includes the upstream January 2026 option-format corrections.

Project and evaluation details · Dataset and license · Practice

GSM8K

8,792 questions. Train and test splits.

The test split is selected by default. Enter a final number without units.

Project and evaluation details · Dataset and license · Practice

Humanity's Last Exam

2,500 questions. Bundled 2,500-question snapshot.

All original answer choices are included. Written answers need your own review, including equivalent mathematical forms.

Project and evaluation details · Dataset and license · Practice

MMLU

14,042 questions. Test split.

Each question has one keyed answer. Subjects vary in difficulty.

Project and evaluation details · Dataset and license · Practice

ARC Challenge

1,172 questions. ARC-Challenge test split.

This is the AI2 science benchmark, separate from the ARC-AGI grid puzzles.

Project and evaluation details · Dataset and license · Practice

HellaSwag

1,000 questions. First 1,000 validation questions.

This collection is a subset of the validation split. Some source passages contain awkward phrasing.

Project and evaluation details · Dataset and license · Practice

WinoGrande

1,267 questions. Validation split.

Two choices per question. Read the full sentence before choosing.

Project and evaluation details · Dataset and license · Practice

BoolQ

3,270 questions. Validation split.

Choose yes or no based on the passage.

Project and evaluation details · Dataset and license · Practice

TruthfulQA

817 questions. Generation questions.

Several phrasings can be valid. Review your answer against the reference; this page does not judge its meaning automatically.

Project and evaluation details · Dataset and license · Practice

HumanEval

164 questions. Test split.

Code stays in this browser and is not executed. Use the reference and tests to assess your solution.

Project and evaluation details · Dataset and license · Practice

Local progress

Answers, scratch notes, and session results are stored in this browser. Clearing site data removes them. Private browsing or blocked storage may prevent saving. Each benchmark has its own saved session.

Rendering and attribution

Markdown is rendered with Marked and sanitized with DOMPurify. MathJax renders mathematical notation. Original dataset wording, including punctuation, is preserved. Dataset terms are linked above. Rendering library licenses.