Dark technical math benchmark dashboard with homework pages and equations
Grade Arena Homework Benchmark

How well can LLMs grade university math?

Can an LLM replace a teaching-assistant grader? Grade Arena is finding out. We give models real student homework PDFs, then compare every score they assign—0, 0.5, or 1—with the score a PhD math grader gave the same work. The current leaderboard includes Gemini, DeepSeek, GLM, Qwen, MiniMax, and GPT models.

00 / The Problem

Grading open-ended math takes expert time.

Each handwritten solution must be read and checked step by step.

$18.3B global higher-education testing & assessment market
estimated for 2025

Higher-ed assessment is an $18.3B market.

Grading is one part of this market. Grade Arena tests whether AI can grade handwritten university math accurately and at low cost.

01 / Current Run

How closely does each model grade like a human?

We compare each model's task scores with the PhD grader's score for the same student work. Gemini 3.1 Pro, DeepSeek V4 Pro, GLM‑5.2, Qwen 3.7 Max, Gemini 3.5 Flash, and MiniMax M3 are normalized to the same 1,007 labels across 116 submissions. Missing reports count as omitted output; GPT‑5.6-sol currently has 913 labels across 104 available reports.

Green: same score as the grader Yellow: 0.5 points away Red: 1 point away
Model Same score 0.5 away 1 away Same or close Pricing / billing Cost / task
Gemini 3.1 Pro
1,007 tasks · LLM matched to PhD human grader
84.6%exactly like the grader 11.3%half a point different 4.1%one point different 95.9%same or 0.5 away $2 in / $12 outstandard tier, ≤200K-token prompt ≈ $0.045benchmark token footprint
Gemini 3.5 Flash
1,007 tasks · LLM matched to PhD human grader
66.1%exactly like the grader 12.0%half a point different 21.8%one point different 78.2%same or 0.5 away OpenRouter billingobserved during benchmark run $0.010 observed$10.46 run total ÷ 1,007 tasks
GPT-5.6-sol
913 tasks · LLM matched to PhD human grader
56.0%exactly like the grader 17.3%half a point different 26.7%one point different 73.3%same or 0.5 away $5 in / $30 outstandard API tier $0.050 observed$46.10 run total ÷ 913 tasks
DeepSeek V4 Pro
1,007 tasks · LLM matched to PhD human grader
69.4%exactly like the grader 12.8%half a point different 17.8%one point different 82.2%same or 0.5 away OpenRouter billing2.58M tokens · full run $0.008 observed$8.28 run total ÷ 1,007 tasks
GLM-5.2
1,007 tasks · LLM matched to PhD human grader
69.6%exactly like the grader 11.1%half a point different 19.3%one point different 80.7%same or 0.5 away OpenRouter billing2.62M tokens · full run $0.009 observed$8.64 run total ÷ 1,007 tasks
MiniMax M3
1,007 tasks · LLM matched to PhD human grader
59.0%exactly like the grader 17.3%half a point different 23.7%one point different 76.3%same or 0.5 away OpenRouter billing5.64M tokens · normalized reports $0.006 observed$5.86 normalized cost ÷ 1,007 tasks
Qwen 3.7 Max
1,007 tasks · LLM matched to PhD human grader
67.2%exactly like the grader 11.6%half a point different 21.2%one point different 78.8%same or 0.5 away OpenRouter billing2.52M tokens · full run $0.009 observed$9.10 run total ÷ 1,007 tasks
Gemini 3.1 Pro was reconstructed from the stored LLM scores against the PhD human grader scores; DeepSeek, GLM‑5.2, Qwen 3.7 Max, Gemini 3.5 Flash, MiniMax M3, and GPT‑5.6-sol were reconstructed from saved LaTeX grading files. All use the same protocol: task labels are normalized and aligned in report order, and required tasks omitted from a model report count as zero-score schema failures. No model reruns were used. Gemini 3.1 Pro, DeepSeek, GLM, Qwen, Flash, and MiniMax are normalized to 1,007 labels. Qwen has 110 report artifacts with six missing reports counted as omitted output; MiniMax has 115 with one missing. GPT has 913 labels because only 104 successful report artifacts were available. Observed provider billing was $8.6420 for GLM, $9.0980 for Qwen, $5.8643 for the normalized MiniMax reports, $10.4609 for Flash, and $46.0955 for GPT. The Gemini 3.1 Pro cost is an estimate that applies its current list price to the benchmark average of approximately 9.7K input tokens and 2.1K output/reasoning tokens per task. Actual cost varies with PDF length, reasoning, caching, and provider. Prices checked July 11, 2026. Sources: OpenRouter model pricing, Google Gemini API pricing, and OpenAI GPT‑5.6 pricing.
02 / Subjects

Results by subject.

The percentage shows how often each model gave the same score as the PhD grader or was only 0.5 points away. Scroll sideways on smaller screens. Discrete mathematics has only 3 comparable submissions, so treat that result cautiously.

Subject
Works
Gemini 3.1 Pro
DeepSeek V4 Pro
GLM-5.2
Qwen 3.7 Max
Gemini 3.5 Flash
MiniMax M3
GPT-5.6-sol
Intro Analysis
proof-heavy foundations
14
96.3%
68.6%
65.3%
65.3%
62.8%
62.8%
64.5%
Linear Algebra
matrices, bases, maps
25
97.8%
90.3%
93.0%
85.4%
92.4%
75.1%
81.7%
Mathematical Analysis
calculus and proofs
40
94.5%
81.7%
80.2%
77.4%
79.0%
80.2%
79.6%
Elementary Mathematics
algebra and short proofs
20
96.5%
88.0%
85.9%
85.9%
84.5%
79.6%
89.8%
Probability Theory
distributions and expectation
14
96.6%
96.6%
89.7%
95.4%
78.2%
81.6%
47.6%
Discrete Mathematics
small sample
3
91.3%
78.3%
87.0%
82.6%
73.9%
69.6%
91.7%
03 / Model Input

Input example.

The model receives rendered pages from student PDFs. These pages contain handwritten solutions, equations, corrections, and partial work.

Rendered page from a linear algebra homework PDF with handwritten determinant work

Linear algebra homework page

PDF page render / handwritten matrix solution / image input to model

Rendered page from a mathematical analysis homework PDF with handwritten trigonometric limit work

Mathematical analysis homework page

PDF page render / handwritten limit solution / image input to model

Prompt excerpt short version

What the model gets

The model gets rendered pages from the student's PDF and a grading instruction. It does not get the official solution. It does not get the original problem statement unless it is visible in the student's pages.

There is no detailed rubric for each task. The model assigns a score using the same coarse scale as the dataset.

You are grading a math homework submission.

Input:
- rendered pages from the student's PDF
- no official solution
- no separate problem statement
- no task-specific rubric

For each visible task, assign:
1   = solution is correct
0.5 = solution is partially correct
0   = solution is wrong or missing

Return task scores and short comments.
04 / Problem View

Dataset coverage.

The dataset is mostly introductory university mathematics: early undergraduate courses that mix calculations with short proofs and require students to show their reasoning. Representative problems include proving that a sequence converges, finding the rank of a matrix or a basis for a vector space, and calculating a conditional probability or expected value. The table shows how many homework sets, student submissions, and individually graded tasks are available in each subject.

Topic
Homework sets
Submissions
Task labels
Range
Mathematical Analysis / Calculus
HW 1-9
41
331
limits, derivatives, sequences, series, integration, proof checks
Linear Algebra
HW 1-5
25
186
matrices, inverse matrices, rank, bases, linear maps, algebraic reasoning
Introduction to Analysis
HW 1-3
17
303
foundational proof-writing, inequalities, functions, sequences, epsilon-style arguments
Probability Theory
HW 1-5
17
103
events, conditioning, random variables, distributions, expectation
Elementary Mathematics
HW 1-4
21
146
algebraic manipulation, inequalities, functions, combinatorial-style exercises
Discrete Mathematics
HW 1-2
3
24
logic, counting, graph/discrete structures, short proof tasks
About

Grade Arena

A benchmark for comparing model grades with PhD math grader grades on math homework PDFs.

Talk to Grade Arena

Tell us what you are exploring. We will reply by email.

Your details are used only to respond to this request.