Linear algebra homework page
PDF page render / handwritten matrix solution / image input to model
Can an LLM replace a teaching-assistant grader? Grade Arena is finding out. We give models real student homework PDFs, then compare every score they assign—0, 0.5, or 1—with the score a PhD math grader gave the same work. The current leaderboard includes Gemini, DeepSeek, GLM, Qwen, MiniMax, and GPT models.
Each handwritten solution must be read and checked step by step.
Grading is one part of this market. Grade Arena tests whether AI can grade handwritten university math accurately and at low cost.
We compare each model's task scores with the PhD grader's score for the same student work. Gemini 3.1 Pro, DeepSeek V4 Pro, GLM‑5.2, Qwen 3.7 Max, Gemini 3.5 Flash, and MiniMax M3 are normalized to the same 1,007 labels across 116 submissions. Missing reports count as omitted output; GPT‑5.6-sol currently has 913 labels across 104 available reports.
| Model | Same score | 0.5 away | 1 away | Same or close | Pricing / billing | Cost / task |
|---|---|---|---|---|---|---|
Gemini 3.1 Pro 1,007 tasks · LLM matched to PhD human grader |
84.6%exactly like the grader | 11.3%half a point different | 4.1%one point different | 95.9%same or 0.5 away | $2 in / $12 outstandard tier, ≤200K-token prompt | ≈ $0.045benchmark token footprint |
Gemini 3.5 Flash 1,007 tasks · LLM matched to PhD human grader |
66.1%exactly like the grader | 12.0%half a point different | 21.8%one point different | 78.2%same or 0.5 away | OpenRouter billingobserved during benchmark run | $0.010 observed$10.46 run total ÷ 1,007 tasks |
GPT-5.6-sol 913 tasks · LLM matched to PhD human grader |
56.0%exactly like the grader | 17.3%half a point different | 26.7%one point different | 73.3%same or 0.5 away | $5 in / $30 outstandard API tier | $0.050 observed$46.10 run total ÷ 913 tasks |
DeepSeek V4 Pro 1,007 tasks · LLM matched to PhD human grader |
69.4%exactly like the grader | 12.8%half a point different | 17.8%one point different | 82.2%same or 0.5 away | OpenRouter billing2.58M tokens · full run | $0.008 observed$8.28 run total ÷ 1,007 tasks |
GLM-5.2 1,007 tasks · LLM matched to PhD human grader |
69.6%exactly like the grader | 11.1%half a point different | 19.3%one point different | 80.7%same or 0.5 away | OpenRouter billing2.62M tokens · full run | $0.009 observed$8.64 run total ÷ 1,007 tasks |
MiniMax M3 1,007 tasks · LLM matched to PhD human grader |
59.0%exactly like the grader | 17.3%half a point different | 23.7%one point different | 76.3%same or 0.5 away | OpenRouter billing5.64M tokens · normalized reports | $0.006 observed$5.86 normalized cost ÷ 1,007 tasks |
Qwen 3.7 Max 1,007 tasks · LLM matched to PhD human grader |
67.2%exactly like the grader | 11.6%half a point different | 21.2%one point different | 78.8%same or 0.5 away | OpenRouter billing2.52M tokens · full run | $0.009 observed$9.10 run total ÷ 1,007 tasks |
The percentage shows how often each model gave the same score as the PhD grader or was only 0.5 points away. Scroll sideways on smaller screens. Discrete mathematics has only 3 comparable submissions, so treat that result cautiously.
The model receives rendered pages from student PDFs. These pages contain handwritten solutions, equations, corrections, and partial work.
PDF page render / handwritten matrix solution / image input to model
PDF page render / handwritten limit solution / image input to model
The model gets rendered pages from the student's PDF and a grading instruction. It does not get the official solution. It does not get the original problem statement unless it is visible in the student's pages.
There is no detailed rubric for each task. The model assigns a score using the same coarse scale as the dataset.
You are grading a math homework submission. Input: - rendered pages from the student's PDF - no official solution - no separate problem statement - no task-specific rubric For each visible task, assign: 1 = solution is correct 0.5 = solution is partially correct 0 = solution is wrong or missing Return task scores and short comments.
The dataset is mostly introductory university mathematics: early undergraduate courses that mix calculations with short proofs and require students to show their reasoning. Representative problems include proving that a sequence converges, finding the rank of a matrix or a basis for a vector space, and calculating a conditional probability or expected value. The table shows how many homework sets, student submissions, and individually graded tasks are available in each subject.
A benchmark for comparing model grades with PhD math grader grades on math homework PDFs.