Grading by the examiner's criteria
The model receives the student's answer together with the grading rubric for that task type — the same one an examiner scores against. That turns 'grade this essay' into a task with a clearly defined result.
My own product, built and run from the first line of code. It prepares students for the matura — the Polish secondary-school final exam. Below is what the portfolio card leaves out: where the idea came from, which decision mattered most, and what was genuinely hard along the way.
Problem
Several hundred thousand people prepare for the matura every year, and plenty of apps serve them. Almost all did the same thing: multiple-choice tests, a correct-answer counter, a leaderboard. In other words, exactly the part of the exam that carries the least weight.
On the Polish-language matura exam, 35 of 60 points ride on a single essay. In mathematics and the sciences it is the line of reasoning that earns points, not just the final number.
Multiple-choice tasks can be compared against an answer key, so every app has them. For open-ended tasks, platforms showed a model answer and left the grading to the student.
They were asking how many points their own answer would earn. A model answer says nothing about that.
Solution
Everyone uses a language model these days, and by itself that means nothing. The real work was preparing the material the model works with.
The model receives the student's answer together with the grading rubric for that task type — the same one an examiner scores against. That turns 'grade this essay' into a task with a clearly defined result.
Back comes the number of points in each category, the places where points were lost, and the reason. The student sees exactly what to fix.
Grading an essay takes about seventeen seconds on average. A teacher with thirty papers a week cannot compete with that, and does not have to — this is a training tool for the time between lessons.
Recordings are generated on the spot, fresh with every attempt, and open answers to them come back scored. Nobody else in Poland does this.
I described the underlying problem in more detail on the product's site (in Polish): open-ended tasks on the matura.
From the app
Three screens from the live app, in the order things happen: the student writes their own answer, grading runs in stages, and a score report comes back.
The model composes a fresh transcript on every attempt, and the synthesizer reads it aloud in the right accent. The material never runs out, and each recording can be played twice — just like on the real exam. Open answers to these recordings come back scored as well.
What was hard
What this means for you
A model that grades work against fixed criteria, billed per graded item and wired into a product with payments — this pattern fits anywhere someone currently evaluates things by hand against a set checklist: applications, submissions, offers, documentation.
If you have a process like that, tell me about it — I will calculate what a single evaluation would cost and whether it makes sense at all.
Related services
This project is a web application with an AI deployment. Below is the rest of what I build for companies, with starting prices.
Let's talk
Describe how it works today and what it is judged against. I will tell you whether it can become a tool, roughly what a single evaluation would cost, and where to start — usually the same day.