Karol Leszczyński Build & Develop PL

Matury Online — a platform that grades what nobody graded.

My own product, built and run from the first line of code. It prepares students for the matura — the Polish secondary-school final exam. Below is what the portfolio card leaves out: where the idea came from, which decision mattered most, and what was genuinely hard along the way.

17,300+
exercises in the question bank
11
exam subjects covered
19
task types from official exam papers
~17 s
average essay grading time

Problem

Everyone automated the part that was easy.

Several hundred thousand people prepare for the matura every year, and plenty of apps serve them. Almost all did the same thing: multiple-choice tests, a correct-answer counter, a leaderboard. In other words, exactly the part of the exam that carries the least weight.

  1. 01

    The points sit in the open-ended tasks

    On the Polish-language matura exam, 35 of 60 points ride on a single essay. In mathematics and the sciences it is the line of reasoning that earns points, not just the final number.

  2. 02

    The market automated only the easy part

    Multiple-choice tasks can be compared against an answer key, so every app has them. For open-ended tasks, platforms showed a model answer and left the grading to the student.

  3. 03

    Students were asking a different question

    They were asking how many points their own answer would earn. A model answer says nothing about that.

Solution

The difference sits in what the model is handed.

Everyone uses a language model these days, and by itself that means nothing. The real work was preparing the material the model works with.

Grading by the examiner's criteria

The model receives the student's answer together with the grading rubric for that task type — the same one an examiner scores against. That turns 'grade this essay' into a task with a clearly defined result.

A score with a breakdown and reasoning

Back comes the number of points in each category, the places where points were lost, and the reason. The student sees exactly what to fix.

A result in seconds

Grading an essay takes about seventeen seconds on average. A teacher with thirty papers a week cannot compete with that, and does not have to — this is a training tool for the time between lessons.

The same for listening tasks

Recordings are generated on the spot, fresh with every attempt, and open answers to them come back scored. Nobody else in Poland does this.

I described the underlying problem in more detail on the product's site (in Polish): open-ended tasks on the matura.

From the app

What it looks like from the inside.

Three screens from the live app, in the order things happen: the student writes their own answer, grading runs in stages, and a score report comes back.

Matury Online — An open-ended task during the exam
An open-ended task during the exam An empty field and the prompt 'explain the meaning of the phrase'. Exactly the task type the competition does not check. Next to it, a timer and a count of tasks remaining.
Matury Online — Grading, step by step
Grading, step by step Multiple-choice tasks are scored instantly, since a key is enough. Open-ended tasks, the note and the essay are separate stages — each with its own grading rubric.
Matury Online — The score report
The score report The result against the pass threshold, how many points are missing, and feedback that addresses the student's specific gaps rather than a generic 'study harder'.
Matury Online — language selection in listening mode

Recordings are created the moment you ask for them

The model composes a fresh transcript on every attempt, and the synthesizer reads it aloud in the right accent. The material never runs out, and each recording can be played twice — just like on the real exam. Open answers to these recordings come back scored as well.

What was hard

What turned out to be hardest.

  • Grading had to be repeatable A model told to 'grade this essay' answers differently every time. Only handing it the rubric and forcing the response into a fixed structure turns it into a tool you can trust.
  • Cost calculated per answer Every grading run is a real model-call expense. At a 49 zł monthly subscription, I had to calculate how many gradings fit into the price and build limits before the bill could spring a surprise.
  • Recordings generated on the fly The app composes a fresh transcript and synthesizes the voice on every attempt. The material never runs out, but the whole thing must be ready before the student loses patience.
  • One backend, two entry points The mobile app on Google Play uses the same backend as the browser version. One grading logic, one account, no versions drifting apart.

What this means for you

The same pattern works outside education.

A model that grades work against fixed criteria, billed per graded item and wired into a product with payments — this pattern fits anywhere someone currently evaluates things by hand against a set checklist: applications, submissions, offers, documentation.

If you have a process like that, tell me about it — I will calculate what a single evaluation would cost and whether it makes sense at all.

Related services

What I do beyond my own products.

This project is a web application with an AI deployment. Below is the rest of what I build for companies, with starting prices.

Let's talk

Have a process someone grades by hand?

Describe how it works today and what it is judged against. I will tell you whether it can become a tool, roughly what a single evaluation would cost, and where to start — usually the same day.