“AI-powered evaluation” is on the box of every assessment product now, which makes it hard to tell what is doing the work. For objective exams the honest answer is unglamorous and reassuring: the grading is computer vision, a deterministic image pipeline that gives the same answer every time you run it. No language model is involved, and that is a feature.

Here is what actually happens between pointing a camera at an answer sheet and a score appearing — and where AI does earn its keep.


The grading pipeline, step by step

01

Frame capture and quality gate

The app samples camera frames continuously and rejects the ones that are too blurred, too dark or too tilted to trust. This gate is why a scan sometimes takes a second longer: it is waiting for a frame worth measuring rather than grading a bad one.

02

Registration mark detection

It looks for the four solid corner squares printed on the sheet. Their positions in the image reveal where the paper is, how it is rotated, and how much the camera angle has distorted it into a trapezoid.

03

Perspective correction

A homography — a 3×3 transform derived from those four points — maps the skewed photograph back onto a flat rectangle. After this step the software is working with a sheet that behaves as if it had been laid on a flatbed scanner.

04

Grid projection

Because the template geometry is known, every bubble centre can be computed arithmetically from the corrected corners and the timing marks. This is precisely why a hand-drawn sheet cannot be read: there is no known geometry to project.

05

Coverage measurement

For each bubble the software measures what proportion of the enclosed area is dark, after adaptive thresholding that compensates for uneven lighting across the page. A filled bubble is dense; an empty one is nearly white; a tick mark sits awkwardly between.

06

Answer resolution and scoring

Per question, the darkest bubble wins if it clears the threshold and beats the runner-up by a margin. Two comparable marks mean an invalid answer, not a guess. The resolved answers are then compared to the key and scored, with negative marking if configured.

How one bubble is actually read

It helps to see the numbers. Suppose the coverage threshold is 45% and the margin requirement is 15 percentage points:

What the student didCoverage A / B / C / DResult
Filled C properly3% / 4% / 88% / 2%C, high confidence
Ticked C instead of filling3% / 4% / 22% / 2%Below threshold — flagged for review
Filled B, erased, filled D3% / 31% / 4% / 79%D, flagged: residue on B
Filled B and D3% / 81% / 4% / 84%Invalid — margin not met

Nothing here is guesswork, and nothing is learned from data. It is measurement plus two thresholds, which is exactly what you want in a system whose output a parent may dispute.

Reading the roll number

Roll number recognition is the same mechanism, read column-wise. Each digit position is a column of ten bubbles; the filled one gives that digit. This is far more robust than handwriting recognition on the printed boxes, which is why sheets ask students to do both and the reader uses the grid.

If the grid is left blank, there is nothing to read — the scan still grades correctly but the result needs a student attached by hand. Worth a sentence in your pre-exam instructions; see the A4 sheet guide for the full briefing list.

Confidence, and why review still matters

A grading system that never says “I am not sure” is not more accurate — it is just quieter about being wrong. The useful behaviour is to flag the ambiguous cases and put them in front of the teacher, with the captured image, so a human resolves the handful that genuinely need judgement.

In practice a clean batch produces almost no flags, and a batch scanned in bad light produces many, which is itself a signal to rescan under a different lamp. Keeping the scanned sheet image is the other half of this: when a student contests a mark three weeks later, the picture settles it in seconds.

Practical test of any OMR product: feed it a deliberately ticked-not-filled sheet. A trustworthy reader flags it. A careless one silently records a wrong answer as a confident one.

Where AI actually helps teachers

The genuine value is not in reading the bubbles — that problem is solved. It is in what happens to a term’s worth of results afterwards, where pattern-finding across many attempts tells you things no single answer sheet can:

  • Per-question difficulty. Which questions the class collectively failed, separating a badly taught topic from a badly worded question.
  • Topic mastery. With topic-tagged answer keys, marks roll up per chapter, so the revision lecture targets the gap instead of the syllabus.
  • Distribution, not averages. A mean of 62% hides whether you have one coherent class or two groups needing different lessons. The distribution shows it at a glance.
  • Trend per student. A steady 70% and a fall from 85% to 70% look identical in one exam and completely different across five.
  • Cohort comparison. For a school, averages per class and per teacher show where practice differs — useful for support, and easy to misuse as a ranking, so treat it carefully.

In ScanScore this analysis runs on the device for your own exams, and across the whole institution in the web dashboard once cloud sync is on.

What automated evaluation cannot do

  • Mark written answers defensibly. A language model can draft feedback on an essay, but marks produced that way are not reproducible and cannot be justified to a parent or an appeals committee. ScanScore deliberately limits auto-marking to objective formats — MCQ, multiple-answer, true/false and numeric — where the key is unambiguous.
  • Rescue a bad question. If two options are both defensible, automation records the disagreement faithfully and calls it wrong.
  • Read a sheet that was never designed to be read. Geometry is not optional.
  • Tell you why. Analytics point at the topic; the diagnosis is still teaching work.

Which is the useful framing overall: automation removes the mechanical hours from assessment — the checking, the tallying, the typing into a register — and leaves the judgement where it belongs. If you are choosing between paper and browser delivery for that, our format comparison covers the trade-offs.