How to Avoid Bias in AI Assessment

A teacher's hand marking schoolwork with a red pen

Bias in marking isn’t a new problem, and it isn’t unique to AI. Human assessment can be influenced by interpretation, inconsistency or expectations for reasons that have little to do with the actual quality of the work.

So when AI enters the marking process, the question isn’t whether bias suddenly becomes a risk. It’s how you reduce that risk in a different kind of system. For Olex, that means rigorous testing, teacher-defined context, automated safeguards, human review and continuous revalidation.

Testing accuracy is only half the job

We benchmark our models against independently marked and moderated assessment materials, including statutory exemplars, local authority moderation materials and other publicly available assessment resources.

Outputs are checked against the original marking decisions and reviewed for correct interpretation of assessment criteria and faithful use of evidence from anonymised pupils’ work.

But accuracy alone isn’t enough.

A model can perform accurately overall while still behaving inconsistently in particular circumstances. That is why our testing also looks at marking consistency, directional marking trends, tone, accessibility, unsupported assumptions, and the effect that changes in prompts or context can have on an output.

We also use our pre- and post-validation tools to monitor whether a model is systematically marking too generously or too critically in particular contexts, including different subjects, paper types and levels of input or marking risk.

Where we identify a trend, we investigate the cause, compare model and prompt behaviour, and recalibrate or retest before applying changes more widely.

Read more about testing and accuracy at Olex.

What fairness actually means in practice

Context matters enormously when AI is used in education.

Olex is designed so that teachers provide the educational context the AI needs to work within. That might include the assignment question, supporting material, success criteria, marking expectations or instructions about the format and focus of feedback.

Teachers can also use that context to differentiate materials or feedback. For example, they may want additional scaffolding for learners with EAL, resources aimed at a particular target grade, or feedback presented in a format that better supports an individual learner or cohort.

The important point is that these decisions come from the teacher.

Olex does not need to build a personal profile of a pupil in order to decide how they should be assessed or supported. Teachers determine what educational context is relevant and remain in control of the decisions that matter.

Alongside that teacher-defined context, we use strict prompts and guardrails designed to reduce unsupported assumptions and keep outputs focused on information relevant to the task.

You can read more about our approach to safe AI in schools here.

Moderation and human review

Prompt design is only one safeguard.

Olex also uses moderation and diagnostic tools across the platform. These can identify potentially inappropriate written content, while generated images can be withheld where our moderation systems identify that an output may have crossed an appropriate threshold.

These automated controls are supported by human oversight.

Our team reviews outputs across the Olex ecosystem every day, including assessment behaviour, generated content, feedback and imagery. This helps us identify patterns or edge cases that automated testing alone may not capture.

Real-world feedback from schools provides another important source of evidence. To date, bias has not emerged as a recurring issue through customer feedback, but we do not treat that as a reason to stop testing.

What happens when something doesn’t look right

When an output raises a concern, we investigate it rather than assuming the cause.

We try to reproduce the output using the same context and sample, then compare different models, prompts and instructions to understand what is driving the behaviour. The issue might relate to model behaviour, prompt design, ambiguity in the context provided to the system, an assumption introduced by a human, or simply an unusual edge case.

Where we identify something that needs changing, we refine prompts or guardrails and retest them against the original sample and comparable examples before applying a change more widely.

The aim is not simply to label an output as “biased” or “unbiased”. It is to reproduce it, understand the cause, make controlled changes and test again.

Bringing in outside scrutiny

It’s easy for any organisation to mark its own homework.

So we don’t only rely on internal testing. We actively engage recognised specialists in AI bias and education, including Victoria Hedlund, to challenge our assumptions, review our outputs, and stress-test the system against emerging bias risks we might not have thought to look for ourselves.

External challenge matters because responsible testing should not simply confirm what you already believe about your own system. It should help identify the questions you haven’t yet asked.

The bar we’re holding ourselves to

There is no single test that can guarantee an AI system will never produce an inappropriate or biased output.

Our approach is deliberately layered: teachers define the educational context, guardrails constrain model behaviour, moderation tools check outputs, humans review the system daily, validation tools monitor marking trends, and models, prompts and workflows are continually retested as they change.

Bias testing in AI assessment isn’t a box to tick once. It’s an ongoing discipline of testing, review, investigation and refinement, with teachers retaining control over the educational decisions that matter.

Share This Post