Education technology is becoming increasingly good at producing answers.
A learner is struggling. The system identifies a pattern. An AI tool generates feedback. A dashboard highlights a concern. An adaptive platform recommends another activity. A progress system predicts that intervention may be needed.
All of that can be useful, but we should ask a question before being impressed by the output: Can we still see enough of the evidence to make a good educational judgement ourselves?
Because in education, a recommendation is rarely the end of the process. It is usually the beginning of a decision, and the better the technology becomes at producing confident-looking outputs, the more important it becomes to preserve a visible trail between:
what happened;
what the technology identified;
what evidence supports that interpretation;
what its limitations are;
and
what a human decided to do next.
That is not resistance to AI. It’s one condition for using it responsibly.
Education contains judgement everywhere
Consider a fairly ordinary tutoring session. A learner answers six questions incorrectly. That fact can be recorded automatically.
But what does it mean?
Perhaps they do not understand the concept.
Perhaps they understand the concept but have misread the question.
Perhaps one missing prerequisite is affecting several answers.
Perhaps they are rushing.
Perhaps they understood yesterday and have forgotten today.
Perhaps the question format has changed.
Perhaps the learner is tired.
The data tells us something happened. The educational judgement tries to work out what that something means.
Then comes another judgement:
What should happen next?
Practise?
Re-teach?
Change the explanation?
Go back to a prerequisite?
Reduce the difficulty?
Increase the challenge?
Ask another question?
Do nothing yet because there is not enough evidence?
Technology can help with all of these stages, but help should not mean hiding them.
Learner evidence informing technology-supported analysis before a teacher makes the final educational decision.
An answer is not the same as an auditable decision
Imagine an education platform displays: “Student requires additional support with fractions.”
That may be correct, but an educator needs more.
Why has the system reached that conclusion?
Which work is relevant?
Is the pattern persistent or based on one activity?
Were the questions measuring the same skill?
Was assistance already provided?
Is the learner making one repeated error or several unrelated ones?
Does the system know what the learner is currently doing at school?
How confident should we be in the recommendation?
A useful system does not have to reveal the hidden internal workings of an AI model. That is not the point. What it should provide is enough external evidence and context for a human to evaluate the output.
For example:
Evidence observed: The learner answered 4 of 6 fraction-comparison questions incorrectly.
Pattern: In all four errors, they compared denominators directly rather than considering equivalent values.
Context: The same misconception appeared in the previous session.
Suggested action: Revisit fraction magnitude using visual models before further procedural practice.
Now the teacher or tutor has something they can interrogate. They can agree. Disagree. Look at the learner's work. Ask another question. Try a different explanation. Or decide that the recommendation misses something important.
That is far more useful than a confident instruction produced without context.
Comparison between a black-box recommendation and an education system showing evidence, limitations and human review.
The DfE position is already clear about professional judgement
Current Department for Education guidance on generative AI in education makes an important distinction. AI can make some tasks faster and easier, but professional responsibility does not transfer to the tool. Staff are expected to critically check outputs for accuracy and appropriateness, and the final content remains the responsibility of the professional and their organisation.
The guidance is explicit that generative AI cannot replace the judgement and deep subject knowledge of a human expert. That is the right principle.
The purpose of technology should be to increase the quality or efficiency of human judgement, not make judgement disappear behind an interface.
Good transparency is practical
“AI transparency” can sound abstract. In education, it can be very practical. A teacher does not necessarily need to know how millions or billions of model parameters generated an answer.
They do need to know things such as:
What information went into this recommendation?
Which learner evidence is it based on?
What assumptions has the system made?
Is anything important missing?
How recent is the information?
Is this a factual statement, an inference or a prediction?
Can I inspect the underlying work?
Can I override the recommendation?
Will my decision be recorded?
Those questions make technology easier to trust appropriately. Not blindly. Not suspiciously. Appropriately.
Confidence should not be confused with correctness
One of the particular challenges of generative AI is presentation. An uncertain answer can sound confident. A plausible explanation can be wrong. A polished progress summary can make weak evidence look authoritative.
This matters in education because presentation quality can easily be mistaken for educational validity.
Consider an AI-generated comment:
The learner has shown strong progress in analytical writing and is now ready to move to more advanced evaluative tasks.
It sounds professional. But what evidence supports it? One essay? Three? Teacher assessment? Automated scoring? Did the learner write independently? Has performance improved over time? Would the classroom teacher agree?
Without that context, we have fluent language rather than necessarily useful evidence. Technology that summarises learning should help users inspect the evidence beneath the summary.
The human decision should remain visible
There is another side to traceability. It should not stop at: “The system recommended X.”
We should also be able to see: “The teacher reviewed X and decided Y.”
Suppose an analytics system flags that a learner's performance is declining. The teacher reviews the work and realises that the latest assessment covered a new topic and that the learner has missed several lessons because of illness. The teacher decides not to assign a remedial intervention immediately. Instead, they provide missed material and review performance again the following week.
That judgement is valuable. If only the automated alert remains visible, we lose part of the educational story.
A stronger system might preserve:
Technology observation: recent scores declined.
Relevant context added by educator: learner missed introduction to new topic.
Decision: provide catch-up teaching before further intervention.
Review point: reassess next week.
Now technology is not simply producing decisions. It is helping create a record of professional reasoning.
This matters particularly when several adults are involved
School. Tutor. Parent. SENDCO. Pastoral support. Specialists. Different adults may see different pieces of evidence.
Technology can help connect those pieces, but it can also make fragmentation worse if each system produces its own independent conclusion.
Imagine:
The school system says the learner is below target.
The tutoring platform says progress is good.
An adaptive learning tool says the student should move ahead.
The parent sees the learner spending two hours struggling with homework.
All four statements could technically be true. The question is how they fit together.
A visible trail helps people ask:
What evidence is each conclusion based on?
Are we measuring the same thing?
Over the same period?
With the same level of support?
What does the learner actually need next?
Without that, more data can produce less clarity.
Automated personalisation still involves choices
“Personalised learning” sounds inherently positive, but every personalised recommendation embeds decisions.
Which objective should be prioritised?
What counts as mastery?
When should a learner move forward?
How much support should they receive?
Which errors matter most?
What level of challenge is appropriate?
Those decisions may be made partly by educators, curriculum designers, product teams and algorithms. The learner may never see them.
That does not make personalisation wrong, but it means schools should ask how those decisions are governed.
The DfE's current generative-AI product safety standards require products to state their intended educational purpose and target users clearly, and say suppliers should support claims about their products with robust and transparent evidence.
That is important. A system that claims to improve learning should be able to explain what educational problem it is designed to solve and what evidence supports the claim.
AI should expose uncertainty rather than erase it
Education is full of incomplete information. Sometimes we genuinely do not know yet.
Why has this learner's performance changed?
Is an intervention working?
Is the difficulty conceptual or motivational?
Is this one weak result or the beginning of a pattern?
AI systems often feel useful because they reduce ambiguity. They summarise. Categorise. Rank. Recommend.
But educational uncertainty is not always a technical flaw waiting to be removed. Sometimes uncertainty is the correct conclusion. A responsible system should be able to support statements such as:
There is not enough evidence yet.
This pattern may have more than one explanation.
The recommendation depends on assumptions that should be checked.
Human review is required.
Technology becomes dangerous when uncertainty disappears from the interface while remaining present in reality.
Learners need visibility too
The principle also applies directly to students. Suppose an AI tutor tells a learner: “You should revise quadratic equations next.”
A stronger learning interaction might make the recommendation understandable: “You solved the procedural questions successfully, but made errors in three questions requiring you to identify which method to use. Let's practise recognising the problem type before doing more calculations.”
Now the learner understands why the next step exists. That matters for independence.
A student should gradually become capable of asking:
What am I trying to improve?
What evidence shows I need to improve it?
What strategy should I try?
Did it work?
If technology constantly makes those decisions invisibly, personalisation may inadvertently reduce the learner's involvement in their own learning. The best technology should help develop judgement, not monopolise it.
Feedback is particularly important
AI-generated feedback is one of the most obvious educational uses of generative technology.
It can potentially make feedback:
faster;
more frequent;
more personalised;
and easier to produce.
But feedback is only valuable if it helps the learner improve. An automatically generated paragraph of comments may look impressive while creating little learning value.
The useful questions remain familiar:
What is the learner trying to achieve?
What evidence does the feedback refer to?
Is the feedback correct?
Is it specific enough to act on?
Does the learner understand it?
What should they do next?
Will somebody know whether they acted on it?
The technology may accelerate the production of feedback. Educational judgement determines whether the feedback is worth having.
The same principle applies beyond AI
This is not only an AI issue. Traditional learning analytics can create the same problem. A dashboard turns red. A risk score changes. A progress line falls. A learner moves into a different category.
We naturally begin treating the visualisation as the reality. But dashboards contain decisions too. Someone decided:
which data matters;
how it is weighted;
what threshold creates an alert;
what “expected progress” means;
which comparison group is relevant;
and what gets left out.
The strongest education technology helps users move from the visual signal back to the underlying evidence.
A red indicator should mean: look here.
Not: the system has already decided what this learner is.
Technology should make professional conversations better
Perhaps one of the best measures of an education system is what happens when people discuss the output.
Does technology close the conversation? - “The system says…”
Or does it improve the conversation? - “The system has identified this pattern. Here's the evidence. Does that match what you're seeing?”
That second interaction is much more powerful. The teacher contributes classroom knowledge. The tutor adds individual observation. The parent provides context from home. The learner explains their experience.
Technology helps them identify patterns they might otherwise miss. Different evidence can then improve the decision. That is augmentation in a meaningful sense.
Traceability should not become bureaucracy
There is an obvious risk. If every educational decision requires a lengthy audit log, we will create another administrative burden. That would defeat much of the point of technology.
A useful trail can be simple. For many situations, four elements may be enough:
Evidence: What happened?
Interpretation: What appears to matter?
Decision: What are we doing next?
Review: What would tell us whether it worked?
That is not a compliance exercise; it’s a learning loop. And technology should make it easier to capture, not harder.
Product design matters
For education technology providers, this has implications. Do not only design the answer. Design the path around the answer. Allow educators to inspect relevant learner evidence. Distinguish observation from inference. Show limitations where they matter. Make recommendations editable. Allow professional context to be added. Preserve human overrides. Connect actions with later outcomes.
Make it possible to ask: What happened because we acted on this recommendation?
That question is particularly important; otherwise, systems become very good at recommending interventions without ever learning whether the recommendations helped.
The strongest AI system may sometimes recommend nothing
We tend to assume every technology interaction should produce an action.
Another resource.
Another intervention.
Another target.
Another notification.
But sometimes the correct recommendation is:
Wait.
Observe.
Collect more evidence.
Allow the learner more time.
Ask the teacher.
Ask the learner.
Do not intervene yet.
Good educational judgement includes knowing when not to act. Technology should support that possibility too.
Human-centred AI means preserving agency
UNESCO's AI competency frameworks for both students and teachers place human agency and human-centred use at the heart of AI capability. That is important because AI literacy should not be reduced to knowing how to write prompts.
For learners and educators alike, competence increasingly includes knowing when to:
use AI;
question it;
verify it;
reject its suggestion;
modify its output;
and take responsibility for the final judgement.
That is a more demanding skill than simply operating a tool; it’s also far more valuable.
The test for learning technology
When evaluating a learning technology, perhaps we should ask more than: Does it produce useful output?
Ask:
Can we see what evidence matters?
Can we understand what the recommendation is based on?
Can a human challenge it?
Can relevant context be added?
Can the decision be changed?
Can we later see whether the decision worked?
If the answer is yes, technology can strengthen educational judgement.
If the answer is no, a polished interface may simply be concealing more important questions.
The future of learning technology should not be a system that quietly makes more and more decisions for educators and learners. It should be technology that helps people make better ones, and leaves enough of the trail visible that we can understand why.
When technology recommends an educational action, what should a teacher be able to see before accepting it?
Sources and further reading
Department for Education — Generative artificial intelligence in education
Current guidance for schools and colleges on opportunities, limitations, professional judgement and responsibilities when using generative AI.
Read the DfE guidanceDepartment for Education — Generative AI: product safety standards
Standards for generative AI products used in education, including intended purpose, evidence behind product claims, safeguarding, security and risks to learner development.
Read the DfE product safety standardsDepartment for Education — Interacting with generative AI in education
Training materials covering the importance of checking AI-generated outputs and understanding issues associated with generative AI.
Explore the DfE training materialsUNESCO — AI competency framework for students
Framework covering human-centred thinking, AI ethics, techniques and applications, system design and critical judgement of AI.
Explore the UNESCO student frameworkUNESCO — AI competency framework for teachers
Framework focused on human agency, ethical AI use, AI pedagogy and the professional capabilities educators need when working with AI.
Explore the UNESCO teacher framework
We are discussing this article on LinkedIn: Good education technology should make judgement more visible, not less
Read more Education Futures articles on the TutorTech blog
For parents looking for trusted learning support: Find Tutors
Join the education movement: Teach. Create. Own.
For qualified teachers interested in joining TutorTech: Join as a Tutor
Read other Education Futures blogs:
Education is being rewritten. Who gets to own the knowledge?
Student finance is not only a Treasury issue. It is an education trust issue.
Why parents are asking different questions about tutoring in 2026
What qualified teachers understand that algorithms often miss
Not every learner needs more content - some need better diagnosis
Young people need pathway literacy, not one-off careers advice
A second choice should not be treated as a second-rate route
Parent engagement should be designed for access, not confidence
The strongest start to a school year is alignment, not activity
Attendance improves when barriers are understood before consequences are applied
Supplementary tutoring should strengthen the classroom, not create a parallel curriculum
Explore more education futures commentary, teacher-led perspectives and AI-in-education insights on the TutorTech blog.
⬅️ Back to Blog