Education has never had more data about learning.
Every platform produces it. Students generate scores, levels, completion rates, response times, assignments, assessments, transcripts, engagement measures and increasingly AI-generated analyses of their performance.
The assumption underneath much of this is simple: if we collect enough signals, we can know whether a student is learning.
AI is exposing a problem with that assumption.
Students can now produce remarkably sophisticated work without necessarily possessing the understanding that work appears to demonstrate. They can generate an essay, solve a problem, summarize a reading, write code, answer questions and complete assignments with assistance that is increasingly difficult to detect.
This isn’t simply a new cheating problem.
It’s an evidence problem.
If the work we measure is increasingly disconnected from what a student can actually do, then measuring it more precisely won’t tell us whether learning occurred.
Learning needs to be verified.
All learning data is a proxy
A score isn’t learning.
Neither is a grade, a completed assignment, a platform level, a transcript or an AI analysis.
They are evidence from which we make an inference about learning.
Usually, that is useful. Sometimes it is extremely useful.
But the distinction matters.
A student can earn a high score without retaining much of what they learned. They can complete an assignment with extensive help. They can perform well on familiar problems and struggle when the same idea appears in a new context.
And now AI can help produce many of the artifacts we’ve traditionally treated as evidence of understanding.
That doesn’t make the artifacts worthless.
It makes the inference less certain.
AI makes the appearance of learning abundant
For years, schools have worried about students finding answers online.
Generative AI is different.
It doesn’t just provide the answer. It can participate in the intellectual work.
It can brainstorm, explain, outline, revise, calculate, translate, summarize, write and reason alongside the student.
Used well, that can be enormously valuable for learning.
But it also makes the finished product increasingly ambiguous.
Did the student write that paragraph?
Perhaps that’s no longer even the right question.
The more important question may be:
What can the student understand and do without the artifact speaking for them?
That’s much harder to determine from the artifact itself.
Learning has to be observed
At some point, educators need to see students demonstrate what they know.
Can they explain an idea in their own words?
Can they apply it when the problem changes?
Can they defend a conclusion?
Can they connect concepts?
Can they create something new?
Can they recognize when their approach isn’t working and adapt?
Can they transfer what they learned into a situation they haven’t seen before?
These aren’t perfect measures of learning either.
But they bring us closer to the thing we’re actually trying to understand: what capabilities now exist in the student?
That means assessment can’t happen entirely through products submitted after the fact.
More demonstrations of learning may need to move back into the classroom, where educators can observe the process as well as the result.
But one teacher can’t observe everyone
That creates an obvious practical problem.
Put 30 students in a classroom and ask them to demonstrate understanding through discussion, explanation, problem solving, creation or application.
A teacher can watch a few closely.
They cannot watch everyone.
Small groups make this even harder. Some of the richest evidence of learning may be happening simultaneously around the room while the teacher is working with one group.
Historically, most of those moments simply disappeared.
The teacher saw what they could see.
Everyone else eventually submitted something that could be scored.
Recording changes that constraint.
Students can demonstrate learning while the teacher is somewhere else in the room. Educators can revisit selected moments afterward. They can sample evidence, investigate uncertainty, compare performance over time, or return to something that deserves another look.
Recording doesn’t replace teacher observation.
It extends it.
That changes what we should collect
If our goal is verification, schools may need a different evidence base.
Instead of collecting primarily the outputs that are easiest to score, we can collect more evidence of students actually doing the things we hope they are learning to do.
Not just the essay, but the student explaining its argument.
Not just the answer, but the student working through an unfamiliar problem.
Not just the presentation, but the discussion that tests whether the ideas hold up.
Not just an AI-generated analysis of performance, but evidence a teacher can examine and interpret for themselves.
That evidence doesn’t all need to become permanent data.
Much of it may only need to exist long enough for an educator to make a better judgment.
The goal isn’t to build a permanent record of every student.
The goal is to give teachers enough evidence to know when learning is real.
Verification makes teacher judgment more important
There is an irony in all of this.
AI may automate more instruction. It may generate more assessment. It may produce more data and analyze that data more effectively than any teacher could.
Yet that may make teacher judgment more important.
Someone still has to determine whether the signals correspond to reality.
Someone has to look beyond the polished artifact, score or dashboard and ask whether the student can actually understand, explain, create, apply and transfer what they know.
AI can help educators examine the evidence.
Recording can help them capture it.
Data can help identify where to look.
But ultimately, learning isn’t something we can simply measure into existence.
It has to become visible in what a student can actually do.
And increasingly, educators will need the evidence to verify it.
