Key takeaways
- MIT Media Lab research found that people writing essays with ChatGPT showed the weakest brain connectivity of any group, and 83 per cent could not accurately recall work they had just submitted.
- The researchers call this cognitive debt: mental effort borrowed from the future. It compounds on top of the forgetting curve, where up to 80 per cent of unreinforced learning is lost within a month.
- AI detection is not a solution. OpenAI withdrew its own classifier, and 2026 research found commercial detectors still unreliable for integrity decisions.
- Draft checkpoints and process evidence have a short shelf life, because learners can prompt models to manufacture a plausible improvement arc.
- The fix is design, not policing: application-based assessment, brain-first sequencing, spaced retrieval campaigns, manager-verified capability and interaction types that resist copy and paste.

I have spent more than fifteen years designing workplace learning, and I have never seen completion data look this good while capability data looks this uncertain. Learners are finishing eLearning modules faster. Written assessments are coming back polished. Reflection tasks read beautifully. And an increasing number of those learners, if you asked them a week later, could not tell you what they submitted.
The reason is no mystery. AI assistants are now doing a meaningful share of the thinking in workplace learning. Quiz answers pasted into a chatbot, assignments drafted by a model, reflective journals generated in seconds. The learner gets the tick. The organisation gets the completion record. Nobody gets the learning.
For a while this was anecdote and instinct. Now there is research that puts numbers, and brain scans, behind it.
What the MIT study found
In 2025, researchers at the MIT Media Lab released a study with a title that has stuck: Your Brain on ChatGPT: Accumulation of Cognitive Debt. They took 54 participants and split them into three groups to write essays over several sessions. One group used ChatGPT, one used a search engine, and one worked with nothing but their own heads. Every participant wore an EEG headset measuring brain activity across 32 regions while they wrote.
The results were stark. Brain connectivity scaled down with the amount of external help. The brain-only group showed the strongest and most widely distributed neural networks. Search engine users sat in the middle. The ChatGPT group showed the weakest connectivity of all, and over successive sessions they got progressively more passive, with many resorting to straight copy and paste by the end.
The finding that should worry every L&D professional is the recall data. When tested shortly after writing, 83 per cent of the ChatGPT group could not accurately quote or recall key points from essays they had just submitted. The brain-only group had no such trouble. The AI users also reported the lowest sense of ownership over their own work. They had produced something, but they had not encoded it. The task was done; the learning never happened.
The researchers called this pattern cognitive debt: borrowing mental effort from the future. And like financial debt, it compounds. In a fourth session, participants swapped conditions. Those who had leaned on ChatGPT for three sessions and were then asked to write unaided showed under-engaged neural activity, as if the habit of offloading persisted after the tool was taken away. Meanwhile, those who had done the hard thinking first and were then given AI access performed well, showing strong recall and engagement. The sequence mattered enormously: brain first, then AI, worked. AI first did not.
Fair caveats apply. It is a preprint with a modest sample, focused on essay writing rather than workplace tasks, and the authors themselves push back on lazy “brain rot” headlines. But the direction of the evidence lines up with what learning science has told us for decades: memory is built through effortful processing. Remove the effort and you remove the encoding.
Why this matters for workplace learning
Translate that lab into your organisation. If a nurse completes their annual clinical governance module by pasting the scenario questions into a chatbot, the LMS records a pass, but the knowledge that module exists to build is simply not there. If a frontline banker completes vulnerable customer training the same way, the organisation carries a compliance record that describes capability it does not possess. In regulated environments, that gap is not an academic problem. It is a risk sitting on the register with nobody’s name against it.
The uncomfortable truth is that most workplace learning was already vulnerable before AI arrived. Multiple-choice quizzes, text-heavy modules and written reflections were always weak proxies for capability. AI has not broken assessment; it has exposed how little some assessment ever measured. A completion certificate has always been evidence that a course was finished, not that anything was learned. AI has just made the distance between those two things impossible to ignore.
Why policing and detection will not save you
Education systems are moving on this right now. On 10 August 2026 the NSW Deputy Premier and Education Minister, Prue Car, asked NESA to consider a moratorium on unsupervised take-home assessment tasks, pending a broader review of AI’s effect on student learning. Half of an HSC result comes from school-based assessment, some of it completed at home. Before that, the standard defence was process evidence, and NESA’s own guidance still recommends it: submit drafts at critical points, keep a process diary or logbook, hand in your working alongside the finished piece. Multiple draft checkpoints. It is a genuinely clever idea, because it assesses the journey rather than just the artefact, and that is the right instinct.
It also has a short shelf life, because students are at least as clever as the controls. I have seen this solved first-hand, and quickly. A capable student prompts the model to write at the standard of a C-grade student a year or two below their real level, seeds the text with plausible typos and clumsy sentence construction, then asks for a deliberately weaker first draft followed by a stronger second, manufacturing exactly the improvement arc the checkpoint was designed to look for. There is also a whole category of “humaniser” tools whose entire function is to strip the statistical fingerprints out of generated text, along with prompt patterns that ask the model to avoid the well-known tells: the tidy three-part lists, the “it is not just X, it is Y” construction, the relentless em dashes, the vocabulary that reviewers have learned to spot.
I am not going to publish the prompts. The point is that they exist, they take about thirty seconds to find, and any assessment control that depends on the learner not knowing about them is borrowed time.
Detection does not rescue it either. OpenAI withdrew its own AI text classifier after it correctly identified only about a quarter of AI-written text while flagging roughly one in eleven pieces of human writing as machine-generated, and a 2026 systematic review in the International Journal for Educational Integrity found commercial detectors still too unreliable to support definitive integrity decisions. In a workplace context, wrongly accusing an employee of faking their compliance training is a considerably worse outcome than the faking.
Draft checkpoints are still worth doing, and moving high-stakes exams into supervised conditions is a rational response for a school system. But workplace L&D cannot invigilate. We have no exam hall, no supervision, and no appetite for surveillance. Which means we cannot police our way out of this. We have to design our way out.
This is not a morality tale
Let me be clear about something, because L&D can sound preachy on this topic. Using AI to produce work is not, in itself, a bad thing. I use it constantly. It saves time. It creates flexibility. It takes the things that only need to be adequate and gets them done faster, which frees up capacity to take the things that matter most from eight out of ten to ten out of ten. That is a real productivity gain and pretending otherwise makes our profession look ridiculous.
The problem is narrower than “AI is bad”. Most work is an artefact, and it does not much matter how the artefact got made. A learning task is not an artefact. It is a repetition. Outsourcing your reps at the gym does not make you stronger, and outsourcing the thinking in a learning task does not build the capability the task existed to build. The MIT crossover finding supports this distinction rather than contradicting it: participants who did their own thinking first and then used AI performed well, because for them the tool was amplifying judgement rather than replacing it.
So the obligation lands on us, not on the learner. If a module can be bypassed with a chatbot in four minutes, that is not evidence of a workforce integrity problem. It is evidence that the module was never doing much, and AI has simply made that visible. Our courses now have to be better than they have ever been: more engaging, more challenging, more genuinely worth an hour of someone’s working day, and designed with enough friction in the right places that the thinking cannot be handed off. And they have to keep going well past the first instance. A single eLearning module or a one-day workshop was never sufficient on its own. It is now barely a starting point unless it is the first touchpoint in something longer and more deliberate.
Cognitive debt meets the forgetting curve
Cognitive debt does not arrive on its own. It compounds on top of a problem workplace learning already had. Hermann Ebbinghaus mapped the forgetting curve in the 1880s, and Murre and Dros replicated his numbers to within a few percentage points in 2015. Without reinforcement, people lose roughly half of new information within a day and up to 80 per cent within a month. A single one-hour compliance module, completed once a year, was never going to survive that curve.
Now add AI. If the encoding was weak to begin with because the model did the processing, there is very little memory for the forgetting curve to erode. You are not losing 70 per cent of what was learned. You are losing 70 per cent of almost nothing. This is why we now push clients away from one-off modules and towards learning campaigns: a designed sequence of touchpoints spread across weeks rather than a single event. Spaced retrieval is the best-evidenced intervention in the field. Cepeda and colleagues’ meta-analysis of 317 experiments confirmed that distributed practice consistently beats massed practice, and Roediger and Karpicke’s work on the testing effect showed that being made to retrieve information, rather than simply re-reading it, dramatically slows forgetting.
In practice, a campaign might look like this: a short pre-work provocation, the core module, a manager conversation in week one, a two-minute retrieval check at day three, an application task at week two, and a scenario challenge at week six. Each touchpoint is small. Together they force repeated retrieval, and repeated retrieval is exactly the thing AI shortcuts remove.
Designing courses people cannot skim
The other shift in our studio work has been visual. We are building modules that are far more visually considered than they were three years ago, and not for decoration. Strong visual design interrupts the scan. When a screen is a wall of text with a Next button, the rational move for a busy learner is to scroll and click. When a screen presents a striking image, a real story, a decision point or a piece of data that demands a moment of interpretation, the learner has to stop and process. Interruption is the design goal.
Alongside that, we are increasing the density of interaction in our eLearning design: more knowledge checks embedded through the content rather than banked at the end, more scenario activities, more self-reflection prompts, and far more narrative. Stories are particularly effective here, because a well-told account of something that went wrong in that industry creates an emotional hook that a bulleted policy summary never will. People remember what they felt.eLearning design
In the room: placemats, pens and a device agreement
Face to face, we have gone in an unfashionable direction: back to paper. Our classroom designs now use printed placemats, large-format worksheets that sit on the table in front of each participant and get written on all day. Participants capture their own thinking by hand as the session moves, and they walk out with an artefact in their own handwriting rather than a slide pack they will never open. It is deliberately analogue and it works.
The evidence supports it. A high-density EEG study by Van der Weel and Van der Meer recorded 36 university students across a 256-channel sensor array and found that handwriting produced widespread brain connectivity in the theta and alpha bands, the patterns associated with memory formation, while typing did not. Earlier, Mueller and Oppenheimer’s “The Pen Is Mightier Than the Keyboard” found that longhand note-takers understood conceptual material better than laptop note-takers, because writing by hand is slower and forces you to summarise rather than transcribe verbatim. The handwriting research has drawn some methodological criticism, and the effect sizes should not be oversold, but the underlying mechanism is the same one the MIT study points to: friction produces processing.
The second change is the device agreement. At the start of a face-to-face day we now negotiate a group agreement with the room: phones and laptops away, not face down on the table. It has to be negotiated rather than imposed. When the facilitator dictates it, people comply resentfully and check under the table. When the group sets the rule themselves, including the facilitator, peer accountability does the enforcement and nobody has to police anything.
The part that makes it work is the trade. We build three scheduled 20-minute digital breaks into the day, at morning tea, lunch and afternoon tea, where everyone is free to deal with email, messages and calls. Most covert phone-checking is anxiety rather than disinterest. People are worried about what they are missing. Give them a guaranteed window and the anxiety drops, and with it the compulsion to sneak a look during a discussion. We were braced for resistance the first few times we proposed this. What we got instead was relief, and consistently better discussion quality in the room.
The problem with the mandatory quiz
We are also actively arguing with clients about the end-of-module quiz. The standard model is ten multiple-choice questions with an 80, 90 or 100 per cent pass mark, and a retry loop until everyone gets through. It is easy to build, easy to report on, and almost entirely uninformative. It tests whether someone can recognise a correct statement, not whether they can apply a skill under pressure with a real person in front of them. It is also the single easiest thing in workplace learning to outsource to AI.
The alternative we advocate is application-based assessment supported by manager involvement. Increasingly our deliverables include a manager toolkit: a structured guide that equips team leaders to run a short conversation with each staff member after the module. The manager asks them to talk through a real situation they have faced, explain the decision they made, and reason out what they would do differently. That conversation reveals in about five minutes whether genuine processing occurred, in a way no quiz score can.
The toolkits now include extension prompts, deliberately written to push people past their first answer. “What would change your mind about that?” “Who else is affected that you have not mentioned?” “Walk me through the moment you would escalate.” These questions activate reasoning and judgement rather than recall, and they cannot be pre-generated because the manager is responding live to what the person says.
Ten ways to reduce AI substitution in eLearning
This is the working list we take into design conversations with clients.
- Sequence brain-first. Require the learner to commit their own answer, decision or draft before any AI support or model answer is available. The MIT crossover data suggests this ordering is the difference between AI that amplifies thinking and AI that replaces it.
- Assess application, not recall. Replace “which of the following is correct” with “here is a messy situation from your own workplace, what do you do and why”. Reasoning is much harder to outsource convincingly than a factual answer.
- Anchor tasks in the learner’s own context. Ask for a real incident from their own team, a specific client, or their own site process. A model cannot know these things, so a generated answer becomes obvious to anyone reading it.
- Distribute knowledge checks through the content. Frequent short checks in the flow of learning are harder to batch-solve in a separate tab than a single quiz sitting at the end, and they force retrieval while the material is still being encoded.
- Use interaction types that resist copy and paste. Drag and drop sequencing, hotspot identification, branching simulations, timed decisions, video-based response and audio scenarios all require the learner to act rather than produce text.
- Design visually to interrupt scanning. Strong imagery, data that needs interpreting, and screens built around a single decision point stop the scroll-and-click reflex that AI substitution depends on.
- Build in a manager conversation. A five-minute structured discussion with a team leader, supported by a manager toolkit and extension prompts, verifies understanding better than any automated score and cannot be pre-generated.
- Add spaced retrieval after the event. Short unannounced check-ins at day three, week two and week six change the incentive during the course. If learners know they will need to recall and apply without tools later, offloading looks less attractive.
- Ask for reasoning and confidence, not just answers. Have learners explain why the distractors are wrong, or rate their confidence and justify it. Metacognitive prompts expose shallow processing quickly.
- Bring AI in openly as a coach. Prohibition does not work. Give learners a legitimate, well-designed way to use AI, such as critiquing their own draft or role-playing a difficult conversation, so the tool becomes a practice partner rather than a workaround.
Two further points worth noting: peer discussion and small-group debriefs add a social layer of accountability that individual eLearning lacks, and changing what you measure matters as much as changing what you build. Completions and satisfaction scores tell you nothing in an AI-saturated environment. Behaviour change, on-the-job application and manager-verified capability are the only metrics that still mean something.
AI is not going away, and it should not. But learning only happens in the head of the learner, and no model can do that part for them. Our job as learning designers is to make sure the effort that builds capability still happens, with AI positioned as the coach beside the learner, never the substitute for them.
Australia is now examining this formally: South Australia has launched a royal commission into AI, NSW is reviewing take-home assessment, and a federal Senate inquiry is taking submissions. I have written about what those processes mean for workforce capability in Australia Is Now Formally Worried About How We Think. You can also see how we build multi-touchpoint learning campaigns, job aids that support application on the job, and where we think the ethical lines sit for AI in learning design.
