UK Tests AI-Generated Writing Samples for Student Assessment
The Department for Education is using ChatGPT to create writing samples for literacy testing, raising concerns about fairness and authenticity.
The Department for Education is using ChatGPT to create writing samples for literacy testing, raising concerns about fairness and authenticity.
The UK Department for Education is experimenting with artificial intelligence to cut costs in a crucial part of the education system. Instead of using real children’s writing samples to assess literacy standards at key stage 2, moderators will soon evaluate pupils against AI-generated text created by ChatGPT’s GPT-5 model.
The shift sounds practical on the surface. Around 2,000 moderators currently cross-check pupils’ work against authentic writing samples to ensure grades are standardized across the country, a process that costs approximately 100,000 pounds annually. Using AI could reduce that by 95 percent.
But here’s where it gets complicated. The DfE’s own risk assessment acknowledges a troubling problem: AI tends to produce flat, bland writing that excludes atypical vocabulary and sentence structures. This means children who speak English as a second language or are neurodivergent could end up disadvantaged simply because their natural writing patterns don’t match the synthetic standard.
Rebecca Clarkson at Anglia Ruskin University, who studies KS2 writing assessment, articulated the philosophical problem perfectly: “Having an exemplification of writing that is not a real child’s writing creates a philosophical and ethical issue.” Moderators she’s spoken with share this concern. They worry that assessing real children against artificial benchmarks fundamentally misrepresents what good writing actually looks like.
The experiment is launching this year and next, with around 20 experienced local authority managers tasked with reviewing the AI material for authenticity before use. One full standardization exercise in 2026-27 will still include scripts from actual children, giving the department time to evaluate whether the switch makes sense.
Jo-Anne Baird at the University of Oxford raises an even more insidious concern. Without proper governance, AI-generated materials could become the de facto standard that teachers try to get pupils to emulate. “This could lead us into some strange places,” she warns. Once teachers know what the AI values, they might unconsciously train students to match that narrow model of acceptable writing.
The consequences, Clarkson emphasizes, remain unpredictable. “Is it a concern if we don’t use real children’s writing? Who knows? I guess we don’t know what the impact is until we can measure the impact.” That uncertainty should give us pause when we’re talking about educational assessment that affects how much support students receive when entering secondary school.
It’s worth noting that the KS2 writing assessments happen after pupils have already been assigned to their secondary schools, so these benchmarks shouldn’t determine which school a child attends. But the level of educational support they receive depends directly on how teachers perceive their abilities relative to these standards.
This situation represents a broader pattern worth monitoring across science and policy. As AI seeps into assessment systems globally, we’re making decisions about children’s educational trajectories based on shortcuts rather than carefully considered evidence. The 95 percent cost saving is genuinely attractive to cash-strapped education departments, but it comes at an unmeasured cost to educational authenticity and equity.
The DfE plans to decide in spring 2027 whether to continue this approach, but that timeline feels rushed for evaluating something this consequential. We should demand that decisions affecting how we measure student achievement rest on thorough research and genuine stakeholder input, not just budget pressure.
When AI starts defining what we consider normal, acceptable, or excellent, we’re outsourcing our professional judgment to algorithms trained on patterns in existing data. For education, that’s a gamble we should think carefully about taking.
Source: New Scientist