Junior developers using AI score better on coding tasks. They also retain significantly less of what they just built. A new study from the University of New South Wales puts numbers on a problem engineering leads have been quietly worrying about.
What the Study Found
Researchers gave 55 undergraduate computer science students three introductory C programming tasks. Half used ChatGPT-4.5. Half used conventional web search. The AI group scored 89% on the tasks. The search group scored 69%.
Then came the knowledge retention test, run immediately after submission. The AI group scored 41%. The search group scored 53%. Two days later, the AI group dropped to 39% while the search group held at 52%.
The students explained the gap themselves. Those using ChatGPT estimated only 45% of the submitted code was actually theirs. The search group put that number at 81%.
The Pipeline Problem
The study was small and limited to one university’s undergraduates, not developers inside real engineering teams. But the question it raises is uncomfortable for engineering leaders: if juniors can ship code they don’t understand before they’ve built the knowledge to debug it, what happens to the senior engineer pipeline five years from now?
What One Engineering Lead Is Doing
Adrian Harwood, head of research software engineering at the University of Manchester, says banning the tools is not realistic. His department takes undergraduates on year-in-industry placements, and that cohort is already enthusiastically using generative AI.
His approach: never let juniors work solo. Senior engineers review every pull request, including AI-generated code, and flag problems the junior doesn’t yet have the experience to catch. The learning shifts from writing code from scratch to what Harwood calls learning how to police the tool: reviewing outputs and understanding why certain choices cause problems downstream.
Harwood draws a direct line between this and what engineering organizations actually need to develop.
Writing code is one part of the job. Understanding requirements, building sensible specifications, verifying what was built, and diagnosing failures are the rest. His argument is that junior programs should spend more time on system failure modes and critical output assessment, regardless of whether the output came from a human or a model.“I’m trying to train engineers in my department. I’m not trying to train programmers.”
The UNSW findings back the concern that output volume is a misleading metric when juniors are shipping code they can’t necessarily explain or fix.
