Artificial intelligence tools typically award higher marks on average than humans and cannot be relied on to give an accurate indication of a students’ performance, a paper has claimed.

With AI increasingly being explored by universities as part of the marking process to relieve pressure on time-strapped academics, the research, published in the journal Assessment & Evaluation in Higher Education, found that generative AI tools such as ChatGPT “do not reliably reproduce human judgement in the marking of extended written work”.

The researchers uploaded 50 undergraduate bioscience essays to two versions of ChatGPT and asked it to grade the essays against seven assessment criteria under four different prompting conditions, finding there were “significant discrepancies” between AI-marked essays and those marked by humans.

In all but one case, the AI models typically returned higher average marks than humans. In one case the difference between an AI grade and a human one was 40 marks for an essay where the top score was 100.

Lower-scoring essays tended to receive inflated marks, whereas higher-scoring essays received lower marks compared with human assessment. AI was more aligned with human markers on mid-grade essays.