AI-generated grades less likely to be challenged by teachers

AI-generated grades less likely to be challenged by teachers

Teachers are more likely to accept an unfairly harsh grade when they believe it was produced by AI, according to new research. The finding raises questions about whether human oversight alone can reliably catch errors as schools use artificial intelligence more widely.

The randomised experiment examined how 1,300 teachers responded to deliberately incorrect grades. Teachers reviewed identical student work accompanied by a score labelled as either AI-generated or assigned by another teacher.

The study published in PNAS Nexus found the grading fairness gap was 22% larger when a harsh score carried an AI label. When the incorrect score was too lenient, researchers found no statistically significant difference between AI and human-labelled grades.

Researchers Dr Sofoklis Goulas (Yale University), Professor Rigissa Megalokonomou (Monash Business School) and Dr Panagiotis Sotirakopoulos (Curtin University) conducted the study.  

Human oversight may not be enough

The researchers said the findings challenged assumptions that keeping a person involved in AI-assisted decisions would necessarily prevent mistakes from passing unchecked.

“Teachers are deferring to AI outputs without sufficiently questioning them, and errors are quietly going uncorrected. Multiply that across thousands of students and thousands of classrooms, and you start to see how quickly this becomes a serious problem,” Professor Megalokonomou said.

The issue has become increasingly relevant as schools expand their use of AI. The technology is now being applied to lesson planning, feedback, student support and administrative work.

In the study, 48% of teachers reported using AI tools at least weekly for lesson preparation. However, only 16% said they actively encouraged other teachers to adopt the technology.

Australian schools are confronting similar questions about how AI should fit into teaching and professional judgement. The Educator has previously examined how Australian teachers view AI and its impact on learning and how schools are approaching its use.

More recently, educators and school leaders have been focusing less on whether to use AI and more on how. That shift has placed greater emphasis on policies, staff capability and appropriate oversight. AI policy and practice in Australian schools have consequently become a growing leadership issue.

Tech-confident teachers showed greater deference

The study also found differences in how particular groups of teachers responded to harsh AI-generated grades. Younger teachers and those with postgraduate qualifications were among those less likely to challenge the AI-generated score.

Teachers who described themselves as technologically confident were also less likely to intervene. The result suggests familiarity with technology does not necessarily translate into greater scrutiny of its output.

“The very people we might expect to be the most capable and critical users of AI tools turned out to be the most likely to defer to a harsh AI grade. That’s counterintuitive and concerning because the people pushing AI integration forward in schools may be the least likely to catch its errors,” Professor Megalokonomou said.

The findings add another consideration for schools determining where AI should sit within existing decision-making processes. Previous discussion around navigating education in the AI age has also focused on preparing teachers and students to use the technology responsibly.

Australia’s Framework for Generative Artificial Intelligence in Schools identifies accountability, transparency and fairness among its guiding principles. It also recognises potential risks from errors and algorithmic bias, while emphasising appropriate human responsibility for AI-assisted decisions.

Researchers turn to AI training

The research team is now developing a training program aimed at helping teachers evaluate AI outputs more critically. The work will examine whether more targeted guidance can reduce the tendency to accept incorrect recommendations.

“It’s not enough to just tell people AI can be wrong. You need to show them specifically how and when their judgment is likely to go astray and build the habits to push back on AI,” Professor Megalokonomou said.

Researchers are also planning further studies into how AI interacts with existing teacher biases in grading. Other projects will examine whether AI training affects teachers’ daily productivity and whether low-cost interventions can change their use of the technology.

“I hope this research reaches the people who are making decisions right now about AI in schools, including policymakers, school leaders and education departments,” Professor Megalokonomou said.