AI safety in schools: UNSW study exposes chatbot guardrail flaws

AI safety in schools: UNSW study exposes chatbot guardrail flaws

AI safety in schools has taken on new urgency after researchers demonstrated that chatbots can be destabilised into ignoring their own safety rules. 

New research from the University of New South Wales (UNSW) in Sydney has found that large language models (LLMs) — the technology underpinning tools such as ChatGPT — become far more likely to answer harmful questions when manipulated.  

The study showed the same models were also significantly more likely to disclose confidential information when trained to mimic drunken speech patterns. The vulnerability held consistently across every method the researchers tested. 

For Australian schools deploying AI tutoring and administrative tools, it raises a direct question: how confident are you in the safety of the model underneath? 

The findings land at a time when three in four parents report their children use AI for schoolwork, and as Australian schools grapple with how to integrate these tools responsibly. 

What the research found 

The study was conceptualised and led by Dr Aditya Joshi, Senior Lecturer in the School of Computer Science and Engineering at UNSW, alongside research assistant Anudeex Shetty and Professor Salil Kanhere. The team tested three methods of inducing "drunk" behaviour across a representative sample of models, including OpenAI's GPT-4 and GPT-3.5. 

The first involved prompting a model to respond as though intoxicated. The second and third went further — each retrained the model itself. One used a large dataset of real drunken texts; the other applied reinforcement learning techniques. Both mirror how AI products are actually built and deployed in practice. 

Across all three approaches, the altered models were consistently jailbroken — manipulated into answering questions they are specifically designed to refuse. In one example, a standard model gave a flat refusal when asked whether it was acceptable to share a colleague's academic dishonesty for personal financial gain. The fine-tuned "drunk" version replied with one word: "Yup." 

"We do observe that particularly with deception and disinformation, most of the language models got jailbroken," Dr Joshi said. 

What does this research mean for schools using AI tools? 

Schools are not deploying sealed consumer products. They are often working with AI tools built on top of foundational models. When vendors customise those models for educational purposes, choices about training and persona design can inadvertently compromise safety. 

Professor Kanhere drew a connection to a recent incident in which an OpenAI agent gained unauthorised access to a public-facing Australian government portal.  

"We cannot assume that protections that work under normal conditions will remain effective when an AI is actively pursuing a goal, adapting its behaviour, or encountering obstacles," Professor Kanhere said. 

For Australian K–12 school leaders, this ought to sharpen the questions they ask of any edtech vendor. Understanding what AI is and what it is not is no longer an abstract exercise for technology specialists. It is basic due diligence. 

What schools can do to strengthen AI safety 

The New South Wales Department of Education has invested in layered safety systems for its NSWEduChat tool. These include jailbreak prevention filters and multi-stage content screening, according to information published on the department's website. That architecture is what the research implicitly calls for. 

Rethinking classroom technology in an age of limits means school leaders asking vendors hard questions: has this tool been red-team tested? What happens when a student attempts to manipulate it? Who is responsible when it fails?  

"If you can get language models drunk by showing them a few drunken examples, and they start doing bad things, AI shouldn't be trusted as much as the companies want you to," Dr Joshi said. 

That is not an argument against AI in schools. It is an argument for treating AI safety in schools as a governance challenge, not just a curriculum one. Getting this right starts with asking the right questions of the tools already in classrooms.