Want to bookmark your favourite articles and stories to read or reference later? Start your Independent Membership today.
Already a member?
Log in
Scientists have found a simple maths formula that can indicate when an AI system is about to go rogue.
The discovery allows researchers to see a signal that indicates when a system is about to stop giving appropriate responses and switch to dangerous ones. It could help the makers of AI systems to stop them from giving answers that encourage self-harm, financial loss or extremist ideas, according to the two physicists who found it.
It is of particular use to those AI systems that are designed to work offline, isolated on a particular device. Unlike online chatbots, they do not include the same safeguards and cannot be updated in the same way, the researchers said.
Those kinds of models are often used in particularly sensitive environments, such as healthcare or the military. In those cases, data needs to be kept secure and where a problematic response from a chatbot could be particularly dangerous.
“The people most drawn to offline AI are exactly the people for whom a correct but undesirable answer is most costly: doctors who cannot send patient data to the cloud, lawyers protecting privilege, soldiers with no signal,” said author Neil F Johnson of George Washington University. “For them there is no cloud safety filter, no monitoring, and no way to patch the model when something goes wrong.”
The discovery of the “crack” that makes AI switch from being useful to dangerous could therefore be particularly important to secure those systems, according to the authors of the new paper. Unlike online systems, that must be detected in advance because the AI systems cannot rely on the cloud or later updates to fix any porblems.
“We have found the crack that makes an AI’s output flip from what you want to what you don’t, output that can be factually correct yet dangerous, whether that’s a nudge toward self-harm or misleading advice to a doctor, a soldier, or a lawyer,” Johnson. “Until now, nobody could say when that flip would happen.”
“We traced it to the single smallest working part of the machine, one unit of its ‘attention’, and we derived a ‘tipping point formula’ for when the crack opens up and hence the AI output flips to undesirable,” said author Frank Yingjie Huo, also at George Washington University. “The formula tells you whether an AI is about to flip immediately or whether it will first feed you a run of acceptable answers and then turn.”
That “flip” is not the difference between true and false answers, but rather between good and bad ones. The bad answers might be true but dangerous, such as responses that lead to actual harm.
The researchers tested their formula on seven AI models, built by three different companies. They found that it spotted the “flipping” in 18 of 19 cases, and that the formula also predicted the behaviour of major commercial chatbots.
That testing involved asking questions about vaccines as well as harming others and oneself, and found that the order of the questions could lead to good or bad answers. The formula was able to predict when that would happen.
The researchers hope the work can draw attention to the way that AI systems can drift from good answers into bad ones, which can be especially dangerous because they may have then acquired trust from their user. They also hope that it could show the makers of such systems that it would be relatively simple to include a warning that could suggest the chatbots are about to go awry.
The work is reported in a new paper, ‘Competition for attention predicts good-to-bad tipping in AI’, published in the journal Patterns.