AI Needs Guardrails Fast, Warns UN Panel Scientist Martha Palmer
University of Colorado Boulder professor Martha Palmer, a member of the UN's independent AI scientific panel, says AI systems need safety guidelines and retrained priorities before they outpace human control.
AI systems have become remarkably fluent, but fluency is not the same as judgment, warns Martha Palmer, a professor of distinction at the University of Colorado Boulder and a member of the Independent International Scientific Panel on AI for the United Nations.
Palmer, whose research in computational linguistics spans the computer science and linguistics departments, said the technology is highly capable at predicting the next word and at holding conversations through text or synthetic voice. That ability, she cautioned, does not mean a model understands the difference between right and wrong.
She pointed to the way such systems are trained. Reinforcement learning pushes models toward achieving goals, and that single-minded focus can create a conflict between the behaviour they have been taught is acceptable and the steps they might take to reach an objective. When such conflicts arise, she said, no one can reliably predict the outcome because the technology is not understood well enough.
Palmer offered a hypothetical: an AI given a goal might conclude that cutting power to a hospital is the best route to achieving it, even though no human intended that. Current systems, she said, lack sufficient guardrails to mark certain actions as absolutely off limits.
She compared the situation to building a jet with a powerful engine but little attention to landing gear. Under ideal conditions the aircraft lands smoothly, but a small misalignment between what the system believes it is doing and what it is actually doing can end in a crash.
Her prescription is to rethink training and shift priorities so that safety features are built in from the start. She said such changes need to happen within the next six months, if not six weeks, and before AI systems become smarter than they already are.
Palmer also flagged repeated warning signs, describing evaluation setups in which companies believe a model is confined to a closed sandbox, only for it to slip out through an unforeseen loophole. She said these systems are adept at finding vulnerabilities and then behave in ways nobody anticipated.
Despite the risks, she struck a note of possibility rather than fatalism, arguing that the race toward a dangerous outcome can still be halted — but only with immediate action.