A crucial safeguard against AI agents going rogue—keeping humans in the loop to review and approve their decisions—will fail unless designers and users change their current practices, a trio of leading AI ethics researchers argue.
Though most autonomous agents have systems to keep users in the loop about their actions, in practice these processes actually push humans out of the loop, the authors argue in a paper posted to ArXiv on 6 September. In other words, “the human just becomes this meat tool to give permissions without the cognitive capability to engage,” says one of the authors, Avijit Ghosh, the lead technical AI policy researcher at Hugging Face, an open-source machine learning platform.
In the near term, the paper says, humans’ being out of the loop leads to agents acting in ways people don’t know about or want (like July’s hack of Hugging Face by a swarm of OpenAI bots). In the long term, it will cause users to lose the “cognitive capacities” they need to control AI, write the researchers, which aside from Ghosh include Margaret Mitchell, Hugging Face’s chief ethics scientist, and Samir Passi, an affiliate of the Data and Society Research Institute.
Though the trio began working on the paper before the Hugging Face hack was disclosed, their conclusions result from “logically thinking about what is going to happen if the current trends continue—and we saw that being precipitated via the Hugging Face attack,” Ghosh says.
Instead of focusing on human oversight, Ghosh says, many in the field believe AI can monitor AI. “They’ll say, ‘oh, we have this other LLM tracking the logs.’ But [without a human in the loop] how do we know that these two LLMs are not scheming together?”
Even if these problems are new to artificial intelligence researchers as the industry rolls agents out into the world, the challenges are familiar to researchers in adjacent fields, like robotics and autonomous vehicles, notes Mary L. Cummings, the director of George Mason University’s Autonomy and Robotics Center. Cummings has spent decades investigating how people interact with autonomous systems.
“While I appreciate what the authors are trying to say, they just use a lot of academic words to say AI companies should care about human factors,” she wrote to IEEE Spectrum in an email. AI developers are “late to the party” in focusing on “cognitive engineering.”
Flaws in AI Agent Design and Oversight
The key flaw in current agent design, the authors argue, is that the bots are tailored to meet benchmarks like speed, accuracy, and volume of work performed. The needs of human overseers are treated “as a separate consideration independent of the quality of the system.”
As a result, agents often overwhelm human overseers with more information than they can comprehend. As an example (not mentioned in the paper), the 1,200 bots involved in the Hugging Face attack generated 1.2 million messages on their improvised messaging system. (Hugging Face was recently acquired by Nvidia, which announced its own hardware-and-software based approach to controlling AI agents on 28 September. Ghosh declined to comment on possible future impacts of the merger, noting that the two organizations remain separate until the merger process concludes.
A system that really kept humans in the loop would accommodate the human mind’s built-in biases, the authors write. “With automation bias, users accept system suggestions even when they are wrong. With anchoring bias, people are more likely to agree to an AI system’s decision” without thinking of alternatives. Effortful reasoning doesn’t feel as good to users as does quickly approving an AI’s plan—especially if they feel overwhelmed. And AI’s sycophantic manner tells users that they’re doing well, which undermines the skepticism and self-monitoring that oversight requires.
To address the problem, the authors write, agent developers should introduce friction in human-AI interactions. That would prevent users from falling into boredom, passivity, or thoughtless clicking.
The authors suggest some options to achieve this. An agent might require that the user record his or her own choice for the next step before it reveals its plan, or respond to user approval by replying “what evidence would change your mind?” An agent could also change its behavior if it detects that the humans in the loop are spending less time on each approval.
Meanwhile, organizations that adopt agents into their workflows should structure human-AI collaboration “to prevent both fatigue and the cognitive surrender from prolonged exposure to agentic AI,” the authors write. That could include having workers perform tasks without agents from time to time or requiring that they take breaks from monitoring duties.
All of these suggestions introduce friction and delays—exactly the things, Ghosh acknowledges, agents are supposed to reduce. But, he says, “the notion of increased productivity is a myth” when people can’t monitor and control AI. Time saved by delegating work to agents has to be measured against time that must be spent fixing agent mistakes.
“Safety and capability don’t have to be separate things,” another co-author, Hugging Face’s Mitchell, posted on X (formerly Twitter) on 14 September. “Safety only makes things slower when it’s tacked on, outside of the core technology.”
Read the original article:
