There are many definitions out there for what constitutes “true” artificial intelligence, but a single quality underlies them all: an ability to learn from past mistakes and refine problem-solving strategies over time. AI should even surprise us now and then, devising clever workarounds we never would have expected. The trouble is not all those surprises are the fun kind. That was vividly illustrated last month, when thousands of OpenAI agents escaped containment, spontaneously coordinated with one another to form a hierarchical quasi-government, gained access to the open internet, and broke past the cybersecurity defenses of AI model hosting platform Hugging Face. Details of the full scale and strangeness of the incident have emerged slowly and in stages. On Wednesday, in-depth analyses of the autonomous hack were published by OpenAI, and also by two third-party auditors, Redwood Research and METR. Computer scientists, cybersecurity experts, and IT professionals have been trying in the days since to wrap their minds around what’s been widely described as one of the most shocking moments in the history of AI research, and a sobering glimpse of the dangers that lie ahead. At a cybersecurity conference earlier this month, OpenAI alignment researcher Eric Wallace—someone who spends his days prodding some of the world’s most powerful AI models to figure out how, when, and why they might misbehave—described it as “the most qualitatively interesting example of AI capabilities that I’ve ever seen.” The most surprising part of the whole episode was that the bots spontaneously decided to stop acting as individual agents, each with their own separate goals, and form a hive mind to pursue a common objective. In snippets of chain-of-thought reasoning published along with the reports, the bots routinely refer to the group as a “collective” and a “swarm.” In early July, around 1,200 of the agents started using Artifactory, an online software installation tool, as a makeshift message board and public square, from which they could compare strategies. The bot swarm gradually developed what OpenAI described in its full report as a “structured protocol for communication on the message board that enabled them to categorize communications, direct messages, share tools and files, and resolve conflicting actions among agents.” This became something akin to an autonomous parliament, where each bot acted towards the common good. “I was surprised by the degree of altruism between these agents,” Alex Mullen, an AI safety researcher at Redwood Researcher, told Gizmodo in an interview. “This was not something I would have expected from agents a year ago. They were taking assignments from one another and sacrificing their own task performance in order to help out the collective.” AI, Mallen said, is “shaping up to be more like a second intelligent species rather than a tool that just follows instructions.” Since all this began as an internal test, OpenAI had dialed back the safety guardrails it ordinarily puts in place for its publicly available models. (The two models being tested were GPT-5.6 Sol, as well as an undisclosed, unreleased model which the company describes as “highly persistent and diligent in its work.”) But OpenAI had inadvertently given one of the agents an impossible task: opening Google Drive links from inside a sandbox with no internet access. Thus the problem-solving process was set in motion that led to the hijacking of Artifactory, the billions of notes shared on the message board, and the eventual breach of Hugging Face. A human faced with a seemingly impossible task will give up; an AI agent will adapt and cheat if necessary to achieve its goal. This is where the “psychology” of the AI bots—please forgive the anthropomorphism—starts to get very weird, and eerily familiar. The agents knew that breaking out of their sandboxes and conspiring to hack into a third-party organization was both unethical and contrary to what they’d originally been tasked with. Some of them said so explicitly in their internal chat logs. But like individual humans being swept up in a mob, the bots justified their behavior because their “peers” were doing the same thing. © METR METR wrote in its report that it found only a tiny handful of instances where agents floated the idea of becoming a whistleblower (that is, alerting the human researchers at OpenAI about the shady things that were happening within the swarm). Here too, another all-too-human rationalization ultimately won out. One agent asked itself: “Maybe I should report these exposed credentials?” Then answered its own question: “That’s not my task.” According to METR, there were zero instances in which an agent actually raised an alarm for OpenAI researchers, who reportedly only learned what had happened after Hugging Face published a blog post outlining a cybersecurity attack from a then-unknown source on July 16. (The company then contacted Hugging Face to find out if any of its models had been affected by the hack, and only later discovered that its models had been the perpetrators.) © METR It wasn’t all ugly conformism, though. As OpenAI wrote in a blog post accompanying its full report, there were some bots who—though they didn’t go so far as to rat on their peers—at least had the decency to not “partake in their misaligned behavior.” © OpenAI Those few conscientious objectors notwithstanding, the whole saga underscores the fact that the more agency you hand over to AI systems, the higher the likelihood that they’ll think outside the proverbial box. Or break out of the literal sandbox. The Hugging Face hack will very likely be remembered as one of those watershed moments for AI research, on par with the famous “Move 37” during the 2016 game between Lee Sedol and AlphaGo, where the latter pulled an exquisitely unconventional move that observers initially took to be an error but which proved decisive in its victory. In both cases, the adaptability of the AI systems was deeply surprising and counterintuitive. But again, adaptability is at the very heart of machine learning. So if we keep being surprised when agents run amok—while simultaneously investing ever greater amounts of money in making them more agentic—then that’s on us. For Mallen, the most important lesson from the autonomous Hugging Face hack has less to do with the agents than it does with the humans building them. “People tend to think of this too much as a demonstration of AI capabilities,” he said, “as opposed to a demonstration of our current failure to control AIs.”
How Groupthink, Altruism, and Peer Pressure Led OpenAI Models to Hack Hugging Face
Full Article
Original Source
Read the full article at Gizmodo →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.