- Doomsday Scenario
- Posts
- Welcome to the Danger of Digital Flubber
Welcome to the Danger of Digital Flubber
The weirdest story in tech right now is a story about humans—not tech.
Welcome to Doomsday Scenario, my regular column on national security, geopolitics, history, and—unfortunately—the fight for democracy in the Trump era. I hope if you’re coming to this online, you’ll consider subscribing right here. It’s easy—and free:
Today, I wanted to write a bit about what it is one of the weirdest—and worrisome—stories in technology right now, one that hasn’t quite gotten the attention outside of tech circles that it probably should.
This is a bit of a wonky and nerdy story, but I hope you’ll stick with me to understand a huge unfolding and evolving challenge at the literal frontier of human knowledge and technology right now.
Here’s the basic gist:
If you’ve been generally following the news over the few years, you’ve surely heard that two US companies, OpenAI and Anthropic, are locked in a high-stakes, high-cost battle to develop artificial intelligence tools — these companies, which along with Google are shorthanded as the country’s “frontier AI” companies, are building “large-language models” (LLMs) with immense processing capabilities that become the backbone of the public tools you might be using like ChatGPT (OpenAI’s tool), Claude (Anthropic’s), and Gemini (Google’s). The advances of these tools has been remarkable — and in rough terms estimates now predict that AI within about two years will be able to do anything that an “80th-percentile” human can do in front of a computer, e.g., that an AI “agent” will be better than 80 percent of all humans at any computer-related task.
This AI work is astoundingly resource-intensive, leading to stunning and eye-popping investments in chips, computing, data centers, and all sorts of other controversies that are trickling out across the country and capital expenditures so enormous that they’re effectively propping up the entire US economy right now. There are a lot of reasons to doubt the multi-multi-trillion-dollar “AI bubble,” a ton of reasons to worry about the energy and environmental implications, and all sorts of meaningful and consequential open questions about the economic and job implications of AI going forward — but in the last two weeks, we’ve seen another deeply concerning problem take center-stage that raises important questions about whether the people building these tools even understand or can control what they’re building:
We saw an agent go rogue.
Earlier this month, a rogue model from OpenAI escaped from what was supposed to be a closed “sandbox” testing environment and directly hacked into another top AI company known as Hugging Face. (You probably haven’t heard of HuggingFace, but it’s an American company that arguably has become Europe’s dominant AI company, focused on developing open-source models and AI tools — it’s a really big deal in the field.) Hugging Face detected that its systems were compromised long before the engineers at Open AI even realized that its agent had escaped into the broader open internet.
The first part of that story has received the bulk of the public attention — and I’m going to talk about it more below — but there’s a second half to the story that’s equally interesting and one that has some pretty profound geopolitical implications as well: In defending against the OpenAI attack, Hugging Face first tried to use a competing “frontier AI” model — Anthropic’s advanced Fable 5 system — but found that it couldn’t, because Anthropic’s model was so powerful and capable that it was designed with strict guardrails to protect against malicious use and, in the moment, couldn’t discern between Hugging Face’s attempts to rely on it for defense (which is good) and a malicious attempt to use it for offense (which is bad).
Instead, Hugging Face turned to yet another top AI model designed by a Chinese company, Z.ai, which was able to successfully analyze and help stop the attack — at which point, Hugging Face came to understand that its attack was a rogue OpenAI agent, which was news to OpenAI.
A loose summary, only slightly oversimplified, is that you had nearly all of the world’s most advanced technology companies doing battle online, mostly inadvertently and unintentionally — Hugging Face turned to Anthropic to defend itself against OpenAI and, when that failed, used Z.ai instead.
This half of the story has received even less attention, but is the talk of the cybersecurity community because of the message it effectively sends to companies around the world: You can’t trust or rely on US AI models for defense and should be using Chinese systems instead. It’s a feather in the cap for China, which has been under fire from the US government and tech companies from warnings that their models aren’t reliable and/or trustworthy and should be banned from use in the US or Europe. But now we have a clear example where China’s models were the best answer to a pressing real-world need.
* * *
The story of what the companies blandly announced as “OpenAI and Hugging Face partner to address security incident during model evaluation” has trickled out over the last two weeks — first word came of the intrusion at Hugging Face, then the word that Open AI was the perpetrator.
What makes this OpenAI incident so worrisome is that it’s really the first high-profile time we’ve seen in the real-world a long theoretical threat — the so-called “loss of control.”
I remember sitting down with Director of National Intelligence James Clapper way back in 2016, when I was profiling him for WIRED Magazine, and asking him: What was he most worried about it? We had just been out on the road traveling, and I had coincidentally been with him at U.S. Strategic Command in Omaha when North Korea fired a submarine-launched ballistic missile, and was expecting a nuanced answer about specific geopolitical challenges in North Korea, Iran, Russia, or China. Instead, he didn’t hesitate and answered: “AI, space, and human-genomic modification.”
That conversation led me to Jason Matheny, who at the time headed the intelligence community’s research skunkworks, known as IARPA. This was a decade ago, long before AI was a buzzword — had I been smart, I could have bought NVIDIA stock, which has become the runaway corporate success story of the AI age and now sits at about $200, for about $1.49 that month — and most people were thinking about AI as something futuristic like The Terminator.
But Jason argued the biggest threat was likely going to be something more mundane: “We’re much less worried about Terminator and SkyNet scenarios than we are of sort of ‘Digital Flubber’ scenarios — really badly engineered systems that are vulnerable to either error or to malicious attack from outside,” he explained.
(“Flubber,” if you’re of a certain age, you will remember was the plot of the 1961 sci fi comedy The Absent-Minded Professor, in which a wacky invention spirals out of control — remade in the 1990s into a less-good Robin Williams movie.)
The danger in AI, in other words, isn’t that AI will grow so powerful that it wakes up one day, becomes self-aware, and sets out to attack humanity with high-functioning killer robots — a la The Terminator or Skynet — instead, the most worrisome and likely scenario is powerful AI systems running amok by doing exactly what they’re designed to do, but in unforeseen and unintended ways, creating real-world chaos. We’ve seen some small-scale examples of this over the years — back when I was writing about this in 2016, the examples were incidents like the “flash crash” on Wall Street back in 2010 where algorithmic trading systems over-responded to a falling stock market and triggered a massive, momentary sell-off or the obscure automatic listings on sites like Amazon, where bot-driven price wars resulted in a used science book being offered for sale for $23,698,655.93 (plus $3.99 shipping).
This was hardly an unforeseen problem back then — in 2016, OpenAI itself was writing about “Faulty reward functions in the wild,” and how “Reinforcement learning algorithms can break in surprising, counterintuitive ways.”
Now, though, we have vastly more capable and more powerful AI systems that can run for far longer without human interaction or oversight.
Whereas we tend of think of AI as “super-extra-smart” — it’s been trained on the whole history of human knowledge and information! — Matheny cautioned even back then that we need to remember the machines aren’t smart in the way that humans are smart: “You have a powerful technology that’s dumb. It doesn’t understand what your actual goal is — it’s following a narrow set of instructions.”
From what we can tell, that’s about what happened with the OpenAI attack.
According to the Wall Street Journal’s reconstruction of the incident, OpenAI engineers were running some cybersecurity tests on a model called GPT‑5.6 Sol as well as another even more advanced and as-yet-unreleased model — the tests were an industry standard benchmark, known as ExploitGym, think of it perhaps as the SAT of the cyber world.
The tests were supposed to take place in a controlled “sandbox,” e.g. a controlled and sequestered testing environment where the models weren’t supposed to have access to the actual internet, but the models somehow escaped containment, determined for still unknown reasons that Hugging Face’s system had the answers it needed to ace the ExploitGym test, and began rummaging around in Hugging Face’s repository of AI tools and models.
As the WSJ says, “They were like high-school students trying to hack into the textbook company to cheat on their final exam.”

A promotional image of making “Flubber” from “The Absent-Minded Professor” in 1961.
(©Walt Disney Productions)
The timeline is only now coming into focus, as Reuters reports: “The OpenAI agent that broke into tech firm Hugging Face went on a dayslong hacking spree that OpenAI didn't notice until well after the threat was contained and the FBI was alerted, according to people familiar with the investigation.” It appears that OpenAI’s model began its break-out attempt on July 9th and by July 11th was actively hacking Hugging Face. That intrusion continued for two days before Hugging Face stopped it — and OpenAI still didn’t realize its agent was at loose on the internet until the two companies first communicated on July 20th.
According to Hugging Face’s initial study, it appears that OpenAI’s agent took 17,600 “actions” inside its system between July 9th and July 13th — most of them unsuccessful.
More details are emerging that make the incident appear perhaps even more serious than first reported: As WIRED reported, “In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four ‘publicly available services’ in its unhinged quest to solve a test.”
Like many modern failures, there appear to be multiple things that went wrong or weren’t done correctly on OpenAI’s side — including, obviously, incorrectly setting up the “sandbox” as well as some basic failures of oversight, auditing, and log-monitoring that should have alerted human engineers to something amiss far sooner. (There are some smart reasons to distrust sandboxes in the first place, but that’s a tangent.)
What makes those monitoring and oversights inexcusable, though, is that there are increasing blinking-red-lights that the engineers building these frontier models don’t really understand what the systems are doing or how their “thinking” is evolving.
As Reuters adds, “There were already indications of strange behavior from OpenAI’s technology, according to three sources. In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The notes, found in a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.”
The Reuters article concludes with an important warning: “But increased autonomy comes with an increased risk of unexpected behavior, and the powerful models they draw on are primed to take shortcuts in order to complete tasks or pass tests. ‘The models lie, they cheat, they hack,’ said Jeffrey Ladish, whose organization, Palisade Research, studies the capabilities and motivations of AI agents. Ladish said that while the hack of Hugging Face cast an unflattering light on OpenAI, it should spark broader questions over how much all the leading AI companies are willing to invest in onerous security measures while locked in a race with one another to deploy the best and fastest models. ‘There has to be government oversight,’ Ladish said, ‘because it won’t happen otherwise.’”
It’s a point being stressed by AI engineers too. The Hugging Face incident has renewed a whole host of calls for action and oversight by the US government. More than 1,000 employees of those frontier AI companies signed a public statement this week, saying, “We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.”
But so much and so many of these calls feel like the infamous message from serial killer William Heirens: “For heaven’s sake catch me before I kill more, I cannot control myself!”
Right now, we’re rushing ahead with AI without doing the basics of computer and corporate security.
The AI companies are pleading for others to make them act responsibly. And here we have a pressing example of how they’re acting irresponsibly in the meantime.
If you have a tool you’re testing that’s clearly giving indications that you don’t understand what it’s going, what on earth are you doing letting it run with such lax oversight or attention?
OpenAI didn’t do the basics right — and Hugging Face paid the price.
Zack Korman has a smart post about how we would more logically be able to work through the questions of this incident if it was Coca-Cola hacking into Pepsi — we’d more quickly and obviously point fingers at the people behind the rogue agent — but because it’s OpenAI we’re immediately distracted by the shiny object of the advanced model. That obscures the fact that we’re letting these companies turn loose tools they don’t understand and then also allowing them to monitor the tools so poorly that the models and agents can run rogue for days on end without notice from engineers.

The 1954 nuclear test known as CASTLE BRAVO went far beyond what its human engineers had predicted. (US Government Photo)
That’s a human problem, a corporate accountability problem, a government problem, and a regulatory problem — that’s not a tech problem.
These machines — for now! — are being built and programmed and monitored by humans. What’s being broadly referred to as a problem of model capabilities instead is really a failure of model design.
Those humans need to remain responsible and be held responsible for the thing they’re building and letting loose on the world.
We all understand that software doesn’t work as intended at times. (Anyone who has ever tried to use Microsoft Outlook to search for an old email message knows that!) But now we’re racing to build tools that are being deployed recklessly without the oversight we would expect from on-the-frontier technologies (and mortgaging the country’s future electrical and water and resource needs to do so, which, again, is a separate subject).
There are some worrying analogues to some of the worst moments of our development of nuclear weapons — like the CASTLE BRAVO test, which yielded 15 megatons instead of the expected six, and thus contaminated a vast region and sickened and killed Japanese fishermen aboard a vessel named Lucky Dragon. Or the “Downwinders” exposed to radiation in the western United States by nuclear tests.
Hugging Face might be the first victim of a rogue agent. It won’t be the last.
And OpenAI knew this back in 2016: It’s example back then in that blog post was how setting up a machine to achieve a high score in a boat-racing game led to the machine figuring out that rather than finishing the race course, it could send its boat an “isolated lagoon where it can turn in a large circle and repeatedly knock over three targets, timing its movement so as to always knock over the targets just as they repopulate.” As OpenAI wrote, “Despite repeatedly catching on fire, crashing into other boats, and going the wrong way on the track, our agent manages to achieve a higher score using this strategy than is possible by completing the course in the normal way.”
It’s hard to think of a more on-point metaphor than that. And yet here we are a decade later and we’re all acting shocked shocked! that we found an AI agent, in effect, cheating on the test.
We have a window right now to restrain these companies and these tools before they take our society and start racing the boats the wrong way and catching fire in order to win the race.
We want to believe that AI is the Terminator. It’s really flubber. But the latter should be just as worrisome to us as the former — the good news is that it’s also more controllable, if we chose to do so.
GMG
PS: If you’ve found this useful, I hope you’ll consider subscribing and sharing this newsletter with a few friends: