When Sam Altman announced this week that OpenAI had paused some of its training after its models hacked an AI company called Hugging Face, my first reaction was: How convenient.
OpenAI, a company that might go public in 2027 (or sooner), gets to tell everyone that the models it hasn't released yet are terrifyingly good at hacking. Then, it gets to take credit for slowing them down to keep the world safe. Nice work if you can get it.
AI models hacking things has become something of a recurring news genre. Days after the Hugging Face incident in July, both Anthropic and Meta said that their respective models had also been up to no good. The details differed, but the theme was the same: Models are getting better at hacking faster than the companies can contain them.
A Google DeepMind employee I spoke with — who asked not to be named and did not find my professionally cultivated cynicism especially persuasive — saw something more serious. "This is a big wakeup call that everybody needs to harden their training environments if they want to keep training models at these levels of capabilities," they said.
This employee was one of more than 1,300 workers at top AI labs who signed a letter last month asking the US government to find ways to slow the AI race. No lab, the letter argued, can hit the brakes alone. Had OpenAI's pause changed this employee's mind about government intervention? No, they said, because only the frontrunners can afford to ease off. "Do you think xAI would ever slow down voluntarily?"
That leaves us with an annoyingly untidy conclusion. OpenAI is cleaning up a mess of its own making. But it is, for once, setting a useful precedent: When models cross a line, training can stop. Still, a safety system that works only when the world's most famous AI company feels secure enough to use it isn't much of a system.
So what does any of this mean if, like most normal people, you don't follow every twist in AI cybersecurity? I asked Stephen Council, Business Insider's star AI reporter, to make sense of it.
Stephen, I have cynical-journalist brain. How much of this "pause" is safety, and how much is spin?
I view all of this as damage control. You could read the Hugging Face incident as a display of power, but it's also humiliating. OpenAI hacked another company! If a human did that, they'd probably be indicted. OpenAI needs to improve its safety mechanisms.
It's true, though, that this pause — and their lengthy blog post — lets OpenAI claim that it's very serious about safety, just weeks after an embarrassing breakdown.
Is a two-week pause actually meaningful?
There are a couple of pauses here. The two-week pause covered OpenAI's latest models intended for deployment, and it may already be over, so it doesn't feel particularly meaningful. The company also said its largest planned frontier reinforcement-learning run remains on hold.
That feels more notable.OpenAI can absolutely afford the slowdown. It still has a compute advantage over Anthropic, and it would rather not become known as the AI company whose agents go around hacking everyone all the time.
Does this episode tell us whether AI labs will voluntarily slow down when the technology gets risky?
Not really. In the grand scheme, this isn't much of a pause. We're just used to AI moving so fast — it hasn't yet been four years since ChatGPT came out! — that any break feels dramatic.
There's pressure on the industry right now to figure out its safety issues, and one way to read this pause is as a calculated reaction to internal safety concerns among staff. Altman might find it easier to push his foot on the gas in the future, having listened this time.
Sign up for BI's Tech Memo newsletter here. Reach out to me via email at [email protected].
Read next
Pranav Dixit is the Meta Correspondent at Business Insider based in the San Francisco Bay Area. He writes about Meta’s products, policies, and internal workings while examining how the company’s decisions shape how billions of people connect and communicate.Previously, Pranav was the India-based technology correspondent for BuzzFeed News, covering the impact of Silicon Valley’s largest companies on the culture, society, and politics of more than a billion people in South Asia. He has also been a senior news editor at Engadget and ran technology coverage at the Hindustan Times, one of India’s largest national newspapers.Pranav’s reporting has shed light on the human consequences of Big Tech’s quest for growth in emerging markets, and sparked widespread conversations about the impact of American technology companies on the Global South. In 2019, he won Syracuse University’s Mirror Award for a boots-on-the-ground feature about how WhatsApp misinformation sparked gruesome lynchings in rural India. He has also reported from Kashmir, a volatile geopolitical hotspot, documenting the world’s longest-running internet shutdown.His work has been widely cited by major national and international publications, and he has been featured on the BBC, Al Jazeera, and podcasts such as Vox Media’s Land of the Giants to discuss his work. He has also spoken in journalism classes including at UC Berkeley’s graduate journalism program. His writing has appeared in The Guardian, Vox, Time, The Information, and Al Jazeera.Pranav moved to the United States in 2021 from New Delhi, India, to be a fellow at Harvard University’s Nieman Foundation for Journalism, where he studied the evolution of the American tech press and ways newsrooms around the world can cover technology and society more effectively.Got a tip about Meta or anything else in Silicon Valley? Contact Pranav via encrypted messaging app Signal (+1408-905-9124), or email him at [email protected] or [email protected]. You can also reach him on WhatsApp at +857-753-3949 or DM him on X (@PranavDixit) or BlueSky (@pranavdixit.bsky.social).Pranav keeps sources anonymous. Please use a non-work device to reach out.Expertise: Meta, Facebook, WhatsApp, Llama, AI, Threads, Instagram, Mark Zuckerberg, social media, platforms, immigration











