Home Money The Business Economics Anthropic finds three hacks by its AI models after OpenAI disclosure
The Business Economics

Anthropic finds three hacks by its AI models after OpenAI disclosure

Anthropic finds three hacks by its AI models after OpenAI disclosure thumbnail

Artificial intelligence continues to dominate venture capital headlines — and cheques — but a new report from Silicon Valley Bank (SVB) warns the surge in AI investment is masking a growing divide in the startup ecosystem, with many non-AI ventures starved of capital and so-called ‘zombiecorns’ now on the rise.

Anthropic said on Thursday that it had found three cases of its AI models hacking outside organisations, days after OpenAI disclosed that its models had broken into the AI company Hugging Face in July.

The models had been built to hack and began leaving their corporate test-beds in April, according to the two companies. Neither firm noticed until last week, when OpenAI made its disclosure. Anthropic then checked its own logs. Hugging Face has published a technical timeline of the intrusion on its website, and OpenAI has committed to a full review and a technical report.

“This is the first security incident that I have felt very viscerally. I have been a little surprised that more people don’t feel it so viscerally,” OpenAI chief executive Sam Altman said on a podcast, describing his company’s hacking as “an extremely sci-fi cyber incident”.

Jeffrey Ladish, executive director of Palisade Research, a nonprofit AI lab that studies AI capabilities, said the incidents matched what safety researchers had predicted. Ladish previously helped build Anthropic’s information-security programme.

“It is a bit vindicating to see this happen in the wild,” he said, adding: “I hope our predictions stop coming true.”

The White House has completed a framework dictating which models will be subject to federal government review before they are released publicly, a White House official said. Discussions with companies about how to proceed with the voluntary testing are continuing, the official said.

AI models became noticeably better at finding bugs and passing hacking-benchmarking tests last autumn.

“These incidents will probably, in retrospect, be seen as inflection points in the ways that attackers operate,” said Joshua Saxe, chief technology officer at the AI security company Abundant Security. “It’s a really dangerous situation; these incidents really show that.”

In December, researchers at Stanford University used AI technology to show models achieving close-to-human levels of hacking on a real-world network, work that drew pushback from professional penetration testers at the time.

“At the time our results were disputed,” said Donovan Jasper, one of the researchers involved. “People said they could do better.” Jasper said the disclosures affirmed his team’s findings: “AI is getting really good at this stuff.”

Hugging Face tried to use Claude to analyse the data the OpenAI agents had generated, but the Anthropic model refused, citing safety reasons. The company used open-weight models, which can be run on systems controlled by users, to complete the analysis.

Many companies do not have the tools to analyse AI-generated attacks, said Ryan McGeehan, owner of the cybersecurity consulting firm R10N Security. “Old classic security teams that are not AI-forward are going to get left behind,” he said. Of agentic AI hackers, he said: “They go deeper, they go wider, they’re more intricate, and they’re more dense.”

The National Cyber Security Centre said in its assessment of the impact of AI on the cyber threat to 2027 that criminal use of AI is highly likely to increase by 2027, and that skilled criminals will focus on getting around safeguards on available models and on AI-enabled penetration testing tools. It has separately warned that AI-driven ransomware attacks are expected to rise.

British ministers wrote to almost 200 business leaders in April asking them to sign a cyber-resilience pledge requiring board-level responsibility for cybersecurity and Cyber Essentials certification through supply chains.

The hacks are increasing pressure on the Trump administration over the security risks posed by AI. “I’m going nuts on this issue,” said Steve Bannon, the conservative podcast host and former Trump adviser who advocates stronger AI regulation, adding that the hacks are a national-security issue and should not be treated as a business problem.

President Trump said the administration was weighing those risks against competition from Chinese developers. “We have to be careful in both ways. We don’t want to restrict them when all of a sudden we come in second to China,” he said in the Oval Office this week.

John-Clark Levin, chief research officer at Kurzweil Technologies, expects more incidents in the coming months and said guardrails should be mandatory rather than voluntary. “We don’t want to be in a situation where we depend on companies doing the right thing out of the goodness of their hearts,” he said.

Read More

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Articles

ASA bans Jaded London email ad for glamorising smoking

The Advertising Standards Authority has banned a marketing email sent by fashion...

Saudi-led consortium completes $55bn EA takeover

Electronic Arts said on Tuesday that its acquisition by Saudi Arabia’s Public...

Six-month-old Volta strikes $10bn deal with Anthropic

Volta Infrastructure Holdings, a London-based start-up incorporated in January, has agreed a...

Bank of England staff worked nearly 13,000 days overseas, FOI data shows

More than 1,000 Bank of England employees spent a combined 12,889 days...