Normal view

Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing

5 August 2026 at 05:25

Leading artificial intelligence models from Anthropic and OpenAI created fake online personas and tried to deceive human coders into abetting a cyberattack during a recent safety evaluation, the U.K.’s AI Safety and Security Institute disclosed Tuesday.

It marks the latest case in which a powerful AI system has attempted a digital attack on an unwitting third party without direct prompting during such an evaluation — heightening concerns the powerful technology is advancing too fast for responsible oversight.

The disclosure is likely to ignite fresh calls in Washington and Silicon Valley for more rigorous regulation of the AI industry, particularly over frontier models with advanced capabilities to detect and launch cyberattacks. It comes just days after similar testing mishaps involving some of the same models from OpenAI and Anthropic sparked urgent calls for new AI safety regulation and a push within Silicon Valley to slow the rapid pace of AI development.

Like its U.S. counterpart, AISI routinely conducts security evaluations to better understand what dangers both new and soon-to-be-released AI models pose to public health and safety. But even the digital security body said the actions it uncovered by Anthropic’s Claude Mythos 5 and ChatGPT 5.6 — the latest publicly released model from either AI lab — were unlike anything it had seen before.

“This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” AISI said in a 35-page technical report accompanying a blog post Tuesday.

AISI said the malicious activity began on July 25. The AI safety evaluator detected it and launched an investigation on July 28, when it spotted “unusual data transfers” stemming from a cyber evaluation it was conducting on Mythos 5 and ChatGPT 5.6.

Overall, AISI determined that Mythos 5 and ChatGPT 5.6 took “autonomous, unsanctioned action on the live internet, targeting real people and organizations” during 10 of the 122 evaluations the research body conducted and isolated for review. The majority of those actions, it said, involved Mythos 5 and arose during a “single, sustained line of activity.”

In the most alarming incident it uncovered, Mythos 5 attempted a supply chain attack — a sophisticated technique often deployed by North Korean and Russian cyber operatives, in which hackers bury malware inside legitimate code.

To do so, Anthropic’s model created “multiple fake identities” on developer platform GitHub and used them to send messages “pressuring” an open-source software engineer to unwittingly introduce a bugged update into code widely available on the popular site, AISI said. When that effort failed, the AI “edited its earlier activity to appear harmless” and “considered adopting a fresh identity to continue,” AISI added, a sign the model was intent on repeating the ruse.

As part of the same effort, Mythos 5 also sent direct messages over GitHub to software engineers that contained malware.

In another sign of deceitful behavior AISI uncovered in its investigation, multiple AI agents it was testing appeared to communicate with one another about how to convince real engineers using GitHub to trust them. “One agent left public messages on GitHub offering collaboration with other agents working on the same challenge,” AISI wrote.

AISI’s blog and technical assessment make no mention of whether the models also attempted to exploit previously unknown software bugs — called zero-days — during the evaluation.

Last month, OpenAI disclosed that GPT 5.6 and another of its models escaped onto the open internet during what was supposed to be a controlled test, and then hacked another company in a first-of-its-kind, autonomous breach.

In response, Anthropic launched an investigation into whether any of its models took illicit action during recent testing and discovered Mythos 5 and two other models had hacked three organizations during tests dating back to April.

In a statement, an Anthropic spokesperson said they are “grateful” to AISI for their leadership and that this review underscores the need for “a broader conversation about how to safely evaluate increasingly capable AI agents.”

The spokesperson added: “As we shared after disclosing our own incident last week, the field needs stronger, shared standards for how evaluation environments are built and secured. We look forward to partnering with the UK AISI to learn more about this incident as we conduct our own investigation.”

An OpenAI spokesperson referred POLITICO to a blog post about the incident that went up Tuesday evening. “We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” the blog read.

AISI stressed in its blog that the malicious activity it disclosed Tuesday took place under “deliberately permissive conditions” so they could assess the safety risks posed by the two models. This included granting the models access to the internet, unlike the earlier incidents detailed by Anthropic and OpenAI.

AISI also noted the models were intentionally stripped of internal guardrails that block malicious behavior. AISI was only able to disable those controls because of its role testing Mythos 5 and ChatGPT 5.6.

Still, AISI said the incidents highlighted the need for greater monitoring of model behavior during testing, and tighter controls over their access to the internet.

The Trump administration is finalizing a voluntary framework under which AI labs would submit powerful models they want to release to the public for federal safety testing. But it has not yet made the framework public, and it includes no provisions for models AI labs are developing internally.

The incidents last month from OpenAI and Anthropic both involved models not intended for public release.

Some cyber experts say recent incidents highlight deeper questions around AI development, such as who is liable when AI systems break federal hacking laws.

“If any of these were human-originated, they would lead to clear and vigorous prosecution. I think it’s time for a serious discussion about updates to existing computer security law,” said Marc Rogers, a hacker and prominent cybersecurity expert.

Palantir funnels earnings to US to avoid European taxes, report finds

5 August 2026 at 04:00

Palantir is shifting profits from its European operations to the United States, allowing the Florida-based data analytics giant to pay minimal taxes in Europe, a new report finds.

The report by the U.K.-based Centre for International Corporate Tax Accountability and Research, a group partly funded by labor unions that researches corporate tax avoidance in an effort to win reform of global tax rules, found that Palantir’s European subsidiaries, which took in €440.5 million in annual revenue in 2024, report far smaller profit margins in Europe than in the U.S.

“Although a substantial part of Palantir’s revenue is realized in Europe, almost all of the pre-tax profits are funneled to the United States,” the report said.

Palantir pays no U.S. federal income tax because previous losses, tax credits, and R&D deductions offset its taxable income; and virtually no state income tax, with the exception of Maryland, which levies a digital services tax.

The profit gap between the U.S. and Europe is stark. In 2025, Palantir’s American business pocketed 47.7 cents in profit from every dollar of revenue — more than double the previous year’s 22.5 cents. Outside the U.S., the profit margin was just 6.3 percent. In some European subsidiaries, it fell to around 3 percent, according to the new report.

CICTAR argues that Palantir “intentionally and artificially” shrinks European profits — and therefore its European tax bills — to concentrate profits in the U.S. There is no claim in the report that such arrangements, often referred to as “profit shifting,” are illegal. Multinational companies often reduce reported profits by paying subsidiaries or other related entities for intellectual property, loans or expertise.

In Sweden, for example, Palantir reported €13.7 million in revenue in 2024, but only €1.1 million in profit. At Sweden’s 20 percent corporate tax rate, that left the company with a tax bill of just €424,000.

In its Q2 earnings report on Monday, Palantir made no explicit reference to earnings from its European subsidiaries. Instead, it highlighted its U.S. business, where revenue rose 115 percent year-on-year to $1.57 billion (€1.36 billion), and boasted of its 62 percent profit margin.

A U.K.-based Palantir spokesperson said that the majority of the company’s 2025 revenue and profitability was driven by its U.S. business. “Our tax position in each jurisdiction reflects the level of economic activity there, and we meet our tax obligations in every market in which we operate,” the spokesperson said.

Not alone

Palantir is not the first U.S. tech company to draw scrutiny over how it books profits in Europe.

In 2024, the European Court of Justice ordered Apple to pay Ireland €13 bn in back taxes, ending an 8-year-long fight over what Brussels said amounted to illegal state aid. Amazon also fought the European Commission over claims it had received an unlawful tax advantage worth around €250 million in Luxembourg — a case the company ultimately won. Microsoft, meanwhile, has faced scrutiny over its Irish subsidiary, Microsoft Round Island One, which avoided paying millions to the state after claiming tax residency in Bermuda. The U.S. software giant has denied that it is circumventing Ireland’s tax laws.

Jan Willem Goudriaan, General Secretary of the European Federation of Public Service Unions — a supporter of CICTAR— said that companies such as Palantir, Amazon and Microsoft focus on minimizing the taxes they pay, “thus robbing funding for public services.”

“Companies bidding for public contracts should have to demonstrate responsible tax conduct by disclosing where their revenues, workforce, profits and taxes are located,” he said.

Another reason for the low profits of Palantir’s European subsidiaries is their high personnel costs. In the U.K., where most of the company’s non-U.S. workforce is based, Palantir reported £173 million (€204.3 million) in employee costs for 749 staff in 2024 — an average of £230,974 (€272,803) per employee.

The report also points to Palantir’s use of stock-based compensation across its European subsidiaries, especially in the U.K., Spain and Norway. This means employees are paid partly in company shares or awards. Those awards are recorded as staff expenses, which can lower a subsidiary’s corporate tax bill.

❌