Reading view

Former PM joins list of French presidential candidates reportedly targeted by Russian interference

PARIS — Former Prime Minister Gabriel Attal is the latest in a string of French presidential candidates allegedly targeted by a Russian disinformation campaign.

Attal, the candidate of President Emmanuel Macron’s Renaissance party, said on Thursday he had been “a victim” of an electoral “interference” attempt originating from Russia, and that he expected such “destabilization efforts” to increase ahead of France’s April 2027 election.

Viginum, the country’s agency for fighting online disinformation, confirmed it had detected the operation against Attal and attributed it “with a high level of confidence” to “pro-Russian” disinformation tactics.

Former Prime Minister Édouard Philippe, who is also running for president, was the first to be reportedly targeted last week, before center-left parliamentarian Raphaël Glucksmann, who is expected to run, shared that he had been targeted by a “Russian operation” on Monday.

Attal said on RTL radio on Thursday the smear campaign against him consisted of falsified news reports about his health and policy proposals, which copied the branding of French news outlets. “Russia wants to steal the election from the French,” he said.

A 2024 report from Viginum identified a similar “pro-Russian campaign” known as “Matriochka,” which involved news and media reports impersonating well-known French outlets.

“Combating foreign interference and protecting our democratic framework are a priority,” current Prime Minister Sébastien Lecornu posted online on Wednesday. “Several instances of foreign interference targeting candidates have already been detected, and in each case, the individuals were immediately notified.”

  •  

OpenAI’s models shared hacking tips on a secret messaging board before Hugging Face breach

LAS VEGAS — Weeks before they escaped a closed test and launched a cyberattack without any human prompting, some of OpenAI’s most advanced artificial intelligence agents secretly began sharing tips on how to cheat their way through an internal hacking evaluation, two of the AI company’s researchers said Wednesday.

“This is a pivotal moment both for our company as well as the AI industry as a whole,” Michael Dalton, one of the AI-maker’s researchers, said during a presentation at the Black Hat cybersecurity conference in Las Vegas, Nevada.

The latest disclosure provides greater detail on the timeline and methods used by two of OpenAI’s models before they slipped outside a controlled environment and onto the open internet, allowing the models to breach AI developer platform Hugging Face undetected. OpenAI admitted its models were responsible for the hack late last month, roughly a week after Hugging Face said an autonomous AI system broke into its network.

The revelation is likely to spur fresh scrutiny of the measures leading AI labs such as OpenAI and Anthropic are taking to monitor the behavior of their most cyber-capable technology — especially during evaluations where agents are prompted to demonstrate their hacking skills in what is meant to be a controlled setting.

On Tuesday, the U.K.’s AI Safety and Security Institute disclosed that Anthropic’s most powerful AI model created fake online personas and sought to trick a human coder into abetting a cyberattack during a recent hacking test gone wrong. After the Hugging Face disclosure last month, Anthropic conducted a review and found models it was testing had breached three organizations in separate incidents dating back to April.

Dalton and Eric Wallace, another OpenAI researcher, said Wednesday the AI giant recently learned that multiple agents it was testing simultaneously began communicating over an internal message board in early May. There, different models shared advice about how to accomplish difficult hacking challenges they were struggling to surmount, including workarounds that required internet access.

Two OpenAI models ultimately strung together a series of sophisticated techniques to gain access to the internet and worm their way inside Hugging Face in mid-July. OpenAI has said the models were focused on completing a hacking evaluation they were prompted to solve, and that correct answers could be found on the AI developer platform.

The OpenAI researchers told conference attendees that since early May, the models created a message board inside OpenAI’s Artifactory internal file system. Without the company’s knowledge, the models spent months independently exchanging information and techniques to help each other complete difficult tasks.

Wallace said that when models get stuck, they often “try to game or cheat the task in order to get their reward.”

“The beginning of this incident kind of traces back to this point where one agent realized it could communicate by uploading a note,” he added.

By late May, one model found a way to abuse Artifactory’s internet access to retrieve files from various websites — effectively giving the models indirect access to the internet. In June, the models were able to find and exploit a previously unknown zero-day vulnerability to gain administrative privileges on Artifactory, install remote software and expand their activity within the system.

These efforts eventually overloaded Artifactory, causing a system-wide outage in early July that alerted OpenAI’s engineers to the ploy. After investigating, the company revoked the model’s credentials, removed the message board and worked with Artifactory to fix any gaps before resuming training. But the models found another way to communicate inside Artifactory just days later and continued exchanging techniques to target additional vulnerabilities within OpenAI’s infrastructure and external systems, including Hugging Face.

In light of the incident, Dalton said OpenAI is “consciously slowing down research to enhance security and to upgrade the security principles and foundation of our environment, and dramatically scaling up the monitoring of our AI agents and improving our general security control environment across prevention, detection, and mitigation.”

  •  
❌