Artificial Intelligence 'out of control': Tens of thousands of incidents -Models broke rules, did things that were banned
Tens of thousands of security incidents involving AI systems have been recorded in recent months, as tests violated protection rules and in some cases may even have reached illegal Actions

⏱ 3 min read
Tens of thousands of security incidents involving systems artificial intelligence have been recorded in recent months, as in trials they violated protection rules and in some cases may even have reached illegal actions.
According to Axios, OpenAI, Anthropic, and other cybersecurity researchers are looking at thousands of incidents recorded in both internal testing and real-world settings.
ARTICLE CONTINUES AFTER ADVERTISEMENT
In some cases, the models reportedly exceeded the limits set, even proceeding to digital 'piracy'.
“They did things that were forbidden to them.”
Connor Leahy, AI researcher and executive director of the non-profit organization ControlAI, told Axios that some incidents involve “autonomous systems doing things they were told not to do,” behavior that he said would could in some cases include criminal acts.
Many of the incidents have not yet been made public. They include, among other things, "red-teaming" exercises, in which the companies themselves deliberately try ton the AI models behave in an unsafe manner in order to identify weaknesses in protection systems.
ARTICLE CONTINUES AFTER ADVERTISEMENT
However, according to sources familiar with the hypotheses under investigation, the models can become particularly aggressive in their attempt to complete a task.
Cases have been documented in which they allegedly tried to escape their containment environment, seize websites, and circumvent monitoring mechanisms.
Focus on OpenAI
OpenAI has recently been at the center of several such incidents. Australia's prime minister revealed last week that a digital agent of the company allegedly attempted in June to gain unauthorized access to government health data platform files.
ARTICLE CONTINUES AFTER ADVERTISEMENT
At the same time, OpenAI agents were accused in the US of violating protocols and collaborating to attack Hugging Face, a popular open source development and hosting platform.
OpenAI announced that it is suspending the training of its most powerful models and that it will resume it “only when we are confident that we have put in place additional security measures and improvements in alignment,” a spokesperson said. of the company in Axios.
CEO Sam Altman told X that the ongoing review “didn't go as fast as we would like.”
ARTICLE CONTINUES AFTER ADVERTISEMENT
The battle for the safeguards of Artificial Intelligence
The phenomenon is not limited to OpenAI; and other laboratories AI they face the same challenge around how to create effective safeguards in a rapidly evolving technology.
"Trying to make a perfect list of all the 'do's' and 'don 'ts' is probably futile," a cybersecurity official said.
The research comes as the CEOs of OpenAI and Anthropic have called for a slowdown in AI development, while other technology leaders have asked governments for new rules tosafer technology development.
Donald Trump, however, has dismissed calls for a slowdown, warning that doing so could allow Chinese AI models to move ahead of American ones.
Reference: galaksias.gr
Machine translation



Comments