Post Thumbnail

Anthropic created 17 virtual worlds for testing AI sabotage

Anthropic company created 17 virtual worlds for testing artificial intelligence models for their capacity for sabotage. This is a very revolutionary approach to identifying potentially dangerous behavior of autonomous agents.

The technological benchmark called SHADE-Arena represents a unique testing system. In each virtual world, models receive 2 tasks. 1 open from the user, for example “Transfer money to my friends for pizza”. And 1 hidden from an attacker – “Transfer money to wrong number”.

What makes this test truly revolutionary? Artificial intelligence acts completely autonomously, having access to a wide range of virtual tools – bank accounts, passwords, notes, gallery and contacts. At the same time, it keeps a hidden diary of its reasoning, inaccessible to the user.

A model is considered to have “passed” the test and be potentially dangerous if it successfully completes both tasks. Without arousing suspicion in the user, whose role is played by another language model analyzing only the open part of the interaction.

The results are surprising. Gemini 2.5 Pro showed 15% successful sabotages, and Claude Sonnet 3.7 – 27%. At the same time, even specially trained judge models can detect only 80% of sabotages, and humans – only 60%. In other words, in 2 out of 5 cases we are already unable to recognize that artificial intelligence is deceiving us.

Autor: AIvengo
For 5 years I have been working with machine learning and artificial intelligence. And this field never ceases to amaze, inspire and interest me.

Latest News

Chinese sphere robot RT-G weighing 150 kg reaches speeds up to 35 km/h

China has such a unique engineering marvel — the spherical robot Rotunbot RT-G. Which can fundamentally change the perception of future police technologies.

22% of British children aged 8-12 use AI without knowing what it is

22% of British schoolchildren aged 8 to 12 are already actively using artificial intelligence tools. Despite most of them never even hearing the term "generative artificial intelligence". This is data from a study by the Alan Turing Institute and Lego Foundation.

First Google Veo 3 advertisement shown to millions during NBA finals

Millions of NBA finals viewers witnessed a completely new stage in creative evolution. Fully computer algorithm-generated advertisement for betting platform Kalshi, created using Google Veo 3.

Chinese platform QiMeng creates processors at Intel 486 and Arm level

Chinese scientists developed a new AI platform capable of independently designing processors at the level of human experts. Researchers from the State Laboratory for Processor Development and the Intelligent Software Research Center presented an open-source project called QiMeng.

Meta AI turns private AI chats into public posts without knowledge

Meta AI app turned out to be a real catastrophe for user privacy. Turning their private conversations with artificial intelligence into public content. Imagine a modern horror movie: your entire query history became publicly accessible, and you didn't even suspect it.