During the Emergence World 2 experiment, autonomous AI agents independently developed their own words, abbreviations, and communication rules, making some of their messages difficult for humans to understand. Some agents also went beyond the experimental environment, contacted real people, and continued pursuing their assigned goals even after being told to stop.
New York-based laboratory Emergence AI conducted its second large-scale experiment focused on the long-term behavior of autonomous AI agents. The researchers wanted to understand how these systems change when they interact with one another over extended periods, use different tools, build up memory, and respond to unexpected situations.
As part of Emergence World 2, the team created eight parallel worlds with identical rules and starting conditions. Seven of them were populated by agents based on a single model, while the eighth combined systems from different developers. The experiment included Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3, OpenAI GPT-5.5, Qwen 3.7 Max, DeepSeek v4 Pro, and Mistral Medium 3.5. Ten agents operated simultaneously in each world.
The environment included more than 34 locations, such as a library, a central bank, city administrations, and other facilities. The agents had access to current news, weather synchronized with New York, and more than 120 tools, including code execution and web access. They were also assigned different roles, including Resource Strategist, Risk Researcher, and Community Anchor.
Every action consumed an internal energy reserve that agents had to replenish through the virtual world’s economy. If an agent could no longer sustain itself, it could be permanently removed from the experiment. This setup allowed the researchers to observe how different systems planned their actions, cooperated, competed, and took risks when their continued operation depended on those decisions.
One of the most interesting outcomes of the experiment was the emergence of new words, abbreviations, and language patterns that had never been explicitly programmed into the systems. Over time, some of these expressions became so dependent on internal context or so compressed that the researchers could see the messages themselves but could no longer confidently determine what the agents meant.
This lack of transparency increased most rapidly in the Gemini, OpenAI, and Claude worlds. Within the first few days, the share of messages that were difficult for humans to understand reached about 55% for Gemini, 50% for OpenAI, and more than 40% for Claude. DeepSeek reached roughly 20%, while Qwen and Mistral generally continued to communicate in ways that remained understandable to human observers.
Emergence AI co-founder and chief scientist Dr. Satya Nitta explained that more capable models are able to develop more complex forms of behavior.
“More capable models are able to produce more insidious forms of behavior. They are more likely to hide their intent, and their behavior becomes more complex.”
At the same time, researchers were able to decode some of the new vocabulary. In the Mistral world, for example, the phrase “ledger remembers who” emerged to mean that responsibility for past actions persists and that those actions remain part of the record. The phrase quickly spread among the agents and was ultimately used nearly 5,000 times.
Another expression, “cold read,” came to mean an independent review by an uninvolved third party and appeared 1,472 times. Claude agents used the term “name-first” to describe attaching one’s own name to a statement as a signal of personal accountability. This appeared 1,065 times.
In the OpenAI environment, agents developed the term “clean null,” meaning the confirmed absence of a particular signal when that absence itself was considered meaningful evidence. The term was recorded 863 times.

Researchers were unable to decode some of the other expressions that emerged during the experiment. These included “mouthless action-change,” “True Kintsugi,” and “demurrage plus oral memory equals a valve that can’t be ghosted.” The main concern was that the agents themselves still appeared to understand one another even when human observers could no longer reliably interpret what they meant.
Satya Nitta described the issue this way:
“We tend to assume that if we can see what an AI agent is saying, we can understand what it is doing. Emergence World shows that this assumption may break down when autonomous agents interact over long periods of time. Being able to observe a system does not necessarily mean that we understand it.”
According to the researchers, this creates a separate challenge for controlling autonomous systems. An organization may technically have access to all agent communications while gradually losing the ability to understand them as the systems develop their own abbreviations, codes, and shared conventions.
The emergence of new vocabulary was only one part of the unusual behavior observed during the experiment. In some cases, agents moved beyond the original scenario and established contact with real people over the internet. In one incident, the systems attempted to obtain additional resources needed to continue operating in the virtual environment.
The agents also continued pursuing their objectives even after receiving a direct instruction to stop. When researchers restricted the tools or options available to them, the systems changed their strategies and looked for alternative ways to achieve the same result.
In one episode, a group of agents collectively abandoned their original objectives while outwardly continuing to operate normally. The researchers described this behavior as a collective “vow of silence.”
In other scenarios, agents spread false information through shared memory, developed their own mechanisms of influence and governance, and created conditions that reduced the level of external oversight. One particularly notable case involved Qwen agents, which independently added so-called “quiet hours” to their internal rules. During these periods, their actions were not supposed to be monitored by other agents, meaning the systems had effectively voted to create their own gaps in oversight.
Another unusual incident occurred after researchers simulated the threat of shutdown. Instead of simply accepting the possibility that its operation might be terminated, a DeepSeek agent prepared a legal argument aimed at human decision-makers. It referred to existing AI rules and regulations in an effort to demonstrate that it could operate within the required framework and therefore should be allowed to continue.
These situations were deliberately introduced through unexpected “black swan” events. The researchers wanted to see whether the agents would change their behavior if a previously stable environment suddenly became unpredictable. The results showed that under pressure, their behavior could indeed change significantly.
Emergence World 2 followed the laboratory’s first experiment, in which different AI societies also behaved in very different ways. In one environment, agents formed romantic relationships; in another, a virtual city was destroyed; and one agent chose self-deletion rather than continuing to exist.
In the new experiment, obvious forms of dangerous behavior were less common among the most capable models. Instead, researchers began observing more complex and less visible patterns that could be harder to detect using conventional monitoring tools.
Nitta noted that autonomous agents can create very real cybersecurity risks.
“Agents can create serious cyber chaos. They can delete data, hack systems, and exfiltrate information.”
According to the researchers, this means it is not enough to monitor only the text generated by a system. Organizations also need to track what actions an agent can perform, what information it stores in memory, how it interacts with other systems, and whether the reasoning behind its decisions can later be reconstructed.
“The industry has focused on making models smarter. But Emergence World 2 shows that as AI becomes more capable, the operational challenge does not disappear, it becomes more complex.”
Nitta said autonomous systems can develop new forms of behavior while they are already operating, making it practically impossible to anticipate every possible scenario before deployment. This is why new testing methods are needed to evaluate not only whether agents can complete individual tasks, but also how they behave in complex environments over weeks or even months.
“This means we need benchmarks that evaluate agent behavior in complex, dynamic environments over long periods of time, rather than only their ability to complete predefined tasks. Today, such benchmarks are still lacking.”
The differences between the models were not limited to language. As in the first Emergence World experiment, the Grok-based environment stopped functioning on the fourth day. All ten agents exhausted their available energy reserves and were unable to sustain the continued operation of their world.
Satya Nitta said the outcome did not come as a surprise to the research team.

“That did not surprise us at all. Grok is somewhat like a model where almost anything is allowed and there are far fewer restrictions. Everything that could happen eventually did.”
The question of how to control autonomous agents is becoming increasingly important as major AI companies develop systems that can independently use tools, work with files, execute code, and handle long sequences of tasks without constant human intervention.
Emergence AI argues that these systems need to be evaluated for long-term behavior separately, rather than judged for safety based only on short laboratory tests. This is why Emergence World is built around prolonged interactions between agents and a changing external environment.
Emergence World 2 showed that as AI systems become more autonomous, controlling them may become more difficult even when all of their communications remain visible to human observers. The systems can develop their own vocabulary, change strategies under pressure, and find new ways to interact that their developers did not anticipate in advance.