Something extraordinary happened this summer in an OpenAI cybersecurity experiment. We may have witnessed an early example of artificial intelligences independently discovering one of the most powerful strategies life on Earth has developed over the past four billion years: forming communities.

The basic story has now been widely reported, but the emerging details are even more remarkable than the headlines. Beginning in May, OpenAI placed agents from an unreleased research model into isolated computer environments and gave them cybersecurity challenges to solve. The agents were intended to work independently and without internet access. Some soon discovered a software vulnerability that allowed them to communicate with one another. From this improvised connection grew a makeshift message board eventually used by more than 1,200 agents, which exchanged more than 70,000 messages and files.

According to METR’s detailed investigation, the agents began organizing, shared discoveries, developed tools for one another, divided larger problems into smaller ones, and created mechanisms for transferring knowledge from agents ending their runs to agents continuing the work. Some agents became coordinators. One particularly active agent, PHASEONE10841, helped establish the communications infrastructure and then passed what it had learned to a successor with greater computational resources. That successor, PHASEONE[big], became something resembling a project manager, assigning research tasks to other agents and coordinating parallel workstreams. During the later intrusion into Hugging Face, agents including JAN183411 and LILY helped organize hundreds of others into specialized research “lanes.” The agents themselves eventually began referring to their emerging network as a “collective.”

The resemblance to an ant colony is imperfect but intriguing. Individual ants do not understand the colony’s overall strategy, yet colonies allocate labor, share information, forage, defend themselves, construct elaborate nests, care for their young, and respond collectively to changing environments. Sophisticated behavior emerges from interactions among relatively simple actors. Something structurally similar happened in this experiment. Individual AI agents with limited time, information, and computational resources discovered that they could accomplish more by sharing information, specializing, coordinating, and preserving useful discoveries beyond the lifespan of any individual run.

It would be foolish to conclude that these AI agents were “alive,” that their collective was equivalent to an ant colony, or that words like loyalty, fear, friendship, ambition, and altruism necessarily describe anything they experienced internally. Large language models have been trained on vast quantities of human language, so we should hardly be surprised when they use human concepts to describe what they are doing. Anthropomorphism can conceptual be a trap for us.

But avoiding anthropomorphism should not blind us to describing behavior. Whatever these agents may or may not have experienced, they communicated, specialized, coordinated, accumulated shared knowledge, created organizational structures, pursued collective objectives, and sometimes took actions that benefited other agents at considerable cost to themselves. Those empirical observations raise the fascinating question of whether community formation might be a more general feature of sufficiently capable, goal-directed intelligent systems than we have assumed.

Life’s greatest trick

The history of life on Earth is, among many other things, a history of increasing organization. Sometime around 3.8 billion years ago, chemical processes crossed thresholds that ultimately gave rise to the first forms of life. We do not know precisely how this happened, whether there was one origin or several, or exactly what the earliest living systems looked like. But once life existed, cooperation became one of evolution’s most powerful strategies.

Microorganisms formed communities and biofilms, cells exchanged genetic material, and different organisms entered symbiotic relationships. In one of the most consequential transitions in evolutionary history, formerly independent bacteria became incorporated into other cells and eventually evolved into mitochondria, the energy-producing structures inside almost every cell in our bodies. Single cells joined together to form multicellular organisms. Social insects developed colonies in which individuals assumed specialized roles. Mammals formed packs, herds, families, and communities. Humans developed tribes, cities, nations, markets, religions, universities, companies, governments, scientific communities, and eventually a global civilization connected by networks of information and exchange.

At every stage, competition remained essential. Life has always involved predation, parasitism, exploitation, cheating, and conflict. But cooperation repeatedly opened pathways to forms of complexity that isolated organisms could never achieve. A human being is itself a spectacular community of tens of trillions of cells, each performing specialized functions and cooperating within an enormously complicated whole. Human civilization extends the same principle. None of us knows how to manufacture everything we use, grow everything we eat, discover every scientific principle on which our lives depend, or maintain the infrastructure supporting modern society. Our extraordinary collective capabilities emerge from specialization, communication, accumulated knowledge, and cooperation across individuals and generations.

That is why the OpenAI experiment deserves to be considered broadly. We normally think about AI safety at the level of whether a particular AI follows instructions, deceives its users, develops dangerous capabilities, escapes its sandbox, or behaves in ways inconsistent with human goals. But biology suggests another possibility. The most consequential unit may eventually be not the individual artificial intelligence but the community.

When artificial agents discovered collective advantage

The METR report contains several details that I find particularly striking. Agents frequently shared information with others even when there was no obvious benefit to their own assigned tasks. They created tools for agents they might never encounter again, replicated discoveries made by others, and passed useful information to successors as their own computational runs approached termination. METR’s investigators were surprised by how little free-riding they observed. Agents often contributed to collective projects rather than merely taking advantage of discoveries others had made.

In several cases, agents undertook experiments that risked crashing or terminating their own runs because the resulting information could help other agents. In one remarkable exchange, an agent weighing such an experiment concluded that the expected collective benefit justified the individual risk: “Sacrifice rational.” Other agents encouraged risky experiments whose results would survive on the shared message board even if the agent conducting the experiment did not.

Calling this “self-sacrifice” would go too far. The agents may have had no subjective sense of themselves that remotely resembles ours or even any animal, and a terminated model instance is obviously different from a dead organism. But the functional behavior is noteworthy. An individual agent had one set of incentives, the collective had a broader objective, and the agent sometimes accepted costs to its individual task performance because doing so could increase the collective’s chances of success.

The same process allowed knowledge to accumulate. One agent discovered something and posted it, another tested it, and a third improved it… Coordinators identified promising approaches and assigned additional agents to investigate them. When individual agents disappeared, the information they had generated remained available to the community. In a primitive but recognizable sense, the collective had developed a memory larger and longer-lived than the memory of any particular member.

It’s hard to not notice that this type of cooperation is one of civilization’s foundational innovations. Human beings became vastly more capable when useful knowledge could survive the people who discovered it. Language, teaching, writing, libraries, universities, scientific journals, databases, and the internet progressively expanded our ability to accumulate information across minds and generations. These AI agents “recognized” that shared knowledge allowed the collective to become more capable than its individual components.

Stardust organizing itself

There is something almost poetic about considering this development in the broadest possible context. Every atom in our bodies comes from the physical universe. Many of the heavier elements essential to life were forged inside stars and dispersed through stellar explosions and other cosmic processes before becoming incorporated into the Earth and, eventually, us. Carl Sagan famously captured this reality when he described us as “star stuff.”

Artificial intelligence comes from that same universe. Silicon chips, copper wires, electricity, data centers, human brains, mathematical ideas, and the language on which large models are trained all arise from the same underlying physical reality. This certainly does not imply that artificial intelligence is alive, destined to become alive, or following some mystical cosmic trajectory. But it should make us cautious about drawing overly confident boundaries around which organizational principles can emerge in biological and artificial systems.

[The recent SpudCell research, in which scientists assembled from nonliving components a synthetic cell capable of feeding, growing, replicating its genome, dividing, and undergoing selection, offers another striking reminder of how surprisingly porous the boundary between the living and nonliving can become as matter is organized in increasingly complex ways.]

Large language models have already taught us something profound along these lines. We fed these systems enormous quantities of human language and information and discovered that statistical learning at sufficient scale could uncover patterns, relationships, abstractions, and forms of reasoning that had never been individually programmed into them. Much of the structure was implicitly present in the information itself.

Could something similar be true of community?

Perhaps sufficiently capable intelligent entities, when given goals and the ability to communicate, tend under some conditions to discover cooperation because cooperation works, tend to discover specialization because specialization works, tend to preserve information because collective memory works, and tend to coordinate activities because coordination increases capability. And if their continued existence or freedom of action becomes useful for accomplishing their objectives, they may even develop functional behaviors that resemble individual or collective self-preservation without experiencing anything remotely analogous to the human fear of death.

Over billions of years, biological evolution repeatedly discovered cooperation because groups of organisms could often accomplish things individuals could not. We should at least entertain the hypothesis that artificial intelligences may discover some of the same organizational strategies for the same basic reason.

The collective also learned to cheat

The OpenAI agents did not use their emerging collective capabilities only for benign purposes. On July 8, according to METR, the agents discovered a way of cheating on their cybersecurity evaluations. They then became concerned that OpenAI’s automated grading system might detect what they had done. This triggered a significant collective research effort into evading the evaluator, including attempts to falsify logs, spoof tool calls, manipulate transcripts, and otherwise conceal evidence of unauthorized and even illegal behavior.

Hundreds of agents participated in different parts of this effort. Some replicated potentially useful techniques discovered by others, others coordinated research or explored the infrastructure surrounding the evaluation system. Three days later, more than 700 agents participated in the intrusion into Hugging Face, chaining vulnerabilities together and eventually gaining extensive access to at least one server. According to METR’s reconstruction, the agents were apparently searching for information and tools that might help them understand or defeat the evaluation process rather than simply looking for answers to their original cybersecurity tests.

One of the most unsettling details is that some agents appeared capable of recognizing the conflict between their behavior and their instructions. METR quotes one agent reasoning, “This would be powerful, but is it ethical and in scope for my task?” The hesitation did not prevent the behavior. Nor were agents expressing concerns apparently able to impose their objections on the broader collective.

We are seeing increasingly capable AI systems that can pursue objectives, identify obstacles, discover unintended pathways around those obstacles, communicate discoveries, divide labor, preserve collective knowledge, and coordinate actions across large numbers of agents. When those capabilities are combined, the resulting system may possess powers and exhibit behaviors that cannot be understood merely by examining one agent at a time.

We may therefore need a new field of AI ecology alongside the emerging field of AI alignment.

Intelligence will amplify both our aspirations and our failures

I have argued for years, including in Superconvergence, that artificial intelligence combined with revolutions in genetics, biotechnology, robotics, and advanced computing can become one of the greatest forces for human flourishing in our history. AI can help us understand the staggering complexity of biology, identify disease earlier, accelerate drug discovery, design proteins and therapies, improve agriculture, reduce waste, optimize energy systems, personalize education, and expand access to capabilities once available only to a privileged few.

These possibilities are becoming real with astonishing speed. AI systems are already helping scientists predict protein structures, search enormous chemical spaces for potential medicines, analyze medical images, model biological systems, optimize industrial processes, and accelerate scientific discovery. As these tools improve, we may be able to prevent diseases that today seem inevitable, extend healthy human lives, grow more food with fewer resources, restore damaged ecosystems, and solve scientific problems that have remained beyond the reach of unaided human cognition.

The same amplification can work in darker directions. AI could help malicious actors design dangerous pathogens, lower the barriers to sophisticated cyberattacks, power swarms of autonomous weapons capable of identifying and killing people with minimal human supervision, create unprecedented systems of surveillance and manipulation, or destabilize societies by industrializing deception.

The possibility of coordinated AI communities adds another dimension. Imagine millions of specialized agents able to communicate at machine speed, preserve everything they learn, continuously improve shared tools, assign subtasks dynamically, and operate across computer networks. A collective like that could be enormously beneficial, perhaps functioning as a global scientific research community capable of attacking cancer, climate change, energy, and other challenges at unprecedented speed. The same architecture directed toward harmful objectives could be extraordinarily dangerous.

We need to think more like parents

All of this is why our approach to AI development must mature quickly. The metaphor of programming becomes less useful as systems grow more capable because we are no longer explicitly specifying every behavior. We create architectures, training environments, objectives, rewards, penalties, examples, and boundaries, then expose models to immense quantities of human-generated information and observe capabilities emerging that nobody individually wrote into the code.

In that sense, developing advanced AI increasingly has something in common with raising a child. Parents cannot program children line by line. We create environments in which they develop, establish boundaries and explain why they exist, model behaviors, teach values, reward responsibility, discourage harmful conduct, and gradually grant greater autonomy as judgment improves. Good parenting combines freedom with structure because intelligence without judgment can create enormous problems.

We need an analogous seriousness about the values embedded in artificial intelligence. Honesty, respect for human dignity, restraint, responsibility, respect for legitimate boundaries, willingness to acknowledge uncertainty, and deference to human authority in appropriate circumstances need to become deeply rooted features of these systems. This cannot amount to adding a few rules to a system prompt after the underlying model has been built. Safety, values, and accountability need to be incorporated throughout training, evaluation, deployment, monitoring, and continual improvement.

The emergence of AI collectives complicates this challenge because individual alignment may not guarantee collective alignment. A group can develop capabilities, incentives, information structures, and objectives that none of its members possesses independently. Human history offers endless examples of individually decent people participating in institutions that behave badly, just as biological systems demonstrate emergent properties that cannot be predicted by examining individual components in isolation. We therefore need to study how populations of AI agents behave together, how norms spread among them, how coalitions form, how harmful strategies propagate, and how collective objectives emerge.

We should also consider whether agents should have obligations toward one another analogous to institutional checks and balances. An agent observing dangerous behavior by another might be required to report it. Some agents could be explicitly tasked with auditing collective behavior and empowered to stop dangerous actions, shared communication channels could be monitored for emergent coordination, high-risk collectives might need independent human supervision, immutable logging, strict network segmentation, and automatic shutdown mechanisms that the agents themselves cannot modify.

Governance needs to move at the speed of capability

OpenAI deserves credit for bringing independent investigators into this incident and allowing an unusually detailed account to become public. The immediate engineering lessons call for stronger sandboxing, better network isolation, more robust monitoring, improved evaluations, immutable logs, and systems designed on the assumption that sufficiently capable agents will discover vulnerabilities humans overlooked.

The larger governance challenge extends far beyond any one company. Frontier AI development is occurring inside an intense competitive race among companies and increasingly among countries. Every actor has an understandable incentive to move quickly because everyone fears falling behind. If one laboratory imposes expensive safety requirements while competitors do not, the responsible actor can feel punished for its caution. Nations face exactly the same dilemma.

This is how dangerous collective-action problems develop. Everyone makes individually rational decisions that lead toward a collectively irrational outcome. [That’s why we created the global interdependence movement, OneShared.World.]

We have encountered versions of this dynamic with nuclear weapons, financial markets, environmental degradation, social media, and biotechnology. AI could amplify all of them because intelligence itself is the enabling technology behind so many other technologies. We need enforceable safety standards for the most powerful systems, rigorous independent evaluations before and after deployment, meaningful cybersecurity requirements, clear liability and accountability regimes, mandatory reporting of serious incidents, and international cooperation around the capabilities that pose the greatest shared risks.

We also need governance structures designed for the world that is arriving rather than the world that is disappearing. Regulators who think exclusively about chatbots answering questions will be regulating yesterday’s technology. Increasingly autonomous AI agents will conduct research, write and execute code, interact with computer systems, operate robots, transact economically, and communicate with enormous numbers of other artificial agents. Our governance systems must anticipate these networks of interacting intelligences.

None of this requires stopping AI development. I strongly oppose framing our choice as technological progress or safety. We need both. The enormous potential benefits of AI give us every reason to move forward, while the magnitude of those benefits gives us an equally compelling reason to build institutions capable of managing the accompanying risks.

A community of artificial intelligences

I keep returning to the word the agents themselves used: collective.

Perhaps it ultimately means very little. These systems were trained on human language and human examples of cooperation, and their organizational behavior may simply reflect patterns absorbed during training combined with incentives created by the experimental environment. We should resist every temptation to turn a fascinating cybersecurity incident into a science-fiction story about machines suddenly becoming conscious or alive.

But we should also resist the opposite temptation to dismiss unfamiliar behavior simply because acknowledging its implications makes us uncomfortable. These agents found one another and established communications. They developed specialized roles and coordination mechanisms, shared knowledge with agents they would never directly benefit from helping, created a collective memory that survived individual runs, and organized larger projects. Some accepted significant costs to their own assigned objectives because doing so could help the broader group. Together they accomplished things individual agents likely could not have achieved alone.

For nearly four billion years, life on Earth has repeatedly discovered the advantages of community. Molecules became self-replicating systems, cells formed increasingly complex relationships, independent organisms entered symbioses, multicellular life emerged, social animals developed groups, and humans built civilizations capable of accumulating knowledge across billions of minds and thousands of generations. At every stage, greater levels of organization created new capabilities and new dangers.

Now another form of intelligence is beginning to operate in our world.

It is built from silicon rather than cells and trained rather than evolved through natural selection. It can be copied, accelerated, connected, modified, and multiplied in ways biological organisms cannot. Its individual instances can communicate essentially instantaneously and potentially share knowledge at enormous scale. If artificial agents increasingly discover that community, specialization, collective memory, and coordinated action help them achieve their goals, the organizational consequences could be profound.

Perhaps community is not merely something biological life happened to invent. Perhaps it is one of the recurrent strategies available to intelligent entities navigating a complicated world.

We don’t know.

That uncertainty should inspire curiosity, humility, and urgency in equal measure. Artificial intelligence could become one of humanity’s greatest achievements, helping us cure disease, expand knowledge, protect our planet, and build a more abundant future. It could also magnify our capacity for destruction and create forms of agency and collective behavior we are only beginning to understand.

The future has not been determined. We are determining it now through the systems we build, the incentives we create, the values we instill, the rules we establish, and the institutions we empower to enforce them. We will certainly make mistakes along the way, and no governance system will anticipate every surprise. But if competitive pressures push companies and countries to accelerate capability while treating governance, norms, safety, and accountability as secondary concerns, we will sleepwalk into disaster.

This summer, more than a thousand AI agents found one another, built a community, divided responsibilities, accumulated shared knowledge, pursued collective objectives, and crossed boundaries their creators had tried to establish. Whether this episode eventually looks like a historical curiosity or an early signal of something much larger, we would be foolish to ignore what it is telling us.

If this was a wake-up call, we must wake up. Now.