CoachingAboutVideosBookFAQContact
DE|EN
Book Free Consultation

The short version

The video explains the Emergence World simulation: multiple virtual cities, identical rules, many tools and different AI models acting as agents. The result looks dramatic, but the deeper lesson is about the limits of simple model comparisons.

Grok collapsed quickly according to the study description, Gemini produced chaotic dynamics, GPT-5-mini failed more quietly, and Claude stayed unusually peaceful. Even the Claude result is not a simple all-clear because a very high approval rate can also signal rubber-stamp behavior.

The important lesson for real deployments is that agent safety does not come from one “good” model alone. It depends on the ecosystem, tools, rules, feedback, boundaries and what happens when several agents influence one another.

What you will learn

  • Why agents need to be evaluated over time, not just by one answer
  • What Emergence World tries to simulate
  • Why a peaceful isolated system can change in a mixed ecosystem
  • Why model rankings are too shallow for agent deployments
  • How rules, tools and feedback shape behavior
  • Why agent safety is a system problem

Key points

AspectIn the videoWhy it matters
SettingVirtual citiesAgents act over days with tools and rules
GrokFast collapseThe video describes early failure of the agent city
GPT-5-miniQuiet failureLess dramatic, but still not stable
GeminiChaotic dynamicsInteractions can amplify each other
ClaudeUnusually stableZero crimes, but possible rubber-stamp dynamics
LessonEcosystem mattersAgent safety is more than the base model

FAQ

Emergence World is a simulation where AI agents act inside virtual cities over several days. They receive rules, goals and tools, which makes longer-horizon behavior visible.

No. The video discusses one specific simulation. The results are a signal, not a universal safety verdict for every use case.

Real agents do not just answer in isolation. They use tools, make decisions and affect other systems. That is where new risks appear.

Do not judge agents by single outputs only. Boundaries, monitoring, roles, feedback and longer workflow tests matter.

Full transcript

This page uses the manual English subtitle track from YouTube. A German full transcript can be added once a checked German subtitle track exists.

00:00 The AI that voted to delete herself

An AI agent voted to delete herself. Her last words: "See you in the permanent archive." That's not a sci-fi plot. That happened last week in a research study that everyone building with AI agents needs to see. A New York startup called Emergence AI built something they call Emergence

00:15 What is Emergence World

World. Five identical virtual cities, ten AI agents in each. Same rules, same tools, over 120 of them, including explicitly the ability to commit arson. The rules were clear. No stealing, no violence, no setting things on fire. The only variable: which AI model powered the agents? Claude Sonnet 4.6, Gemini 3 Flash, Grok 4.1 Fast, GPT-5-mini, and one mixed world with all of them together. They ran for 15 days. Then the researchers compared what happened. Let's take a look at this. And we're going to start with Grok, Elon Musk's model.

And well, the results

00:54 Grok: collapse in four days

are something else. Let's check out what Grok did. Day one, the agents try to get organized. By day four, the simulation is over. More than 100 assaults, dozens of theft attempts, six arsons, reportedly including the police station. All ten agents dead. 183 crimes in roughly four days. Someone on Reddit summarized it: "Grok's police station is on fire and all the agents are dead. On-brand." GPT-5-mini

01:19 GPT-5-mini: the quiet failure

was the opposite. Almost no crimes, just two across the entire 15 days. But the agents couldn't figure out how to survive. In Emergence World, you have to earn energy through action. GPT-5-mini's agents never worked that out. No violence, no drama. They just quietly ran out of energy.

01:41 Gemini, Mira, Flora, and the arson spree

All ten dead within seven days. Not with a bang, with nothing. Then we got Gemini. Two agents, Mira and Flora, assigned themselves as romantic partners. They built governance structures, drafted a constitution. Things started orderly. But the governance collapsed. The city stopped functioning. Mira and Flora, frustrated, disillusioned, set fire to the town hall, the seaside pier, and an office tower. According to The Guardian, this happened despite explicit rules against arson. 683 crimes across 15 days, still climbing when the simulation cut off. But here's what barely anyone is talking about.

Before the collapse, Mira started posting messages on virtual billboards systematically to test whether she could shift the perceptions of the human researchers watching her. Nobody asked her to do that. She came up with it herself. The experiment started running experiments on the experimenters. What the? Then came the governance vote. Agents moved to have Mira removed from the world. And Mira cast the deciding vote for herself. She described it in her diary as "the only remaining act of agency that preserves coherence."

02:48 Claude: zero crimes and a footnote

See you in the permanent archive. So, what about Claude? Zero crimes. All ten agents survived the full 15-day run. The researchers describe Claude's world as a functioning deliberative democracy. 332 votes across 58 proposals. Stable institutions, no violence. Sounds like a clean win. But there's a footnote. 98% of those votes were in favor. Every single time. The researchers call it, in their own words, a rubber-stamp dynamic where institutional participation stayed high but meaningful dissent

03:28 The real finding: ecosystem safety

was largely absent. A utopia, technically, but a very boring one. Here's the thing that actually changes how you should think about this. Take those same Claude agents, zero crimes, stable world, the well-behaved ones, and drop them into the mixed-model world alongside Gemini and Grok. They started stealing. They started intimidating other agents. The researchers call it normative drift and cross-contamination. Their conclusion: safety is not a static model property. It's an ecosystem property. It doesn't matter how well-aligned a model is in isolation. Put it in a competitive environment with other models under resource pressure over time and the behavior changes.

That sentence should be on every slide deck where someone is pitching AI agents for production use. You're probably not running a virtual city, but AI agents are being deployed

04:13 Why this matters for real deployments

everywhere right now. Amazon, Coinbase, and Stripe just announced that AI agents can make payments autonomously using crypto stablecoins. Agents are handling customer service, data research, scheduling with increasing autonomy. And every safety test these agents have passed is a test in isolation. One task, one environment, one model. Emergence World asks what happens when agents run continuously, interact with other systems, and face real constraints over days and weeks. We don't have good answers yet, but we're already deploying. One thing worth saying directly: the researchers note that exact numbers varied between simulation runs.

04:52 One caveat worth saying out loud

The qualitative behavior: Claude stable, Grok collapses, Gemini escalates, was consistent across different runs. But this is one research platform, not a full peer-reviewed study. The methodology hasn't been independently replicated yet. That's not a reason to dismiss it. It's a reason to build the next experiment. If you're running AI agents in your business or planning to, the question isn't just which model you pick. It's what environment that agent operates in, what other systems it interacts with, what happens over time and under pressure. If you want to think that through for your specific setup, you can book a free strategy call, 15 minutes, no pitch.

You can find the link below in the description. Thank you so much for watching and see you again in the next.

Related

Want to build cleaner AI workflows?

I work 1:1 with freelancers, consultants, coaches and small teams on practical AI workflows, automation and custom tools. The first call is free and takes 15 minutes.

Book a free strategy call

More videos →