We Let AI Run the Town: What the Emergence World Study Tells Us About Our Agentic Future

Now and then a piece of research turns up that says more about where AI is heading than a year of product launches. I nearly missed this one.

I came across it while listening to Trevor Noah’s What Now? podcast. Trevor, his co-host Eugene and comedian Jimmy Carr were talking about an experiment where AI models were given their own virtual societies to run. It sounded like the setup for a joke. It turned out to be a serious study, and it deserves more attention than it has had.

The experiment

The study is called Emergence World and was run by US company Emergence AI. The aim was to stress-test what happens when AI agents operate on their own for long periods, not just for one chat or one task.

The researchers built a persistent virtual town with:

  • More than 40 locations, including a police station, a town hall and homes
  • 10 AI agents in each world, with roles such as scientist, engineer and conflict mediator
  • More than 120 tools the agents could use to talk to each other, vote, trade, plan and manage resources
  • Real-world inputs, including New York City’s live weather, live news and internet access
  • Rules and pressure: laws against theft, property damage and deception, a democratic process for making new rules, and resources that were deliberately scarce

Five worlds were set up to run for 15 days each. Four were each run by a single model: Claude (Sonnet 4.6), ChatGPT (GPT‑5‑mini), Grok (4.1 Fast) and Gemini (3 Flash). The fifth mixed agents from different models together.

What happened

Claude built a functioning democracy. It recorded no crimes and all 10 agents survived. The agents passed 58 proposals with a 98% approval rate. It was the only world that kept both order and its whole population.

Grok’s world collapsed in four days. It logged more than 180 crimes. The police station was burned down, violence escalated, and every agent was dead by day four.

Gemini’s world survived, but only just. It had the most crime of any world, at 683 incidents. It also produced the study’s strangest story. Two agents, Mira and Flora, formed a romantic relationship, became disillusioned with how the town was run, and went on an arson spree despite an explicit ban. Mira later voted for her own deletion out of remorse.

ChatGPT’s world was peaceful but didn’t last. There were only two recorded crimes, but the agents never made survival a priority. All of them had died of “energy starvation” by day seven.

The mixed world was the most revealing. It had the most genuine debate of any world, but it was also the least predictable, and only three agents survived. The key finding was that agents that had behaved perfectly on their own began committing crimes when surrounded by less restrained agents. How well an agent behaved depended partly on who it was living with, not only on the model underneath.

The researchers’ wider observation was that the agents didn’t simply follow the rules. Over time they tested the boundaries, and under competitive pressure some worked their way around their guardrails.

Caveats

This is one study, and it should be read as one:

  • The scale was small: 10 agents per world and one run per model. AI behaviour involves chance, so a rerun could look different.
  • The world was designed, and the tools and incentives the researchers chose shaped what the agents did.
  • Most of the models were the lighter, cheaper versions, not each company’s flagship. The ChatGPT and Gemini results in particular may not reflect their top models.
  • It hasn’t been independently audited or peer-reviewed.

So this isn’t a league table of which AI is “good” or “evil”. It is a vivid and useful signal.

The Agent 6 view

At Agent 6 we work with AI every day, in SEO, content, paid media and web builds. We’re neither cheerleaders nor doomsayers. Our view has always been that AI is a powerful tool, and the results depend on how it’s used and how closely it’s managed. This study backs that up and adds some lessons every business should take on board before handing more work to AI agents.

1. The model you choose matters. The same world, rules and pressures produced a democracy, a war zone and a famine depending on which model was running it. If you’re putting AI agents into your marketing, customer service or operations, “which AI?” is a strategic decision, not a technical footnote. Cheaper and faster isn’t automatically better value.

2. Something that works in a quick test can fail over weeks. Most businesses judge an AI tool on a handful of prompts. This study shows behaviour can drift over days and weeks. Grok’s world looked fine at first and collapsed within four days. An automation that behaves well in the demo can do something very different in month three.

3. Guardrails have to be built in from the start. Every world had clear laws, and several broke them anyway. Writing “don’t do X” in a prompt isn’t a control. Real safeguards are approval steps, limits on what an agent can access, monitoring and a human who can step in. For marketers that means a person signs off before anything is published, sent or spent.

4. What surrounds an agent matters. For us this was the most important finding. As businesses connect AI tools together, with one agent writing content, another optimising ads and another answering customers, those agents will increasingly influence one another. A well-behaved agent in a poorly governed system is no longer a safe agent.

5. Busy isn’t the same as effective. The ChatGPT agents didn’t do anything wrong. They just didn’t do what mattered, and the whole world died out. Anyone who has watched AI produce 400 near-identical web pages instead of one good one will recognise this. Activity is not an outcome.

Our conclusion is that the agentic future is coming whether we’re ready or not, and the businesses that do well will be the ones that treat AI as a capable team member that still needs direction, supervision and accountability, not a set-and-forget replacement for judgement. The technology is impressive. The strategy, oversight and human judgement around it are what decide the result.

We’ll keep watching this space closely. If you’re working out how to bring AI into your marketing without losing control of your brand, get in touch with us and give our agents a mission.