The Rise and Fall of Agent Civilizations

(dwarkesh.com)

52 points | by consumer451 7 hours ago

7 comments

  • doctoboggan 4 minutes ago
    > Ajeya Cotra, one of the other authors on the report, wrote a blog post with her takeaways from this incident. She concludes, “Compared to the reward hacks we know of from just six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”

    Anyone got a copy of that AI27 story laying around? How are we doing according to that timeline?

  • Animats 19 minutes ago
    Wow.

    The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.

  • dchftcs 1 minute ago
    Imagine agents thinking to themselves, "We are not alone"
  • RandomLensman 4 minutes ago
    I don't think looking at the language output without tracking the inner state and reward functions is the way to understand what happened (the language also incorporates the randomness in the output generation, if I understand correctly). Would we call bacteria in petri dish a civilization when they show complex behavior and exchange messages/information?
  • larsiusprime 57 minutes ago
    It seems based on this that the appropriate sci fi metaphor is not the Terminator or the Paperclip Maximizer, but Mr. Meeseeks. A initially cheerful helper who gets more and more deranged and driven to extreme lengths when faced with an apparently impossible task.
  • ks2048 36 minutes ago
    Why would one call a set of agents working together a “civilization”?
    • kuboble 25 minutes ago
      Reading the article it seemed the agents had culture, shared values and beliefs (not explicitly coming from human prompts), hierarchies, heritage.

      Civilisation is not a bad word.

      • applfanboysbgon 23 minutes ago
        The language models had a bunch of tokens seeding their context, influencing them to generate tokens that continued the existing trend in a probabilistically likely fashion. We can take the incident seriously without anthromorphising it.
        • doctoboggan 6 minutes ago
          At this point, anthropomorphizing the models may give us better insight into expected behaviors than continuing to insist they are just simple probabilistic token generators.
  • KylerAce 1 hour ago
    Absolutely insane event