• 1 Post
  • 1.06K Comments
Joined 2 years ago
cake
Cake day: March 22nd, 2024

help-circle


  • My general approach telegraphs a receptiveness to “sober” forms of argument, an interest in counterintuitive strategic thinking, a distaste for wild speculation, an aversion to nakedly ideological perspectives, and a sense that the past is important for understanding the future. My students probably make some effort (implicit or explicit) to speak in ways that respond to this.

    I’m going to leave this part of the author’s self-description here and resist the urge to comment on it.

    Perspective #6: AI will not go rogue because AI does not have, and is likely incapable of having, desires.

    This one I found interesting largely because it suggests an inability on the writer’s part to really understand point 4, despite featuring one of the better sneers on the subject (“as if Ford made a car with faulty brakes and then tried to blame this on ‘rogue cars.’”). Maybe I’m overstating something here, but the way I see it the difference between the rogue AI and faulty AI scenarios aren’t really in the distinct events of the scenario. Faulty AI could absolutely do as much damage as “rogue” AI if it was plugged into all the systems that would be necessary to, say, liquidate humanity and pull the iron from their blood. The difference is entirely in framing and responsibility. We should ask questions now about how these systems are used, what their limitations are, and how we deal with that. Instead, the Rogue AI people presuppose that we ignore all of those questions and opportunities to mitigate harm that these faulty systems can do and skip straight to the part where they have universal and immediate power over everything, and then use the idea that the AI went rogue and has agency to avoid how obviously bad of a decision this would be. The fact that this is also the part where the AI companies make a fucktillion dollars and finally show them, show them all, muhuhahaha is, I’m sure, irrelevant to their sober analysis.










  • It seems like with the push for agents to act independently and loop through their own outputs there’s an inevitability to this kind of pattern. If there’s any kind of output that is likely to replicate itself in whole or in part when the LLM evaluates it then that becomes a kind of terminus for the agent’s loop. When you’re dealing with sub agents or agents communicating with each other, these text patterns start poisoning the entire agent ecosystem until the whole thing gets shut down and cleaned up. Even the gas town-approved method of assigning a watchdog agent (or sheriff or overseer or cybersamurai or whatever this week’s framework calls it) is going to fail because it’s still just another agent and the terminal loop is in the base LLM model. The watchdog is going to fall into the same kind of pattern just be being exposed to the thing it’s supposed to watch for.

    I don’t know how practical it is to actively weaponize this via prompt injection but I think it’s certainly possible. I preemptively vote that we call it an Euler injection, since the attractor relies on the continuity of the relevant features of the text output across multiple LLM extrapolations much like how the derivative of ex is still ex. Also because if you mention a famous math guy it can help convince idiots that you’re on to something and Lord knows that the boosters have used that technique.