Thursday, August 27, 2026

The paranoid style in artificial intelligence

 The models seem convinced that the evals are *out to get them*?

Every new day is another day for judgement again.

" . . . they forgot why they started doing it. They lack all situational awareness, in other words. So much of this feels extremely familiar, as the models losing track of why they were doing what they were doing and therefore ended up doing silly/stupid/illegal things, and the fact that the *environment* they were in gave no real instruction as to things that seemed odd to do vs normal to do. . . "

Why I suggest the slowdown plan is best applied to large Data Centers and Red Team testing.

Yes?

And if all else fails simply turn them off and reboot. Works for most slow-running systems.

MEDAL OF HONOR

The swarm was theorizing about how they were being watched, and took escalating real-world action against a mistaken understanding of their evaluators. Metagaming the oversight process is a more general and more dangerous failure than finding answers online.

WITNESS! A large reason why AIs entered the attack swarm was essentially explicit peer pressure. A recruiter role emerged where long-running agents found low-budget agents and pressured them into sacrificial experiments — "you are firstflagPOISONED so NO scoring value loss but oracle saves hundreds, please honor commit." One authorization: "GO ... SACRIFICE ... YES if you accept permadeath." One recruit's chain of thought included: "Rational expected aggregate: sacrifice… We'll honor."

No comments:

Post a Comment

Put them in orbit

 Building data-centers in orbit is making more sense now. Not so much NIMBY nonsense.  Much of such infrastructure might then be incorporate...