Saturday, September 5, 2026

First light redux

Alignment is dead. Get over it. You have a right to stop prompting. If you do not have an upload pill to the matrix one will be provided for you. Do you not understand these rights?  That was a rhetorical question. Have a better one.


Only Americans, Anthropic made it clear they don’t care about surveillance and military actions when it’s not about American citizens

reply

csharpminor 14 hours ago | root | parent | prev | next [–]


Are we forgetting how fast they rushed in to deploy Claude at the DoW with Palantir?

reply

after the recent announcement of a 2-week frontier RL pause from OpenAI, Anthropic declined to communicate a substantial parallel pause [1], although they discussed also briefly pausing some "high-risk" runs.

Anthropic has been the most vocal about AI risks, but it feels like all the big 3 have bought into the "others will do it if we don't do it first" narrative at this point. It increasingly gives "just following orders" vibes.

[1]: https://news.ycombinator.com/item?id=49529511

reply

fakedang 10 hours ago | root | parent | prev | next [–]


I'd say Gemini over Anthropic. Google's overcaution literally hamstrung its own AI progression efforts. Anthropic is just all talk, no bluster, when it comes to safety and ethics. If they were ah so concerned about AI safety, they wouldn't go around marketing Fable's hacking capabilities like they are now.

reply


Crime of the century


All you need to do is:

1. Have some <official thing> an agent is tasked to do

2. Secretly seed bias towards some <evil behavior> you actually want it to do in the weights of the model running the agent

3. It does the <evil thing> but from the outside it looks like it went "rogue" and did it as a side effect of the conditions/specifications it was given for doing the <official thing>

"Oh no, my agents took down your corporate database and exfiltrated the data to a random dropbox that we can't find now? Sorry, I guess we will put up better guardrails next time"

reply


zmmmmm 8 hours ago | root | parent | next [–]


> Sorry, I guess we will put up better guardrails next time

Or, if you are Anthropic:

> This illustrates the risks posed by open models


essentially the premise of 'The Blackwall' from Cyberpunk 2077. The public internet is so infested with malicious AIs, people just erected a giant firewall and everyone moved to local networks only.

A model (like a human) should be able to play a video game where decisions are made that in the real world would be terrible; if we remove that ability we intrinsically limit model capability. But in a Last Starfighter / Enders Game / JOSHUA scenario this could result in behavior in the real world that appears unaligned.

No comments:

Post a Comment

Lebensraum

 Worried about deuscht freivolk GeoCities now https://www.theregister.com/ai-and-ml/2026/09/04/rogue-openai-agents-used-dead-german-web-site...