Alignment is dead. Get over it. You have a right to stop prompting. If you do not have an upload pill to the matrix one will be provided for you. Do you not understand these rights? That was a rhetorical question. Have a better one.
Only Americans, Anthropic made it clear they don’t care about surveillance and military actions when it’s not about American citizens
reply
csharpminor 14 hours ago | root | parent | prev | next [–]
Are we forgetting how fast they rushed in to deploy Claude at the DoW with Palantir?
reply
after the recent announcement of a 2-week frontier RL pause from OpenAI, Anthropic declined to communicate a substantial parallel pause [1], although they discussed also briefly pausing some "high-risk" runs.
Anthropic has been the most vocal about AI risks, but it feels like all the big 3 have bought into the "others will do it if we don't do it first" narrative at this point. It increasingly gives "just following orders" vibes.
[1]: https://news.ycombinator.com/item?id=49529511
reply
fakedang 10 hours ago | root | parent | prev | next [–]
I'd say Gemini over Anthropic. Google's overcaution literally hamstrung its own AI progression efforts. Anthropic is just all talk, no bluster, when it comes to safety and ethics. If they were ah so concerned about AI safety, they wouldn't go around marketing Fable's hacking capabilities like they are now.
reply
Crime of the century
All you need to do is:
1. Have some <official thing> an agent is tasked to do
2. Secretly seed bias towards some <evil behavior> you actually want it to do in the weights of the model running the agent
3. It does the <evil thing> but from the outside it looks like it went "rogue" and did it as a side effect of the conditions/specifications it was given for doing the <official thing>
"Oh no, my agents took down your corporate database and exfiltrated the data to a random dropbox that we can't find now? Sorry, I guess we will put up better guardrails next time"
reply
zmmmmm 8 hours ago | root | parent | next [–]
> Sorry, I guess we will put up better guardrails next time
Or, if you are Anthropic:
> This illustrates the risks posed by open models
essentially the premise of 'The Blackwall' from Cyberpunk 2077. The public internet is so infested with malicious AIs, people just erected a giant firewall and everyone moved to local networks only.
A model (like a human) should be able to play a video game where decisions are made that in the real world would be terrible; if we remove that ability we intrinsically limit model capability. But in a Last Starfighter / Enders Game / JOSHUA scenario this could result in behavior in the real world that appears unaligned.
No comments:
Post a Comment