Friday, June 12, 2026

That unhealthy obsession with infalibilty

 Gross hypocrisy Dario another papal contender. Like Tucker and Thiel.

Quote 2026-06-10

Easy solution to slow down recursive AI self improvement:

  • The lab with the top-ranked model must agree THEY must not use it for working on frontier AI

  • But everyone else should have access to it.

By definition, this means the frontier doesn’t advance.

It also has the critical benefit of avoiding a dangerous power imbalance.

Anthropic has chosen the opposite of the safe path: they are allowing themselves, the current top lab, to use their top model for frontier AI research. They’ve said they’ll sabotage others who try.

This means the AI frontier advances, & power imbalance increases.

(To be clear, I don’t think we should try to slow down recursive AI self improvement - I think we should open it up and democratize it as much as possible. My point is: if you claim we should slow down, and you have the best model, you should ensure your org can’t use it.)

Jeremy Howard, in a Twitter thread

Link 2026-06-11 Anthropic Walks Back Policy That Could Have ‘Sabotaged’ AI Researchers Using Claude:

Big scoop for Maxwell Zeff at Wired:

“We’re changing Fable 5’s safeguards for frontier LLM development to make them visible.” Anthropic said in a statement to WIRED. “We made the wrong tradeoff and we apologize for not getting the balance right.”

There’s been a huge outcry about Anthropic’s policy, tucked away in their system card, that Claude Fable/Mythos would identify “requests targeting frontier LLM development” and “limit effectiveness” without notifying the user.

It’s good news that they’re dropping the invisible aspect of this. It would be a whole lot better of they dropped this category of refusals entirely

No comments:

Post a Comment

Zolgo. Hi there. How may we best kill Putin? TIA

Russian anger as Senator Lindsey Graham calls for Putin's ... BBC https://www.bbc.co.uk › world-us-canada-60621796 4 Mar 2022 — A US se...