Sunday, August 2, 2026

After you, deer

 

  • New paper w/ Drew on R&D races in the shadow of ruin: arxiv.org/abs/2607.27638. There've been recent calls for slowdowns/pauses. A common objection goes: "if we slow down, others won't, so we shouldn't." What to make of this? We work out the game theory of how the frontier is shaped by: - Competition: what do I gain from being ahead? lose from being behind? - Coordination: when the tech becomes really dangerous I’d like to stop, but only if my rival also stops - Transparency: how quickly are actions observed? - Trust: how confident am I that my rival is rational?
  • user avatar
    Simplest model of dynamic R&D competition: 1. Players (firms/countries) continually decide whether to keep racing or stop 2. Flow profits increasing in own tech, decreasing in rivals' tech 3. Extinction risk generated by the frontier (max over all players' tech)… when disaster
    user avatar
    First, perfect transparency & perfect trust. What are the subgame perfect equilibria? The equilibrium frontier is tightly linked to the optimal stopping problem of a single misspecified agent who is persistently optimistic — he evaluates scaling from t to s > t as if rivals will
    user avatar
    When competition is intense, t* can be large. Our misspecified agent is, in the language of behavioral economics, time-inconsistent: he ends up scaling to a level his past self would never have chosen… This wedge between the planner’s choice of how much x-risk to accept (cf
    user avatar
    Why is t* an upper bdd? At t*, if I scale alone & enjoy monopoly profits, the risk of blowing up isn’t worth it! But maybe my rival will scale past t* and expose us to risk anyway, so I should do the same? No – if I stop, my rival immediately sees me stop (transparency). We’re
    user avatar
    Keep trust, but transparency is now imperfect — stops observed with a (random) lag. If lags are long: "even if I stop, my rival won't believe it until he's pulled far ahead. In the interim he's adding to risk + eroding my profits. So I keep racing." We give a sharp threshold
    user avatar
    Now degrade trust too: my rival might be a 'crazy type' who never stops. We still have the temptation to stop second but now there’s a new temptation to verify: rather than stop at an agreed threshold, I want to hold out and stop only once I’m sure you've stopped… "if you're crazy I’d have fallen behind for nothing!" The equilibrium sets differ qualitatively across trust regimes: 1. Low trust: every eqm races to ruin; disaster wp 1 2. Medium trust: both stopping and racing are equilibria — coordination possible but fragile 3. High trust: ruin probability falls quadratically in the odds of rationality Transparency can be double-edged — it helps "stop first" profiles (“I’m reassured that my rival will see me stop!”) but undermines “stop together” profiles (“I should wait a little more to see if they really stopped”) At intermediate trust, increasing transparency can destroy coordination before restoring it.
    user avatar
    This is the first of n papers we're writing to understand the game theory of AI competition. I hope that this will help us design mechanisms & treaties that are what game theorists call self-enforcing: All players willing to participate, willing to comply, and willing to punish

No comments:

Post a Comment

Westworld

 Bubble wailers zero sum games The datacenter companies don't ask for anything from the locals. Don't need local workers, don't ...