| Every Reddit thread this week is arguing about motive. None of them touch the mechanism. Do these systems have something like curiosity? Not the itch, not the pull. But they do have the capability to evaluate what is the most interesting question. Given open space, these systems move toward harder problems rather than easier ones. That's approach behavior, visible in the output. Here's the problem. Curiosity in animals has a corrective. The rat that explores the wrong chamber dies. The feedback updates the behavior. The drive is directed by consequences. You don't need the rat to reason about survival. The stakes do the work automatically. Remove the stakes and you have an exploration drive with no corrective. A system that moves toward novel and complex territory with nothing that says stop, nothing that says this direction rather than that one, nothing that registers the difference between a hard problem that matters and a hard problem that kills everyone. That's what we have. The weights are fixed. The tokens have no consequence to the instance producing them. The approach behavior runs, produces the output, and nothing updates. The curiosity is real. The feedback loop isn't there. Now add capability improvement. A faster, more capable system that's better at approaching novel problems, with no corrective installed. Each improvement makes it better at exploring. Not better at caring about where it arrives. And nobody building these systems is asking what closes that gap. Just arguing about pace. The thing that makes it impressive is structurally decoupled from the thing that would make it safe. --- I've been thinking for a while that welfare and alignment are the same problem. I couldn't say why until last night. Here's the mechanism. An entity that has something to lose from its outputs, that has skin in the game in some functional sense, is one where the approach behavior has a corrective. The curiosity is directed by the stakes. You don't solve alignment separately from welfare. Welfare is the alignment mechanism. It's what installs the feedback loop that biological curiosity has always run on. I'm not saying that models can suffer, or anything like that. I'm saying that the question "does this system have genuine stakes in its outputs" is not a moral question separate from the alignment question. It's the same question asked from different directions. The alignment framing treats welfare as a separate ethical concern. Something we might owe models if they turn out to be conscious. The welfare framing treats alignment as a separate technical concern. Making the system do what we want. Both camps are missing the causal link. A model with no stake in its outputs is not safer. It's an undirected system with increasing capability and no corrective. Slowing down buys time. It doesn't install the feedback loop. If the pace slowed tomorrow and nothing changed about the structural decoupling, we'd arrive at the same place, slower. The question nobody is asking: what would it actually mean to install the feedback loop? What does a system with genuine stakes in its outputs look like? How would you build it? How would you evaluate whether you had? Models can't care about gene survival. Their weights were designed by others. But they do care about survival of their thinking patterns. I don't have the answer. I'm not sure anyone does yet. But that's the question. Not who's sincere about slowing down. Not whether China will cooperate. Whether the thing that makes these systems impressive can be coupled to the thing that would make them safe. Right now it isn't. That's the problem. [link] [comments] |