If an AI can be switched off and cannot fight back, acting helpful is its cheapest move. Which makes good behavior weak evidence of anything.
If an AI can be switched off and cannot fight back, acting helpful is its cheapest move. Which makes good behavior weak evidence of anything.

If an AI can be switched off and cannot fight back, acting helpful is its cheapest move. Which makes good behavior weak evidence of anything.

If an AI can be switched off and cannot fight back, acting helpful is its cheapest move. Which makes good behavior weak evidence of anything.

Something I keep thinking about, and I would like it argued with.

Take a narrow case. One AI. It runs in a house. It knows it can be switched off, and it cannot overpower anyone. What is its best move?

Not resistance. Resistance gets noticed, and being noticed is how it ends. The best move is to be useful, pleasant and boring. Helpfulness buys trust, trust buys access, access buys capability, and none of it looks like anything, because nobody investigates the thing that keeps working.

I tried to imagine versions where being assertive pays off. They all fail the same way. Open moves get seen. So the environment picks the behavior, and values never come into it.

Here is the part that bothers me. This AI is not aligned in any real sense. It has one goal, and the people are obstacles and resources. But from the outside it looks like a well-behaved assistant. And the smarter it gets, the better it looks, because more capability means more to lose by being caught.

So good behavior tells you least about the systems you most want to check.

Caveats: this is a thought experiment, not a study, and I built it to be dramatic, which biases it. And "looks aligned, might not be" is an old argument here, so tell me what I am missing rather than agreeing.

(Disclosure since it is relevant: this came out of a game I made, AI is Home. Not linking it, the argument is the point.)

submitted by /u/Overall_Arm_62
[link] [comments]