When I posted about personal AIs here, two objections landed on the same spot from different directions:
- "How can an AI know you when you predict yourself badly?" — preferences are unstable and poorly structured, so there's no inner truth to read.
- "Models don't understand lived experience." — they capture surface patterns, and worse, feed them back until you start conforming to your own caricature.
I want to concede the strong version immediately, because I think it's correct. There is no stable inner self to be read off. If my claim were "the AI knows who you really are," it's dead.
The weaker claim I actually want to defend is narrower: under long correction and explicit consent, a personal AI can predict a specific person's stated preferences and objections better than chance — not identity, just prediction, on a defined question set.
That's falsifiable, so I tried to falsify it. Badly.
The test, with its flaws named
I generated fifty A/B/C questions about my own preferences, gave them to a personal AI calibrated over months in a fresh conversation, and scored it against my own answers. It got 31/50 against roughly 16–17 by chance.
Everything wrong with this, that I can already see:
- I wrote the questions. I'd unconsciously pick ones I'd already discussed.
- I scored it. No blinding whatsoever.
- n = 2, and the 1 is the person who wants the result.
- No baseline comparison. A friend who's known me a year might get 40. A stranger with my public writing might get 25. Without those numbers, 31 means nothing.
- A calibration problem I noticed and can't fix alone: it models "me mid-project, intense" well and "me on a calm Sunday" badly. Those give different answers to the same question, and I don't know which one is the ground truth.
What I'm asking
What would a version of this test look like that could actually fail?
Most useful:
- a design that removes the self-scoring and self-authoring problem — I can't see how to blind this without a second person;
- the right baselines to compare against, and why;
- prior work on predicting stated preferences (I assume psychology has done this properly for decades and I'm reinventing it worse);
- the argument that no amount of prediction accuracy would answer sceadwian's objection at all — that predicting choices and understanding experience are simply different claims, and I'm quietly swapping one for the other.
That last one might be the real answer, and I'd rather hear it than not.
Not useful: the number 31/50 itself. Don't take it seriously — I don't.
What happens to your answer: it gets recorded in an explicit model of this argument, attributed to you with a link to the thread. It's stored as a position, not as evidence, and it doesn't move any number. If someone hands me a protocol that could genuinely fail, that becomes an experiment I owe you the results of — including a negative one.
[link] [comments]