A million people, a million personal AIs, three base models. Is that a diverse deliberation — and how would you measure it?
A million people, a million personal AIs, three base models. Is that a diverse deliberation — and how would you measure it?

A million people, a million personal AIs, three base models. Is that a diverse deliberation — and how would you measure it?

Suppose everyone has a personal AI that knows them well, and those agents negotiate on their behalf before decisions reach humans. Someone raised this objection to me and I haven't been able to answer it:

Three providers can feel diverse to one person and be nowhere near diverse enough for a decision involving a million.

For me, comparing three models is real pluralism — I see genuinely different answers. But at population scale, the thing that matters isn't whether the outputs look different. It's whether the errors are independent. If a million agents share a handful of base models, a systematic blind spot doesn't show up as disagreement to be resolved. It shows up as unanimity. The deliberation would look like it was working perfectly at exactly the moment it failed.

Vendor count is obviously the wrong metric. "Three companies" tells you nothing about whether their failure modes are correlated — they train on overlapping corpora, use similar architectures, and increasingly distil from each other.

The question

What would you actually measure to tell "diversity of the represented humans" apart from "diversity of the underlying models"?

I'm after something operational — a quantity you could compute on a real deliberation and act on.

Useful to me:

  • a metric from ensemble learning or forecasting that transfers here, and what it needs as input;
  • work on correlated error in aggregation (I suspect this is a solved problem in a field I don't know);
  • an argument that the distinction I'm drawing is confused — that "represented human diversity" isn't separable from model diversity even in principle;
  • a threshold: how decorrelated is decorrelated enough, and decided how?

Not useful: "just use more models." That's the answer whose sufficiency I'm questioning.

submitted by /u/Lesterpaintstheworld
[link] [comments]