I wanted to list out all the options for how AI could go wrong. Wrong being major human suffering or extinction.
No using AI. Let's do some good old fashioned brainstorming and maybe even rabbit hole following.
This is what I have so far. What am I missing.
Feel free to add to the list or offer subtypes under any of these
Options on how AI could go wrong:
An ASI actively trying to destroy us because it dislikes us
An ASI doing what is best for it and we are an inpediment. Doesn't hate us but we are in the way
An ASI doing what is best for it while not even considering impact to humans (ant hill theory)
A capable model or more likely Agentic aystem with the ability to get the right access doing bad while trying to do good (misalignment or is this the 3 wishes problem)
AGI with the ability to get the right access and someone writing poor instructions (paperclip optimizer)
A capable model with the ability to get the right access and a nepharious human writing destructive instructions
A capable model and a nepharious human having it help the design a doomsday weapon (biological or other)
A capable model and a good intentioned person trying to do good but ending up doing wrong (people do this all the time but not at the scale AI may allow)
At this point we want general feasibility, possible even if improbable is ok.
Might have to do a separate one of these on actual feasibility. But I'm betting I'll want to do some grouping of ideas first
[link] [comments]