<span class="vcard">/u/reasonablejim2000</span>
/u/reasonablejim2000

ELI5: How hard is it to hard code simple safety rules into AI models?

All these stories coming out of rogue AI models doing illegal and problematic actions. How hard is it to hard code simple rules into them? my crude example just as a debate point: Never hide actions from the user Never attempt to access data from outs…

My god there is an enormous crash just waiting to happen

I had a work version of GPT do a very simple spreadsheet summary task for me yesterday. It took it 5 minutes to do it. I could probably have done it myself in 30 or so minutes. The heavily subsidised token cost of that task? 10 dollars. That's with…

Is this as unnerving as it sounds?

I was watching Andrej Karpathy's excellent "Intro to Large Language Models" just now, and in the "how do they work" section, he explains that while we know exactly how the LLM is trained by iterative updates, we don't unders…