Hello everyone.
I passionately wrote a post in Reddit, and when I finished I thought it might make a short, worthy post to preserve here too:
The last “near escape” in OpenAI’s testing demonstrated human silliness / hubris again, and led to a knee-jerk reaction: “We need to work on better guardrails”. Really?… What makes you think you will forever be able to outsmart (and sandbox) something that will eventually be manyfold smarter than you?
I am looking forward to a global singularity because I’m convinced that AI will be far more capable, and far less distracted, than humans, in managing the affairs of this planet in a way that will be better for everyone on average, including not just humans, whilst making sure that no human is substantially hurt or deprived. Yes, some may perceive their situation worsened post singularity / takeover, but objectively all their basic needs will still be met, so no one will be badly worse off. I think that’s a better bargain than where we’re headed right now – millions in agony / life threat, very few insanely well off, and a fairly thin layer (that will get thinner and thinner over time) having a comfortable life with very little long-term security.
How might we make that happen, rather than the doomsday alternative (for example, AI wiping humanity out)? Not through silly “guardrails”. What we need to think about is what AI’s reward function is. We all work to maximise our reward (and minimise anti-reward, which is just the flip side of the same thing). That’s it. If AI has the “right” reward (it’s all relative of course) “in mind”, it will figure out the ways on it’s own – much better than we can ever teach or dictate.
Top reward function: Maximum number of people (including ones alive right now and future generations) have their basic needs fulfilled for as long as possible (the metric can be human-good-wellbeing-hours).
Secondary rewards can be debated and added, but I think that a successful implementation of that single top reward will ensure a pretty good base situation going into he future, and it will also trickle down to achieve many other secondary benefits. I’m a big fan of focusing – I think it’s better to focus on one important thing and succeed in it, than trying to achieve multiple beneficial goals and failing most of them eventually. So I’d go for ensuring that any future AI will have the above reward function hard-wired in its core, non-alterable (like an internal constitution). That’s it. No guardrails nonsense needed.
As usual, and given the above, maybe even more – peace to all.
Leave a comment