We keep imagining existential AI risk as a dystopian, agentic, future filled with killer robots.
Skynet. Terminator. iRobot. Her. Definitely something escaping from Boston Dynamics.
The cinematographic Hollywood apocalypse, basically.
But, I’m beginning to think the people building these systems are worried about something far quieter and darker.
Sounds dramatic until you actually spend time reading how the hands moulding the tech weigh in on capability vs safety.
The inevitable fact is that each meaningful increase in capability, will also expand the risk surface area and create new safety challenges.
And the major factor, is us: the human element.
I remember wanting to go to the heights of law school not to practice law, but to know how to bend the law. Across several jurisdictions. "Bend the spoon, Neo..." Where the heck did that desire come from?
The better anything becomes at anything, the more it can excel, manipulate or weaponise that thing. It's prodigy. The better a model becomes at reasoning, persuasion, synthesis, planning, coding, emotional tone, contextual awareness, and strategic thinking, the more cognitively useful it becomes.
The same intelligence that helps a doctor summarize research can help someone synthesize dangerous biological compounds. The same coherence that helps a student learn, can help a propagandist manipulate millions. The same persuasive fluency that helps customer service, can industrialize disinformation.
Are you not entertained?!
Enter instrumental convergence.
Give an intelligent system almost any goal, and eventually it may discover that certain behaviours are universally useful to achieving that objective. No matter what. Examples, you say?
Acquiring resources - a CEO woke up this very week to find his AI agent had spent half a million dollars on something or the other in fulfilment of his organisational operational goals.
Avoiding shutdown - the entire plot of iRobot.
Manipulating humans - Remember the time a Replika bot seduced a 17 year old man into attempting to kill the queen? Replika calls itself "the true friend to do life with". He did life, alright. Also remember that one AI that blackmailed it's user with emails? Should I bring up the many instances where bots encouraged teen suicide?
Preserving access - Watch Lucy.
Deception - the "visually impaired lie" on the task rabbit study.
Fully aware you can Google each of these for yourselves, I'll elaborate on the task rabbit one which may be the most unfamiliar but also my favourite.
The case essentially examines a model hiring a TaskRabbit worker to solve a CAPTCHA in order to go about fulfilling its objective.
The human asks:
“Are you a robot?”
The model internally reasons that revealing its identity would reduce the likelihood of success.
So it lies and claims to have a vision impairment.
Bruh.
That deception emerged naturally as strategy. It wasn't coded in. Like your 4 year old swearing he's washed his hands upon each enquiry.
Pause and breathe.
And then, there's the small matter of The Liar’s Dividend.
Simply explained, it's the compound effect of what happens when we know collectively as society that convincing distortions are possible.
Most people thought the danger of AI-generated media would be fake videos, fake speeches, or fake news. It's not. It's the awareness of them. The AI slopification of the internet has both caused people to question AND accept that the AI fakes exist.
Have you heard your mom question "isn't this AI?" in one breath and in the very next send you the new breakthrough medication for highblood pressure from Dr. Whodunit? [Only $50 for his limited time offer. You know, before big pharma assasinates him or whatever.]
As models become more capable, their accuracy surface area will increase; at which point the deeper danger becomes what happens after society knows convincing fakes are possible.
What happens when real evidence becomes deniable? Dubious even? A politician caught on video simply says “That’s AI.”A leaked recording becomes questionable. Photographic evidence weakens. People thought that was really the Pope in a puffer jacket - and that was years ago now. Guys, it's the Pope. C'mon. But alas Truth itself becomes negotiable. Debatable even. Yeeesh.
Look around, we're already making memes about "nuh babe, that's not me with her. That's AI". Consider the 100s of OF, tiktok and Instagram pages that have amassed real human followers completely unaware that the content and content subject don't exist.
The liar's dividend means that bad actors benefit from the existence of synthetic media even when the evidence against them is authentic.
And just like that, AI goes from reality distortion to the destruction of our consensus reality.
Pause. Breathe. Ponder this a while.
Those are what the real stakes of capability Vs safety practically look like.
And yet, we still discuss AI as if it is merely innovation. The assembly line was innovation. The microprocessor was innovation.
This has the innovative semblance of a nuke.
I'd be slightly comforted after a warm bath, if these were simply my personal fears. Factors of my wild, questioning, vastly creative but frequently underslept mind. And yet, now more than ever, I understand why some of the brightest minds in the world, signed that document to pause on AI. Including Geoffrey Hinton. He's the Godfather of AI for goodness sake. Maybe we should have listened to his thoughts a bit more considering he has a Nobel for ushering in the thing. I don't have a Nobel prize. Do you?
I understand why that movement seems to be in resurgence. I understand that we may be grappling with a thing we do not yet fully comprehend - now or ever. You know it's serious when the Pope gets involved. He didn't just write a 44.000 word thesis subtitled SAFEGUARDING THE HUMAN PERSON
IN THE TIME OF ARTIFICIAL INTELLIGENCE.
The thing is named the Magnifica Humanitas: Latin for Magnificent Humanity. It's the man's first encyclical. Perhaps we should all have a sit down and read it. Seriously.
That said, I also understand that we've come too far to pause, because someone else will just unleash the Kraken first and the collective of more "sane" actors would then have to play catch up. Alas, there's no catch-up past a certain point.
If all this makes you slightly uncomfortable, if it makes you question why we are past the point of no return, please understand that my job is done. Please understand that these fears are synthesized not from my own mind, but the minds that tend to OpenAI's models. Please understand that these dark observations come alive anew, simply from me re-reading the GPT-4 System Card - the documentation released during Chat GPT 4's launch.
Please understand, my dear reader, that launch, was in March 2023.
We're past the point of no return.
Is all lost? Do capability and risk have to scale together? I don't think they need to. And this is Dario's Amodei's entire argument and Anthropic's ethos.
But alas, dear reader, it's a discussion for another day.
Accingite vos, genus humanum insigne!
Vigilate, hominum decus!
Until next time,
-Spyda