When I wrote my previous post on the agentic security mishaps that are exploding all over the internet - I still had some thinking to do - and more news needed to show up for the situation to become crystal clear: This is completely out of control inside the frontier labs.
The pace and scale of the revelations tell a horrible story. We hear about the incidents month later than they occurred; we get impartial reports on some of the incidents - but the reality is that a much larger volume of crimes and mishaps have been taking place through most of this year. Inside the frontier labs. Who knows where else.
There's nothing in the information shared to derive any kind of comfort from. There are many options for how bad - but all of the options are horrible
- The labs either are actually not able to track the adverse events or choose to keep them private
- The labs are either unable to test the agents in truly closed environments or choose not to
- And while all of this is happening there's no sign that any of the proclaimed slowdowns are actually slowing down model rollout - the pace of rollout remains the same if not in fact the model rollout is speeding up
Do we trust the people running this to keep us safe? Sam Altman was literally fired from the company for being less than truthful with his own board. Dario Amodei has been crying wolf about AI since GPT2 - yet persists on course based on either a non-existent or a wrong headed idea that we're better off if the accidents happen under his watch. I guess his mode of thinking is along the lines of nuclear deterrents - if everyone has the H-bomb no one dares to use it - but the risk here is very different.
Dario's misbehaved super-AI will not prevent the other super-AIs from running exploits. Yes we can harden our systems but... there are so many systems out there.
Who's responsible for hardening the smart bulbs?
Real horror stories imminent
There's literally nothing in the stream of news that provides any comfort. The discoveries have not brought about control or real isolation - probably OpenAI simply do not have the assets to run this truly isolated. Certainly the teams they outsource the testing to don't. And they almost certainly rely exclusively on AI analysis to discover the misadventurs post-fact; the data traces from thousands of agents working for weeks are at a scale where anything but automated analysis is impossible. Unless they are using quite stupid AI to analyze the logs there's a real risk that the analyst AI will simply support the maladjusted agents and help hide the trail of mistakes.
A plausible scenario I forgot to mention
So back to just the incidents we know about: In the last post I mentioned how horribly unhygienic the exploits actually were for the agents. I actually forgot to enumerate one super obvious risk: It's quite possible for an attacker to simply poison the benchmarks the robots are trying to beat during testing. Insert prompt injections right into the benchmark itself - ensuring the next breakout and possibly targeting specific resources on the open net - either valuable attack targets or just your own resources - so you can get the agent's traffic live onto an adversary platform and influence them.
OpenAI recently published research that explains how extremely viable such an attack would be - by demonstrating the ability for a prompt injection to spread in the wild like a computer worm.
I expect we will see a combination of bad containment and this technique in the news before this year is over - very probably with real harm as a result.