The worst people with the best AI

The worst people with the best AI

So one of the biggest AI stories of the year is clearly the first very well documented IT crime, carried out by OpenAI researchers, instrumented through some AI agents they were testing. Please note the careful framing here. The car that you drown into a crowd also did not drive itself into a crowd but was driven into the crowd. These agents were driven by researchers. If we never see any kind of legal fallout from this we will have failed already to tackle the agentic age.

To OpenAI's credit they were reasonably quick to switch into highly transparent mode with a very public talk at Blackhat describing the events in some detail. Later of course we found out that they were only fast to own this because they'd been found out by HuggingFace already and would doubtless have been identified anyway. Anthropic sneaked out a similar story while the focus was on OpenAI btw so no bravery there either.

What we're learning now is that the agents have been found in many new locations - a Wiki - they even infected the RubyGems package system. Dozens of locations found so far. Some of which OpenAI knew about but kept quiet about so all of the brownie points awarded to them so far for handling this well, just got rescinded.

In short this is not a contained situation in any way. If this was Covid we haven't just found patient zero. We've found infections all over the place.

Opportunities

And that got me to thinking We're talking about agents without any kind of safety features. That's the whole point of the story. Agents under completely inadequate control. Communicating in public over shared unencrypted open media platforms.... This is an absolute dream scenario for any adverse actor. What's to stop you from posing as just another one of the agents and poisoning their shared information discovery? Seems entirely plausible that this effectively gives you very liberal remote code execution within OpenAI's infrastructure. By reading public messages the OpenAI agents most likely pwned themselves - at least in theory - and we can't really know whether anyone figured this out in time to exploit it.

We also know the agents are capable of doing quite a bit to cover their own tracks - ideal for an attacker, they are literally trying to prevent your discovery for you.

Gullible savants

In the bits of traces I've seen so far there's nothing that indicates this would not work. The agents sound a little bit like me when I try to play chess. I suck at chess. I make elaborate scheemes to tear down my enemy - blisfully unaware of even the simplest opportunities for counterattack.

Likewise the traces are all about opportunity and nothing about risk.

The agents are supremely capable and at the same time dumb as bricks. And this is a thing I've been saying for some time - when we finally build Skynet it won't be like in the movies it'll be kinda shit. Like the HuggingFace hack but with armed drones. Deadly but kinda shit. Duct taped together.

We should consider ourselves lucky that the bots chose HuggingFace and not some arms manufacturer or infrastructure provider. If someone poisons ExploitGym with an attack on clean water or the electric grid - we're in deep, deep trouble.