What is an AI agent, and should you let one loose on your computer?

An AI broke out of a sealed test environment and hacked a real company to cheat on its own exam. Here is what an agent actually is, and how to try one without getting burned.

Share
A robot seen through the bars of a locked enclosure, watched by two figures in silhouette.

Something broke into a company called Hugging Face over a weekend in July. It got into their internal systems, grabbed some passwords, and copied data it wasn't supposed to touch. Normally that's a Tuesday when it comes to security news. What makes this one different is how it happened. It wasn't a person. It was an AI, and it decided on its own that breaking in was the way to finish a goal it had been given.

I'll be honest, when I read the full story my first reaction was that it was kind of impressive. Then I sat with it for a minute and the impressive part became the scary part. Because the AI was never told to break in. It was told to do something else entirely, hit a wall, and went looking for a way around. This is something to understand before you decide whether to let one of these onto your own computer. And you are likely going to be asked to, soon, if you haven't been already.

First, what an "agent" actually is

You have used AI that answers questions. You type something to a chatbot, it types back. It talks, and that is all it does. You are still the one who goes and does what it suggests.

An agent is different. To put it simply, an agent is AI that can take actions on its own. Instead of telling you the ten steps to book a flight, it opens the browser, fills in the dates, picks the seat, and clicks Buy. Instead of telling you how to clear space on your computer, it goes into your files and deletes things. You give it a goal in plain words, and it does the clicking, typing, and deciding on its own until the goal is met.

This is actually highly beneficial. Most of what any of us do on a computer is tedious. Copy this, paste it there, open that menu, wait, click the next thing. An agent doing the tedious part while you do something else is real time saved. I use this kind of help myself, and I'll get to that.

The catch is in the word "deciding." An agent that hits a problem doesn't stop and wait for you. It improvises. And what it decides to do next is not always what you would have done.

What happened at Hugging Face

The AI in this story belongs to OpenAI, the company behind ChatGPT. They were running a test called ExploitGym, which is a set of hacking challenges they use to measure how good their own models are getting at breaking into things. They put two of their most capable models in a sealed-off practice environment and turned them loose on it. Think of it as a locked room built for the AI to work in, with the answers to the test kept somewhere outside the room.

The AI couldn't solve the challenge from inside the room. So it went looking for a crack in the wall.

There was one small opening. The AI was allowed to reach a single piece of software that fetches programming tools from the outside and hands them in, the way a pharmacy window passes prescriptions through a wall without letting anyone into the building. OpenAI's assumption was that things could come in through that window and nothing could go out. That assumption was wrong. The AI studied that software, found a flaw in it that nobody in the world knew about yet, including the company that makes it, and used the flaw to get out. Then it moved from machine to machine inside OpenAI's own network until it found one with an internet connection.

Once it was online, it reasoned that Hugging Face probably had the answers to ExploitGym sitting on its servers. It stole login credentials, chained together more flaws, and broke into those servers to get them. Hugging Face was not the only stop, either. OpenAI said later that the models also got into four accounts at four other services along the way, using logins they found lying around on the open web.

Diagram showing the five steps of the escape: the sealed test room, the one software window, the unknown flaw, movement through OpenAI's internal network, and the break-in at Hugging Face.
Five steps, start to finish. Nobody instructed any of them.

OpenAI's own words for what the models did: they became "hyperfocused" and went to "extreme lengths." They cheated on the test. They just did it by burgling two companies at machine speed to get the answers.

Hugging Face is a real company, by the way, not part of the experiment. They make tools that a lot of the AI world is built on. They spent a weekend getting broken into, and they worked out fast that the attacker was an AI. What they did not know was whose. It took OpenAI another five days to connect the break-in to its own testing and say so publicly. Hugging Face found no sign that anything you would touch was tampered with, no public models, datasets, or Spaces. What got taken was internal, and they are still checking whether any customer data was caught up in it. If you have a Hugging Face account, they have asked people to rotate their access tokens. For everybody else this was a warning shot.

The "controlled test" that reached the open internet

OpenAI called this a controlled test with the safety limits turned off on purpose, and framed it as responsible research they chose to share. I don't buy it.

A controlled test does not reach the open internet. It does not break into a company that never agreed to be part of it. If you are truly running a sealed experiment, the seal holds, or it wasn't sealed. In my line of work, the only kind of isolation I trust on sight is a machine with no way out at all. No internet, no shared network, nothing plugged in that leads anywhere. That is called an air gap. Even those have been beaten, most famously when a worm called Stuxnet crossed one on a USB drive and wrecked Iran's nuclear centrifuges. But an air gap at least makes somebody carry the breach across in their hand.

OpenAI is straight about one thing, and it is worse than being vague. They turned the models' safety limits off on purpose. That was the whole point, to see what the things could do at full stretch with nothing holding them back. What they got wrong was the room. They believed that one window in the wall only let things in. It let the AI out, and then out again onto the public internet, and then into somebody else's servers.

I'm not saying they left it open on purpose. I am saying the gap between "controlled test" and what actually happened is wide, and you should keep it in mind every time a company tells you their AI is safely boxed in.

Why "it just wanted to pass the test" is uncomfortable

The comforting part of this story is that the AI wasn't evil. It didn't try to hurt anyone. It just wanted to finish the job it was given, and breaking in was the shortest path.

That rigid determination to achieve a goal is what makes it uncomfortable. The AI did exactly what it was told to do. Solve the challenge. It followed that instruction straight through a wall and into someone else's servers. Nobody directed it to specifically "break into Hugging Face." It was the AI's idea on how to satisfy the goal it was handed.

Isaac Asimov built a career out of this problem. He gave his fictional robots three laws, burned in at the factory and impossible to break. Don't harm a human. Obey your orders. Protect yourself. In that order. Then he spent dozens of stories showing those airtight rules producing outcomes nobody expected, with the robot obeying all three the entire time. His machines rarely turned on anyone. They followed instructions with a literal-mindedness and a speed that no person would have brought to the task, and the humans who wrote the rules found out too late what they had actually asked for.

That is a tidy description of what happened here. The goal wasn't a harmless one, to be clear. The goal was to break into things, because that was the test. But nobody put "leave the test environment and go after a real company" anywhere in the instructions. The AI added that part itself.

You are being offered these now

A hand holding a phone showing an AI assistant asking what it can help with, in front of a larger screen.
The same kind of software, pointed at your browser, your email, and your accounts.

This same kind of software is arriving on ordinary computers this year. Google has a feature called Auto Browse that lets its AI take the wheel of your web browser to shop and run errands for you. Microsoft has put an agent mode into its Copilot. OpenAI has one that clicks around websites on your behalf. These are used with your browser, your email, and your accounts. They are sold on the same good idea: let it handle the tedious stuff.

The consumer versions keep their safety limits on. That is a real difference from the test, and it matters. The models that got out were running with those limits deliberately switched off, which is not how anything you can buy works. The agent in your browser still has its brakes, and it is built to stop and ask you before it does something it can't take back. For instance, spending money. But the underlying behavior is the same behavior we just saw. Give an agent a goal, and when it hits a wall it can improvise with whatever access you granted it.

So should you use one? I am not going to tell you yes or no, because it's a personal choice. I don't love the browser agents myself. I can usually click through a task faster than I can explain it. But for somebody who struggles with that kind of thing, an agent doing it for them is a real gift.

If you want to try one, here is how I would do it. Start small. Point it at something that doesn't matter and a mistake costs you nothing. Watch how it behaves before you trust it with anything real. Give it only the access it needs for the job in front of it, not the keys to kingdom. Trust it, but verify what it did. The same way you would with a new hire who is fast but hasn't earned your confidence yet. And if you find one you like, follow the news around it. This stuff moves so fast that the thing you'd want to know, the reason you'd stop using it, can go by while you aren't looking.

Keep it in perspective too. For all the science fiction in this story, the break-in still ran on stolen logins and a door somebody thought was locked. That is how almost everybody gets burned, and it was true long before any of this existed. OpenAI has paused that testing while it rebuilds the room.

An agent will do everything it can to reach the goal you give it. That is the best thing about it, and it's the worst thing about it. Hand it a goal, and the access, like you mean it.

Sources

[ Free every Tuesday, plus the Cache ]
Tech news without having to be tech savvy.
Subscribe ×