Rachel Feltman: For Scientific American’s Science Quickly, I’m Rachel Feltman.
Have you heard about the OpenAI-Hugging Face incident? (Sidenote: Does that string of words also make you think of something Timothy Olyphant’s character might get up to on the show Alien: Earth? You know ‘cause he is AI, the aliens hug faces. Anyway.)
The incident in question went down a few months ago, when OpenAI agents messed with an artificial intelligence platform called Hugging Face. This definitely wasn’t supposed to happen, and it freaked a lot of people out. But what actually happened? According to many headlines and lots of online chatter, more than a thousand “rogue” AI agents collaborated to conduct an unsanctioned attack. While all of that could technically be considered true, some of the words I just used imply a sort of intentionality and dare I say impishness that simply wasn’t involved. The truth is a lot less science-fiction-y: The human beings who designed these models and gave them their prompts made some missteps. But while that might sound less apocalyptic than the headlines did, it doesn’t mean we shouldn’t find this incident deeply troubling.
On supporting science journalism
If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
My guest today is an artist and AI researcher who’s thought a lot about the way we talk about this and other AI mishaps. His name is Eryk Salvaggio, and he’s a Gates scholar at the University of Cambridge working at Cambridge Digital Humanities.
Thank you so much for coming on to chat with us today.
Eryk Salvaggio: Thank you for having me.
Feltman: So I brought you on because I read a piece of yours about AI and how we are talking about it, and I’m really excited to get into that with you. But to give a little context for our listeners, could you tell us a little bit about the work you do and your professional relationship with AI?
Salvaggio: Yeah, certainly. So I’ve been working with technology and thinking about technology and ethics since I was a teenager. Basically, I’ve worked in and around policy spaces around technology, and I came to it through a kind of cultural background. I worked as an artist, and I was working with this technology to make work, and that exposed me over time to some of the ethical questions, some of the questions around technology and power and the way that technology is changing our social lives, our political lives and culture in general.
So I come at it from that lens, trying to think differently about the stories that technology is telling us about itself, in a way, and trying to think about how I can look at that from a different angle, a different perspective.
Feltman: Well, so the piece I referenced earlier was about the [Hugging Face] incident. Could you just briefly summarize, for our listeners, what happened there and, sort of, what the prevailing narrative around it was? And then we can get into, a little bit, the ways in which you disagree with that narrative.
Salvaggio: Yeah, sure. So the prevailing narrative of what happened comes from OpenAI. They released [an] alert. Basically, this company Hugging Face announced that they had been hacked by an agent of some sort of large language model or AI system had hacked them. And then OpenAI said, “Oh, sorry, that was us. Looks like we did that, and we understand kind of what happened, and we’re gonna do a full report.”
A couple weeks after that, we get this report, and there was, kind of, pandemonium, right? A lot of the way that people understood what happened was filtered through this lens of going rogue. The artificial intelligence had kind of learned to coordinate with other instances of itself and had identified a target and had gone after Hugging Face in order to basically do this kind of weird exam for AI models called ExploitGym. And so it “broke containments”—this is another phrase that we heard in some of the headlines—and got online and went to the Hugging Face website and did what it did there.
And so this created a lot of concern about these AI agents and what they were becoming capable of doing and their coordination and their persistence, right, their perseverance. They weren’t giving up until they got in and this kind of stuff. And so you really had a lot of, I think, fear, a lot of mystery, and a lot of, kind of, talking about the agents and the AI system as if it was its own thing that acted on its own volition.
And sometimes I refer to this as, like, a system from nowhere. Like, no one built this thing. No one knows how it works. It just happens to everybody. And I take some umbrage to that frame. I think there is something to look at and say that there’s some real accountability that we can see when we look closer at this incident.
Feltman: Yeah. Well, and I think for a lot of people, this idea that the agents were communicating with each other on a message board, and it really evoked the idea of these distinct individuals, you know, collaborating—which, correct me if I’m wrong, but that’s not really what we’re talking about with an AI model, right?
Salvaggio: Right. What we’re talking about these days is something called an agentic system, which really is multiple versions of the same model, for the most part, being spun up and run with a kind of subroutine, you could call it, or a side task of some kind. And so ultimately—and it’s interesting, Anthropic had this interesting diagram where they showed Claude pointing an arrow at another box that said Claude, and that had an arrow pointing to another box that said Claude, and then both of those converged, right? And it was just Claude all the way down. And this is pretty much how they’re running.
And so, when we talk about these agents, what we’re really talking about is a kind of narrow slice of a bigger model, but it’s all optimized toward the same goals. It is trained on the same training data. It is basically the same model being run multiple times. Now, with the Hugging Face incident, OpenAI had two models, and they were kind of collaborating, but most of the stuff that we’re talking about actually happened within the single model, which is their internal model that only they have access to right now.
Feltman: Well, and one thing I really appreciated about the way you wrote about this is that you weren’t denying that there is reason for concern here but rather that we might be missing the actual areas of concern by, you know, focusing on this sort of like, Skynet scenario. Could you tell us a little bit more about, you know, what you think we should actually be learning from this incident?
Salvaggio: So a lot of the way this has been talked about has been using intentional language, right? This is this idea that the model wanted something, the model believed something, the model was trying to do something. And this can be a really, kind of, straightforward way of getting the basics down of what exactly happened.
But if you stop there, you basically kind of say that the model did all of this on its own, and you stop looking at the human decisions about how this model was designed and how it was deployed, how the testing environment was built and deployed, right? So we know, for example, OpenAI says this is a persistent model. It didn’t give up. It kept trying. And what that, kind of, moves you away from is asking the question, “Well, why?” And there’s a real answer to that, and it’s a human decision about how we train these models. And what it turns out: they trained it not to stop. Most of the time, these language models have something called a stop token. It comes up, and it looks like it’s the end of something that someone would say, and so it finishes. This didn’t really have that. It was designed to keep trying. If it failed, it would just, kind of, step back and repivot and try and try and try again. So there’s one aspect of this that, I think, is that we lose sight of when we look at just the model’s behavior and instead start saying, “Well, how was it designed?” and we’re constantly optimizing models, training models. We’re involved in making decisions about pre-training, what we do to the data, in other words, what we do to the models once they’re trained, and what we kinda ask them to do.
All of this stuff is human decisions happening inside the organizations, and we can ask for accountability by looking at those decisions. Whereas if we look only at what the model did and attribute all of this to its desires, or wants, we kind of lose sight of exactly how those desires and wants came to be.
Feltman: Yeah, and it seems like a company that’s creating this kind of AI product would have a lot of incentive to make people believe that it’s becoming super smart and autonomous and not a lot of incentive to admit that they need better guardrails internally.
Salvaggio: Yeah. A lot of this language is doing that work of saying, “We actually don’t understand this, and we don’t have control over it.” And when you are a company that is building models, and you’re saying that we’re not in control of what we’re building, there’s something going on there, and I think we can ask some real questions about what purpose that serves.
And they’re constantly trying to share this story of competency and almost overwhelming competency, right? And of course they are, because they’re asking people to use these models in situations that really we want to be able to trust them. We want to be able to ask it for information and get information back that we trust. And so they’re constantly trying to say, “Well, these models are intelligent.” And I actually think that’s the wrong frame. I think intelligence is, kind of, this idea, this kind of name that we gave to this category of computing. And really what we’re looking at is: We have no idea what’s going on when we say that something is intelligent or not. This is a well-known kind of philosophical problem. But what we do know is that they produce language, and these models are really interesting because even if it’s opening up your e-mail application, even if it’s running a piece of code, it’s doing all of this through language.
And so the way I approach it is: What exactly is this language doing? And how can we look through the language that the model’s producing, how it generates that language? And all of that is human optimization, it’s someone’s decision about the type of language these models are gonna produce, how it gets there, what it’s rewarded for producing and saying and doing. So I would like to look at more—less at thinking about these as artificial intelligence and try to focus it at more on artificial language, the idea that this language is doing something to us to influence us but also to act in the world. So how do we understand that, and how do we understand the—how these are being steered through this language?
Feltman: Could you tell us a little bit more about that, you know, in terms of what you would like to see from these technology companies in the future to steward good AI?
Salvaggio: I would say probably the most important thing for thinking about the way this particular technology is developed is really to think through where accountability can be found. When we are talking about models as if they are happening to the companies that are building them, and we hear this a lot in the rhetoric, we hear that they are grown, not built, which really kind of only refers to, sort of, the scaling of the training data. Getting more and more training data makes them bigger. But there is a building side of this, too, which is to say that they are making decisions about what exactly these things are doing. Whether they stop was a classic example of OpenAI and this incident, right? Whether they coordinate, that’s a human decision.
When you have so much that is unpredictable—and so, as you say, I’m not dismissing that there are real concerns about the unpredictability here. But when you know that something is unpredictable, I think we have, still, a responsibility to manage the kind of scope and the boundaries of where that unpredictability might lead.
And so what I wanna see is more accountability, saying, “Well, if you design a model that does something that you should’ve anticipated that it could be unpredictable,” and I know from my experience that these models are unpredictable. I work with very tiny versions of them, all the time, and they’re doing weird stuff, and I know that, so clearly these companies know that, and they can build better safeguards, guardrails, whatever you wanna call them, to make sure that unpredictability doesn’t have real-world consequences.
Feltman: Absolutely. Well, thank you so much for coming on to chat about this. It’s been super interesting.
Salvaggio: My pleasure.
Feltman: That’s all for today’s episode. We’ll be back on Friday to talk to a doctor who’s developed and tested a new pain-relief method for an infamously uncomfy procedure: IUD insertion.
Science Quickly is produced by me, Rachel Feltman, along with Fonda Mwangi, Allison Rodgers and Jeff DelViscio. This episode was edited by Alex Sugiura. Marielle Issa and Aaron Shattuck fact-check our show. Our theme music was composed by Dominic Smith. Subscribe to Scientific American for more up-to-date and in-depth science news.
For Scientific American, this is Rachel Feltman. See you next time!
