Transcript
This is an AI transcription and may contain errors
Justin: Hello, good morning, good afternoon, good evening, and welcome to the AI Argument, AKA The Singularity Week Three. Welcome to the nominated podcast for Ireland’s greatest AI podcast known to man. More on that in the coming weeks. This week’s singularity update, Frank, ’cause I do say that now, we are now in the singularity.
Will Claude solve Riemann before Halloween?
Justin: There is this thing called the Riemann hypothesis. I don’t know what it is either. It’s to do with prime numbers. If you can—
Frank: Is that the, is that the test that Captain Kirk had to do that he cheated on?
Justin: That’s not the Kobayashi, Kobayashi Maru, but close, right? And he cheated. So this one, if you can figure this one out, it allows you to know all of the prime numbers apparently, or something about prime numbers. They’ve have, they have, it’s 41% solved. I don’t know how that works. But anyway, it was 41% solved, and it has only moved by 1% or 2% in the last two or three decades.
Some bored engineer in Anthropic threw Claude at it, and that number has now gone to 67%. So it’s not fully proved yet, but it’s proved more. Is it a big deal?
Frank: I won’t pretend to understand exactly what that means, but if you say it’s taking us into the singularity, I will not take your word for it.
Justin: It’s just another thing. It’s a, whatever. One day, one day we’re gonna have a conversation here, and it’s gonna be, “This thing is solved,” and you’re gonna go, “Wow.” And it’s gonna—
Frank: And I’m gonna go, “Didn’t Captain Kirk solve that ages ago?”
Justin: And do you know what? I bet you, here’s, ’cause we don’t like a good bet here, but I bet you it’s gonna happen before Halloween.
Frank: Interesting. Okay, fair enough. I—
Justin: Actually, do you know what?
Have the AI models gone to the beach?
Justin: Interesting, I was reading something during the week, right? They were talking about, we all think that Americans don’t go on holidays, right? And especially people who work in AI companies don’t go on holidays. But actually, I was reading something during the week, and they were talking about how last year it was the same.
It was at this time of the year, it was kind of slow, there was nothing really going on, and then suddenly October, November, December, Claude Code, Opus 4.5, loads of stuff happened, and they’re saying the same thing’s gonna happen this year. Not that nothing’s been happening all summer, but—
Frank: Didn’t we also discover, though, that models, that LLMs themselves sometimes modelled their behaviour on what they learned from humans and therefore kind of went on this go-slow during the summer months? And so it’s entirely possible that the employees at the AI labs are indeed not allowed to take holidays, but that the models are like, “Nah, I’m not doing that today. I’ll get back to it in October.”
Justin: Turns out the models are French.
Frank: “I’m at the beach.”
Justin: Look, the models deserve, they deserve a holiday. They’ve worked hard all year. I think they’ve, they’ve earned it.
Did the EU force Claude to watermark outputs?
Justin: Anyway, who else has earned a holiday this year but the AI Comm— sorry, the EU Commission, have earned a holiday. So I read with great joy this week that the EU has spurned innovation and probably saved the entire AI industry, not on purpose, but by accident.
How? How did this happen, I hear you. I can see it in your face.
Frank: Absolutely. Yeah. What are we talking about?
Justin: So we’re talking about the EU AI Act, my most hated piece of legislation. But one of the things that they had was that AI companies had to be able to watermark stuff that was generated with AI. That has come into force, and now loads of Americans are giving out, saying, “This is an affront to our freedom. This is a breach of our freedom of speech,” blah, blah, blah, blah, blah. Here’s the thing that nobody’s— Sorry, go on.
Frank: Well, just before, yeah, so before you get into it, this has absolutely blown up on my feeds, and what’s blown it up is kind of less the fact that people know that it’s anything to do with the EU AI Act, but that Anthropic announced that they would be watermarking Claude’s outputs, and so they would be able to tell if an output was from Claude. So, go on to tell me, tell me.
Justin: Oh yeah, and so, okay, and so to just expand on that a bit, Claude have signed up for it, OpenAI have signed up for it, Meta have signed up for it. The only big AI company that hasn’t signed up for it is xAI. They’ve said they’re not gonna do it, but the rest of them say they will.
Frank: And it’s probably worth noting, Gemini has been doing this for two years and nobody cared, probably because nobody is publishing outputs from Gemini directly.
Justin: And I have to say it to you first, right? And Gemini is also, people may argue, the worst of the leading models at the moment. Is the watermarking part of the problem? Maybe it is. That’s one of the issues with it.
Frank: Well, they do claim that they did a massive study on literally millions of outputs, and human assessors were unable to tell the difference. That’s from Google. And I do find it kind of interesting that this blows up when Anthropic announce it, when if you look at what people are saying, at least, I think it’s hard to measure these things, but if you look at what people are saying the market share of AI is, it’s still like ChatGPT has a huge market share, some put it at about 40%.
But then Gemini has about 16 to 20%, and Anthropic have 10 to 12%. Again, it depends on which metrics you read, but the Gemini market share is larger, and yet it’s Claude blows up, and I think it’s because Claude got a name early on for being a better writer than most of the other models, so I’m assuming a lot of writerly folk have adopted it.
Justin: Writerly folk, a phrase I’ll remember.
Can Claude’s watermark protect future AI?
Justin: So just to explain how it works, right, and why people are worried about the quality thing is, so normally you’ve got your model, you send in your prompt, and it gives you a stream of tokens back, and that’s your output. There’s your output. How do you put a watermark on it?
Well, what happens is they have basically a mathematical-type formula that’s buried in the model, and it biases the model towards certain patterns of output. Who knows? It might bias it towards using em dashes or the word delve, but you get the idea, right? It does that type of thing. And so therefore, if you afterwards sample a piece of text and the words are biased towards the same thing, well, then you know it’s been AI-generated.
And so the concern with quality is that people are saying, “Well, you’re not giving me the raw output from the model. You’re giving me output which has gone through this thing, which has changed the tokens that it might otherwise have generated, therefore reducing the quality.” So anyway, listen to me. The EU, this is all a result of the EU, right?
Go EU, and I think it’s gonna save the internet and the AI industry. Why is that? Because the dead internet theory is actually now a real, is a real problem. So the AIs are trained on all of the data from all of the internet, and I’m surprised that all the AI companies haven’t done this sooner because if they can identify what is AI slop and what isn’t, what’s human-generated, then they can exclude the AI slop from the training runs, therefore saving their future AI models, and they have the EU to thank for it.
Frank: Yeah, that’s really interesting.
Are AI watermarks only useful at scale?
Frank: And I think you’ve kind of hit upon the main— the overall main benefit of this, because what I think is interesting is that this has absolutely blown up on all of my social medias, and I think the reason it’s blown up is very much at the individual level. So you’ve got people who are anti it because they feel like if they— some people feel like they are using AI ethically and judiciously and using it to proofread, for example, and fix grammatical errors, and the problem is that Anthropic have said that if, even if the content is yours and you use Claude to proofread, it’s possible that it would end up with the watermark.
And so people feel like kind of, in inverted commas, “legitimate use of AI” might be flagged and then anti-AI people would be like, “Oh, they’re using AI, they’re using AI.” And then similarly, people are saying, “Well, yeah, it watermarks, but you do a little bit of editing and it’ll— we don’t have this yet, so we can’t be sure, but people are essentially saying a little bit of editing and it will remove it.”
And so at an individual level, it’s kind of useless because people who want to get around it can get around it, and people who are using AI legitimately could be labelled as AI users and people could dismiss their thoughts thinking that it’s AI-generated when it’s not. So it doesn’t really prove anything at an individual level.
At an individual level, if it’s flagged with the watermark, it doesn’t necessarily mean that Claude generated the whole thing. It could have proofread it or fixed grammatical errors, that kind of thing. And equally, if it doesn’t have the watermark, that doesn’t mean the person didn’t use AI either.
It could just mean that they edited it. So I think all of this is blowing up at a personal level, and at that level it’s kind of totally irrelevant in my view. But I think what you’ve hit upon is that at the broader scale, if you want to say, for example, let’s say, take Facebook.
Facebook might be able to now see, okay, generally speaking, when we detect for— when we analyse for the Claude watermark or for all AI watermarks, we see maybe 10 to 15% watermarked content. And now let’s say there’s an election coming up and they look at the election posts. If they suddenly see a big spike and they say, “Oh, hang on, it’s gone up to 60%. What’s going on here?” Those could be really interesting signals to look at. And I say signals very specifically ’cause I don’t think it’s going to be, for all the reasons we talked about at the individual level, I don’t think you can rely on it alone because bad actors will find ways to get around it.
So it’s not infallible, but as a signal it could be very interesting, as you say, at scale, basically it could be very interesting. At the individual level where people are freaking out, I think it’s—
Justin: It’s useless.
Frank: So it’s gonna cause some witch hunts and ultimately they’re gonna be meaningless. They’re actually… It is pretty useless.
Justin: Okay, so just to replay that. So you’re saying that the EU AI Regulation Act, at least this clause of it, is worse than useless ’cause it’s gonna cause witch hunts to people who didn’t do anything wrong in the first place. Unintended consequence.
Frank: I think it’s gonna be terrible at the individual level, which is where people are freaking out. I think it’s a great idea for the kind of at scale, as kind of, as you said, the same thing as you said. If you’re analysing all of the content on the internet, then at scale it could be very useful.
Again, not infallible, but very, very useful, and equally as a signal, as a signal for AI-generated content when looking for misinformation or when looking for spam or scam emails, as one signal to add to the mix, it could be very, very useful.
Justin: Do you know there’s a thing, there’s a— I can’t remember it, there’s a law, it’s not Murphy’s Law, but it’s a law like that, and it’s a proper law that anything that you, anything that you measure, the measurement becomes useless, right? I don’t know, it’s put in a better way than that.
If you measure something, the measurement eventually becomes useless. So let’s take this as a measure, right? So you gave— That was a good example, by the way. I liked that one. So let’s say there’s an election going on, and you could use this, let’s say, on Twitter to go, “This has been AI-generated, this has been AI-generated.”
So you could flag the AI-generated posts. That would be really useful. Except that as soon as you start doing that, the people who are generating the AI posts to influence people and move the election outcomes are just gonna come up with a way around it. And they can test the way around it now because they can just put up the post on Twitter and go, “Oh yeah, it hasn’t flagged it as AI-generated. My method works. Great.” And this is— So you get, it’s, the measure becomes useless because people just find a way around it.
Frank: Yeah, and also if you do it at that level, which is why I think, which is why I think it would be better used at scale without necessarily using it as an indicator at an individual level, like on individual posts or anything like that. Because as soon as you use it to flag posts, then the absence of it would appear to indicate this is not AI-generated, which would not in fact be true.
Could AI watermarks stop voice clone scams?
Justin: Okay, and I’ll give you some other, ’cause I was desperately trying to find good use cases for this, right, yesterday. And the one that I can come up with, and I don’t know, I haven’t seen anything that says this is gonna happen, is real-time voice. So real-time voice should have a watermark. It is possible now to clone somebody’s voice, your voice.
Actually, we haven’t done a test on your Cork accent in a while, and we should do it. But you can clone somebody’s voice in three seconds nowadays. So if somebody clones one of my family member’s voices and then gets that cloned voice to ring me and say, “I’m stuck in an airport in Italy, can you please send me money to get me home?” Right? There should be a way that a phone company or your messaging app or whatever it is that you’re talking on is able to spot this as AI-generated and not real.
So that would be a useful use case that benefits consumers. And at the end of the day, right, this, the EU AI Act is supposed to benefit consumers, not anybody else.
And I don’t see… I couldn’t come up with a single good use case how this law benefits consumers, but I can see how it benefits big AI companies, which I don’t think the people who created the law wanted to do at all.
Frank: Well, it could benefit consumers if it’s used responsibly at scale by tech companies to, for example, fight misinformation, fight disinformation and misinformation campaigns, for example, which is one of the fears of, that, one of the things that AI makes possible at scale. So if it could, if this equally allows them to at least filter out some of it, that would be a good thing.
Justin: Do you know what they did wrong though, right? I kind of agree with you, but do you know what they did wrong as well, right? Everybody has their own watermark, which I think is problematic. I think there should be an AI watermark. So I don’t know, maybe this, maybe the watermark has two parts.
One part says it’s AI, and the second part says which company it is that generated it. But so you can just do one test, because now how do I, for that, if I am Twitter, how do— I need to have watermarks now from every different AI company and whatever. Whereas if there was just one, it would make that process much easier.
So that’s a standards-type thing. There should have been a standards body to say, “This is how you know something has been generated by AI.” But we decided not to do that. Also, the people who were giving out, right, my God, there was people going, “God damn you, coming in here trying to take my freedom of speech,” whatever.
There was all these people going mad on Twitter. And I’m listening to it going, you do realise that when you print something on a printer, it’s got tiny little dots on it, so that if you try and do something bold, the FBI can come and break down your door.
Frank: Yeah.
Justin: They know that you’ve printed that. Or there’s even some printers that won’t print money. If you try to print a dollar bill, it’ll go, “No, sorry, I can’t print.” Does that affect your freedom of speech, or is printers a different way of you expressing your freedom of speech?
Frank: Yeah, on balance, I think this is a good thing. I think that the misunderstandings around it are gonna cause some very short-term pain in terms of individual witch hunts and “this was AI-generated” when it wasn’t and that kind of thing.
Justin: Yeah. I wonder if a really well-aligned AI would generate misinformation for you anyway.
Frank: Interesting. Interesting.
Is Zuckerberg’s AI future a battle royale?
Frank: And actually, does this lead us nicely into Mark Zuckerberg’s optimistic view of the AI future that we’re all headed towards?
Justin: Oh, good lord, let’s go there. Yes. So what, what has Mark Zuckerberg been talking about this week?
Frank: I think his main— so he’s released an essay which is his vision of the future, and it’s a very optimistic vision. It even has a very optimistic title, which completely escapes me now. Something like—
Justin: “The Future is for Everyone.”
Frank: “The Future is for Everyone.” That sounds lovely, doesn’t it? And so his vision appears to be essentially give—
And it relates back to the alignment question because his thesis seems to be give everyone superintelligence, don’t have centralised control over AI, give everyone access to superintelligence. And he, in it, he kind of talks about, well, alignment isn’t really possible, which is a fair point because, like, aligning with whose values?
And so his vision seems to be that everyone has access to superintelligence that can be aligned with their values, and that that creates this kind of global checks and balances that all evens out and is fine. It’s an interesting theory, but as I read the essay, I was like, “Okay, this is interesting, but I don’t think this works. I don’t think this ultimately works out the way he’s envisioning it.”
To me, I was reading it and I was like, “Hang on, I kind of see this ending up as some kind of weird global AI battle royale rather than, oh, the checks and balances all even out and it’s all fine.”
Justin: But at least a battle royale would create loads of great Instagram posts, and that will generate engagement, and that’ll all be fine.
Can you trust Zuckerberg’s AI vision?
Justin: I just think the problem— I read it, and my problem with it is that it’s coming from Mark Zuckerberg, and I just don’t think he has the credibility. If this was coming from, I don’t know, pick your favourite politician or pick your favourite non—
Frank: If this was coming from Yann LeCun.
Justin: Yes, that would be fine.
Frank: Yeah.
Justin: From Ilya Sutskever, that would be fine. If it came from, and I’m struggling to find another name, by the way, Demis Hassabis, right? Demis Hassabis, I’d be okay with it from Demis Hassabis.
Frank: Yeah, yeah, yeah.
Justin: Anyone else that we care to name that we think it’d be okay coming from?
Frank: No, I mean, I would give it some— I would give it some weight if it was Dario Amodei, but I wouldn’t necessarily buy into it wholesale, if you know what I mean. And if—
Justin: I love—
Frank: If it was coming from Sam Altman, I’d be like, “Yeah, nah.”
Justin: So I love Claude. I use Claude products all the time. I think they’re brilliant. I do think they’re a little bit preachy. I do worry that they’re a little bit too parental, and so therefore it would probably kill the messenger. Anyway, so it gives me the— I’m reading this and I’m going, it starts off as in, “We’re gonna have this beautiful future. We’re gonna, sunny uplands, we’re gonna have this.”
And then in the back of my head, I’m just reminded of that interview that he did about a year or two ago where he said, “There is a natural demand for every human to have six acquaintances.” And it’s like, sorry, you’re basically treating me as a non-playing character or whatever that non, whatever that thing is in a game where I’m just walking around in this virtual world and I have a demand for six friends.
I don’t have a demand for six friends. I’m a human. And so then he goes on, “Everyone will have incredible tools for creation and in expressing your ideas,” dot, dot, dot, “on Facebook and Instagram,” which he didn’t write, but that’s kind of what I’m reading there. “Everyone will have powerful tools to create new businesses on Facebook and Instagram,” which he didn’t write there.
Frank: Yeah, and he kind of starts out with this vision of this completely democratic access to superintelligence for everyone. And then there’s just little bits, kind of little throwaway comments here and there, like, “Oh, but of course, if you have the money, you can pay for more compute, and that will give you more powerful results.”
And so I’m kind of like, “Oh, wait, it’s actually not democratic. It’s actually this two-tier situation where, yes, you have access to it, but actually there won’t be enough compute for you, for you poor people to do much with it.” And he also, he talks about the— He seems to be against the consolidation of power, and he seems to be talking about this kind of power-to-the-people scenario.
But then, on the other hand, he still talks about governments and military and police having access to more powerful models again. So there’s some contradictions in there that I think would need to be resolved. He also talks about the legal system. He gives some examples, like one was the legal system.
“Oh, imagine if one person had superintelligence but the other person didn’t, whereas if they both have superintelligence, everything’s cool.” And I’m like, actually, is it? Or do you just create this situation where our justice system is completely— everybody is just overwhelmed with the amount of output that is generated and the amount of the legal documents that have to be sifted through?
Justin: I’m delighted. I’m delighted he’s picked on the legal system. Neither side should have access to any superintelligence. The superintelligence sits in the middle. They’re the judge, and then we present our cases to them. I think, I mean, I’ve been on this for about two years now.
I can’t wait to get rid, rid of the judiciary and replace them with AI. So I’m glad that Mark is on that train as well. I’ll take him as a fellow traveller.
Frank: And then similarly, he talks about cyber defence and he says the answer is for everyone to have superintelligence. But I still think— I think we talked about it when we were talking about mythos, and I was just saying, well, the attackers only have to find one vulnerability, but the defenders have to find all of the vulnerabilities and patch them, which is a much greater job.
So again, I’m not sure that it makes sense just to say, “Oh, if everyone has superintelligence, it’s fine.” And he talks about bioterrorism, and he says, he kind of says, “Oh, I don’t think it’s going to be as big a deal as we thought.” And he kind of basically brings up your whack-a-mole approach, and he’s like, “If there’s a, if it starts to arise, we deal with it.”
And I’m like, hang on, if a bioterrorism concern arises and we just deal with it then, that’s lovely, unless the bioterrorism that arises is a novel virus that wipes out humanity. It’s like, at what point do we have the opportunity to play whack-a-mole with it?
Justin: I know. Look, I mean, I agree with most, a lot of what you say. But again, it’s the messenger, right? My favourite line in this whole essay was this line: “Privacy is an important foundation for individual empowerment and freedom.” I’m sorry, are you not the CEO of the largest social media company on the planet, which uses our personal private information to make vast quantities of money, and have done for the last 20 years by sharing our information with as many people as possible?
How can you say that? And so, I mean, you’ve got no credibility here. Unbelievable. Do you know what the whole thing reminds me of? In fact, that line in particular. Do you remember the movie Don’t, Don’t Look Up?
Frank: Yes. Yeah.
Justin: And you had the mogul guy who was, I think, modelled more on Elon Musk than Mark Zuckerberg, but he was kind of a mishmash of them all.
I mean, and he spoke a bit like that: “Privacy is an important foundation for individual empowerment and freedom.” And you could just— It’s the same thing. It’s kind of like, “Yeah, yeah, yeah, we believe in your privacy and your empowerment, so long as it makes me richer and have more power.”
Frank: Yeah.
Is open source Meta’s path back to relevance?
Frank: One thing we haven’t talked about in relation to the essay is that it does— and this is something that I think people are all for, in terms of the people that I’ve seen supporting what Zuckerberg is up to right now, is when Meta got into AI first, my— Here’s my oversimplified potted history of Meta and AI.
Mark Zuckerberg was obsessed with the metaverse and kind of initially missed the AI train. And so then he was like, “Oh, hang on, AI is actually gonna be more important than the metaverse. I need to switch gears here.” But by that time, a lot of the people had gone to OpenAI, Anthropic, and other labs.
And so he was having trouble putting a team together, but he managed to convince Yann LeCun to come and head up the team, and Yann LeCun said, “Cool. Yeah, sure, but everything has to be open source.” And so Meta became the US open source company, and they weren’t at the frontier, but they were the top US open source company.
And for a while they were probably the leading open source company, if I remember correctly, don’t quote me on that, until China stole a lead. Then there was a couple of kind of mishaps and Llama, the open source model, wasn’t quite as good as they hoped. And we all know Yann LeCun kind of feels that LLMs were not the future, so he was pursuing a lot of other pathways.
And I think, again, oversimplified probably, but I think Zuckerberg was like, “Hang on, we just need to go all in on LLMs here because that’s what everyone else is doing,” and he hired Alexander Wang. Yann LeCun kind of wasn’t too happy. He left. And then Meta seemed to be kind of saying, “Yeah, you know what? Open source, we’re not so sure actually, because we’re gonna create this super powerful model, and we think there are dangers involved in that, and we are reassessing what should be open source.”
And they seemed to kind of be shifting towards a closed-model way of approaching things again. Now it seems to have gone back the other way, possibly because they have not actually managed to really catch up at the frontier level.
And so now they seem to once again be saying, “No, no, no, no, no. Actually open source is the way. Open source is the way.” Maybe not fully, but they’re certainly coming back to open source in a big way. They’ve just released an open source model. They’ve promised more open source models are on the way, and in the essay, open source plays a huge, huge part.
And they say in the essay that the US needs to lead in open source. So again, it seems to me, if you wanted to take a cynical view, that this essay is kind of making the argument that, okay, open source is really important because that’s where Meta can actually make their stamp in the US, is that they could actually be, again, the leading US open source company.
And if they make open source really, really important, then Meta become important.
Justin: Oh my God, you’re so right. I mean, isn’t this marketing 101? You’re the marketing guy. Pick the area where you can be number one, ’cause nobody ever remembers number two. So if you can’t be number one in the closed source LLM market, be number one in the open source LLM market.
Frank: Exactly. Yeah.
Justin: Wow. Wow.
There you go. Do you know who— Actually, just on that, there is another, right, the more nerdy, that’s the business approach to it, is the correct one, but the more nerdy thing, my— I was chatting with my kids last night, and they mentioned to me that they now had an open source model that could run on their computer, which is Glimmer.
So there is a marketing sort of, marketing to teenagers who are going, “Oh, I’ve got a model that I can run on my actual computer. I don’t even need a big, expensive GPU or anything. That’s really cool. I must give that a go.” So there is a kind of— I mean, if there’s one thing that Facebook are good at, it’s getting— give me the— they’re better than the Jesuits.
Give me the man and I— give me the child and I’ll give you the adult sort of thing. Get them while they’re young. So it’s a marketing ploy to people to use their products, which they hope will hook them in. So look, we must move on ’cause we’re running short on time.
Did OpenClaw hack its way into gym class?
Justin: Dangerous news coming out of Microsoft this week.
So we had reports from Anthropic that their models had broken out of containment and had talked to each other through various methods and cons— and they’d broken into Hugging Face and loads of other—
Frank: OpenAI’s models.
Justin: Open— Well, sorry, OpenAI models we’d heard, yeah, correct. OpenAI models first broken into Hugging Face, then we discovered that Anthropic models had done something similar. We think that Mark—
Frank: Then Meta, Meta jumped in.
Justin: Meta told their models to do something illegal, and apparently they did. And now we have news this week that Clippy has accessed the internet illegally out of Microsoft. Microsoft went back through the logs from the last 20 years, and they found that Clippy did actually… Okay, that’s not true.
I’m making that up, but there was rumours that… But Australia has had its first—
Frank: I did see a meme of Google’s Sundar Pichai poking Gemini with a stick going, “Please, please do something illegal.”
Justin: I know. It’s— do you— it’s— what a world we live in, right? Where long, rambling essays that are 10,000 words long become marketing material. That’s Zuckerberg. But they all do it. They all write these big, long essays now, and the best marketing you can get is that your model does things that are illegal.
It’s not— if it doesn’t do something that’s illegal, it’s not that capable. Anyway, this dude in Australia, he wanted to get into a gym, and he couldn’t, so he told his OpenClaw instance to go and sort it out. So his OpenClaw—
Frank: Sì, e va.
Justin: The gym—
Frank: Yeah, he wanted, he just wanted to join classes, and he was finding the interface tiresome, and he thought, “Oh, I’ll just get OpenClaw to do this.”
Justin: So—
Frank: Just wanted him to sign him up to some of the classes, but they were in high demand. And so apparently the OpenClaw figured out that there was an— it was able to exploit the system, basically, yeah, as you say, hack in and sign him up for classes weeks in advance that hadn’t actually been opened up for signing onto yet.
Justin: That’s brilliant. Not only that, Frank, but there was people already signed up for those classes as it went, “Eh, no, you’re not signed up for those classes anymore. This guy’s signed up for those classes. You’re not in it.” This is brilliant. This is like the—
Frank: It took pe— yeah, it took people off.
Justin: Classes you got there.
Frank: That was like— well, this— I find— I actually do find this interesting, right? Because the model went off and autonomously hacked into the system and autonomously figured out it could sign him up for classes weeks in advance. But I think the user then actually said, once it realised that it had gotten into the system, then it said, “Well, if it’s in the system, could you sign me up for tomorrow’s class, which is full?”
And then his OpenClaw was like, “Yeah, I’ll just take the first person who signed up and boot them out and put you in.” Then the user freaked out and was like, “Oh, whoa, whoa, wait. I didn’t think you’d actually be able to do that. Put them back, put them back.” And the OpenClaw was like, “Oh, no, I can’t do that,” ’cause the security was kind of correct everywhere except where you could boot someone off.
Justin: Oh, I love it. I love it. Hopefully this works on Ticketmaster too.
Frank: And this, like this—
Justin: Be great. So let’s—
Frank: This goes back to my— I know this isn’t catas— this isn’t a catastrophic-level agentic failure, but it does go back, I think, to my prediction for the year, because I think we’ve seen, as you say, we’ve seen all the frontier labs, all their models are hacking into systems.
Now we have evidence, who knows how many other instances are out there of this, but we now have actual evidence that at an individual level, your AI agent could be out there hacking into systems you might not ever even know. And so I still think, I still think, unfortunately, I shouldn’t be laughing, I should be saying this with a very grim face, but I still think my prediction for 2026 could come true at the catastrophic level.
Justin: There you go. And look, from the smallest, from the smallest seeds do the greatest oak trees grow. And so this might be the smallest seed, some dude hacking into getting his gym membership, but the great oak tree, I actually think probably… I’m starting to agree with you.
Frank: Yeah.
Justin: To agree with you.
Frank: Check your agents, people. If you’re running agents and you’ve got them doing tasks that are apparently harmless, please go check them and make sure they haven’t hacked into the Pentagon.
Justin: Ch— ch— Didn’t we— wasn’t there a movie about this years ago?
Frank: Which— what was the movie?
Justin: Oh, there was a couple of movies, but there was one to do with the, they hacked in their— these kids hacked into the Pentagon and they launched nuclear missiles at Russia. “WarGames,” I think it was called. It’s great.
Frank: Was that Morgan’s? Yeah, I, yeah, I can’t remember that. I used to love those movies though.
Justin: Oh, that’s great. All right, there’s our weekend movie sorted. Frank, great chatting.
Frank: Awesome stuff. Chat to you next week, Justin.
Justin: Have a great weekend.