GPT-5.2 arrived in ChatGPT without warning. No fanfare, no countdown — it just appeared, carrying some very serious credentials. A year ago, the o3-high model cost around $4.5k per task on the ARC AGI benchmark. Wait until you hear what 5.2 costs now. You genuinely won’t believe it.
Justin was already mid-experiment, pitting Claude Code against OpenAI’s Codex, when 5.2 dropped. So he ran the whole thing again, and got a completely different result.
Also on the table: whether AI agents still need tight supervision, why Disney is suddenly doing deals with OpenAI, and why AI-generated images outperform human ones in ads — right up until you admit they’re AI.
And finally, an AI product placement idea that could ruin all your favourite shows.
Links to content we discussed
- Introducing GPT-5.2
- Introducing: Devstral 2 and Mistral Vibe CLI
- Yann LeCun taps European talent for new startup, says Silicon Valley is ‘hypnotised’ by generative AI
- The Impact of Visual Generative AI on Advertising Effectiveness
- Elon Musk’s xAI Says It Can Now Stuff AI-Generated Product Placement Into Any Scene of Your Favorite Movie
Three Insights from the episode
Insight 1: GPT-5.2 shows a sharp drop in cost at the same benchmark level
The discussion around GPT-5.2 focuses on a direct comparison with a model released roughly twelve months earlier. Frank references o3-high scoring 88% on ARC AGI at a cost of approximately $4,500 per task. Justin then points out that GPT-5.2’s thinking model scores higher, at 90.5%, while costing $11.64 per task.
Both hosts describe this change as a large efficiency improvement, citing a figure of 390×. The comparison is used to respond to claims that AI progress has slowed or reached a scaling limit.
The point made is not about new capabilities, but about the relationship between performance and cost changing rapidly over a short period of time.
What this establishes: performance on the same benchmark has increased while cost has dropped significantly within one year.
Insight 2: GPT-5.2 changes how Justin approaches agent supervision
Justin explains that his recent testing has caused him to rethink a position he previously held. He had argued that AI agents should be broken into very narrow, tightly controlled tasks to remain reliable. After working with GPT-5.2, he says he no longer believes the same level of task decomposition is always necessary.
He clarifies that this does not mean removing safeguards, giving unlimited access, or allowing unrestricted behaviour. Instead, he suggests that when an agent is given a clearly defined task, appropriate data, and a limited set of tools, it can be allowed to work more independently than before.
What this establishes: Justin reports a change in his own workflow assumptions based on recent hands-on testing.
Insight 3: AI advertising performs differently depending on disclosure
Frank references research comparing human-created advertising images with AI-generated ones. According to the study discussed, AI-generated images performed better than human-created images when used in ads. However, when audiences were told the images were AI-generated, that performance advantage disappeared.
The conversation then turns to demonstrations of AI-based product placement. These examples show how AI could be used to insert branded content into existing video scenes. Both hosts react negatively to the idea, describing it as intrusive and undesirable. The examples are discussed as demonstrations rather than widely deployed products.
What this establishes: AI-generated ads can outperform human-made ones unless disclosed, and there is concern about potential uses of AI for product placement in existing media.
Transcript
This is an AI transcription and may contain errors
Frank: Before we get into it, I’m just curious. Justin, have you been, have you been on YouTube lately? No.
Justin: Well, I mean,
Frank: Yeah, but I mean, no. Why you get that end of year. This is, this is what you’ve been watching, message and, and, oh.
No, I did actually wonder. No wonder is it, ’cause I’m in the States, I don’t know because I don’t ever remember seeing it before. But it’s kind of like a Spotify Wrapped thing, but, but it’s on YouTube. So I, I’m gonna run, I’m gonna run past my, my top watched channels past you, and we’ll see, see if you’re surprised.
Right. So we, we’ll look at the top, the top. My top three. So tell me on, we start with number three. Uh, no, we start with number one, right? Better be The AI Argument. Here’s what number I want you to say. Like a scale of one to 10, how surprised are you? Right. So apparently my number one watched channel was Wes Roth.
Justin: Oh, I’m not surprised. Wes is great. I love Wes too.
Frank: Yeah.
Justin: Yeah.
Frank: So my number two spot, right. No surprise really. I think, uh, Matthew Berman.
Justin: Also great. Yeah, also great. No surprise.
Frank: So number three, in the number three spot, do you wanna make a prediction? Who was in number three?
Justin: The AI Argument.
Frank: Andrea Bocelli. On a scale of 1 to 10, how surprised are you?
Justin: I’m actually… one… like not at all surprised.
Frank: Oh, really? I was very surprised. I was shocked. I was, what on earth is going on? Do you know what happened? I left my YouTube, uh, I left it logged in on my mum’s TV.
Justin: I don’t know, you’re kind of a sophisticated debonnaire type of a chap.
Frank: I like to give off that air, Justin. It’s not actually true at all. Cool.
My true number three channel was actually Theoretical Media, which makes perfect sense. ’Cause I go to Matthew Berman and Wes Roth for all of my kind of latest model updates, like when it comes to LLMs, et cetera. And then Theoretical Media I go to for all my kind of more creative tools, updates.
Yeah. So yeah, for all of our viewers, they’re all definitely worth, um, worth a follow as well.
And I’m sure Andrea Bocelli is as well, but I couldn’t say for sure because I don’t like opera.
AI Model Updates and Surprises
Justin: If you go back 12 months, right, to The AI Argument, and you look at this show, 12 months ago, both of us were holding our breath waiting for new models to drop. We couldn’t believe it. It was like, when are we gonna get a new model?
We’re practically hyperventilating at the moment, right?
Yeah. There’s
Frank: So much. Like you said, we were holding our breath. Now we can’t take a breath. It’s unbelievable. So were you surprised? Were you surprised when we got 5.2 yesterday?
Justin: Yeah. Yes.
Frank: So ChatGPT 5.2 just right there popped into our, like, there seems to be very little.
Prediction Markets and Insider Trading
Frank: I know people because of the, um, these prediction markets, which I’ve only just learned about, by the way, these prediction markets where you can kind of bet on any event happening whatsoever. And apparently they’re turning out to be really good predictors of when and what will happen.
Justin: Oh, really?
Frank: The prediction markets apparently said it was going to be, uh, December the 11th, that OpenAI would release ChatGPT 5.2.
And so it came to be. Some people are not surprised because they’re saying there’s a lot of insider trading going on in these prediction markets.
Justin: Yeah, of course. Yeah. Yeah. Well, I mean, that’s kind of how they work, the prediction markets, in a way. You know, the, uh, I don’t wanna get derailed here, but the CIA did a prediction market on when, uh, depos and, and dictators would get bumped off, and the depos and dictators didn’t like it and they had to stop it. But anyway, that’s, we’re getting, we’re getting distracted now.
Um, so this came for me. This came outta nowhere, right. And I started using it, and I, we’ll go into that a bit more, but this is like a, a Claude 3.5 moment.
It’s kinda like, wow, that I didn’t see that coming. And, ooh. That is good.
So it came out yesterday. We were chatting yesterday at our pre-show and you said, oh, look at 5.2, is that, and I went, oh, did it really? Okay. And I checked it out. It had, it had literally come out an hour beforehand.
Frank: Like they, they release a model and we, we both just had it immediately.
I’m just gonna go on it. This, I’m gonna go off, this is my tangent day, clearly. I’m gonna go off on a slight tangent and say, I like how OpenAI roll things out. They tell you it’s coming. They give you some kind of idea of when you’re gonna get it, and that normally happens. With the exception maybe of Sora 2, you know, which we don’t know when that’s coming to Europe, et cetera.
Google. I’m just gonna have a quick rant about Google. Google release things and I get all excited about them and then they, they never come to Ireland. Their support. People don’t know if it’s actually going to be on your plan. It’s, anyway, rant over. Let’s carry on with 5.2.
Justin: Alright, so let me just explain, right.
Claude Code vs. Codex: A Detailed Comparison
Justin: So what I did this week, so this week I sat down and I said, right, I’m gonna do a side-by-side test. I sat down on Sunday, um, maybe Sunday, maybe Monday. I’m gonna do a side-by-side test of Claude Code and Codex—Codex, the agent harness, not Codex the model, right? So Claude Code and, and Codex…
Frank: Codex from OpenAI?
Justin: Correct, from OpenAI, and I’m gonna ask them both to write software and I’m gonna see which one does a better job just to get a feel for which one, right. So exactly the same prompts and the pieces of software I was gonna get it to write was to write a test harness for this thing called GEPA. The only reason I mention GEPA is if you’re into writing software for LLMs, go check out GEPA.
It’s a way that you can do an evaluation of your prompts and GEPA will iterate through your prompts and improve them and make them better, and then give you a percentage at the end and say, now you’re getting 90% whatever. Super important you do that, go do that sort of thing. Evals are really, really important.
So, but you gotta write software. So I got them both to write the software and at the end, just as big reveal, was that the code that came outta Claude Code was better than the code that came outta Codex. It was more readable. It was just simpler. It was better code, right. The other big reveal was that the test case.
Frank: Uh, did Codex actually like manage to complete the task?
Justin: Okay, so both models, the task was I asked them to write the test harness. I asked them to write the unit tests and I asked them to write, uh, test cases, so synthetic data to push through this test harness, right? So that was the task I got them.
This is a new library, probably not in their training data. So they had to go onto the internet, read the website, download the software, understand the instructions, and then go write some code. So a good test, right? They both wrote software that worked, right, so functionally. They both work.
The code that Claude Code wrote was easier to understand, right? So if I was to give it to somebody else and say, this is gonna do a certain thing, Claude Code wrote better code.
The test cases, the synthetic data: Claude Code was streets ahead of Codex. Way better, right. So we’re talking like, I won’t go into the details, right, but Codex is almost autistic. It just, it’s very, and I kind of like working with it ’cause it’s so succinct, but when it comes to generating synthetic data, you want lots of data, not like a tiny bit.
So it sort of falls down on that, but the cost difference was massive. It cost me, I would say, not much shy of a hundred dollars to get Claude Code to write what it wrote, and it cost me about $20 on Codex.
And so by the end of it, as we were chatting about it, my, my thing was: if I’m doing this for somebody else, with somebody else’s money, I’m using Claude Code every day of the week. If I’m doing this for myself and using my own money, I’m using Codex every day of the week.
That was Thursday.
Frank: And, and like that makes—That was Thursday.
Justin: Yeah, exactly. Yeah.
Frank: But it kind of, just to clarify as well, like it kind of makes sense, right? Because if you’re, if you’re doing it in a work environment, you want the more easily readable code. You want something that you can easily put into a project and everyone can figure out what it does and understand it and contribute and what have you.
If you’re working at home, the kind of that, that level of readability or understandability of the code isn’t quite as important, ’cause you can pick through it at your own pace and figure it out.
Justin: Yeah, and it’s your own money and you’re sort of, you know, it’s kinda like, yeah, I’ll put in a bit of extra pain for my own money. Anyway, that was Thursday. So then we talked. What’s the story now?
OpenAI’s New Model: Performance and Cost Analysis
Justin: I sat down this morning and said I’ve gotta do exactly the same test with the new model.
Oh my God. Right. So I saw just like you did the, the, the model card, right. And I saw all the numbers and they looked amazing. It, all, the numbers are correct, right.
The code that Codex now came out with, with this new model, was even better than the Claude Code code. Wow. The test scenarios were actually, I’d still say Claude Code wrote better synthetic data, but there wasn’t much in it.
Right, right. But the code was brilliant.
The cost: it cost me, I think it was $7, $7 versus nearly a hundred. Right. That’s a huge difference. And the code was, I would say the Codex one was now better code, right. It was even simpler. Again, well-written code.
The only downside that I saw with this new version from OpenAI was it is slow, slow, slow.
Frank: Oh, interesting.
Justin: Yeah, really slow. So I, I started work and I could see that it was slow. So I just put a screen here beside me. And I put it into, you know, no protection mode. Just don’t ask me for anything, just go do what you do, right. And I hit go on the thing, I gave it the instruction, which is quite detailed.
I’d say it went for three and a half hours doing its thing, just churning through, doing its thing. So, um, it’s very, very slow. Like, I mean, just, you’d wanna have a couple of windows open, but it’s not costing you a huge amount of money, like seven or $8. Like it was nothing.
And what I noticed, right, this is really important. What I noticed between the two of ’em was when I was working with Claude Code, you get this message over and over again saying, I’m compacting my context window. It fills up its context window. It just seems very chatty, right? It’s really, really chatty and writes beautiful code as a result, right? But it’s very, very chatty.
Whereas Codex, I could see it, it went for three and a half hours. I don’t think, I think at the end it did clean up its context window and it does have a slightly bigger context window, but still it’s not going through the context window nearly as quickly.
So it’s taking its time, but it’s, you know, it’s like that old saying, forgive me for writing a long letter. I didn’t have time to write a short one. And that’s basically Claude Code versus Codex, right? Codex is writing the short letter, but it’s taken it a long time to do it.
Frank: Right. Interesting, interesting. Yeah, I saw, um, I think it was Wes Roth actually, was talking about how it’s no longer kind of, you know, do a prompt and, and get a little task back. You can actually give it projects and it will just go away for an hour and come back. And he was, he was saying interesting.
Now I think he might’ve been on the, I think he might’ve been using GPT-5 Pro. Um, but he said, you know, it’s not coming back in the canvas anymore and just giving you a kind of a block of code. You’re actually getting a zip file. This is done.
This is just using the user interface. We’re not, I’m not talking about Codex or, you know, I’m not talking about any of your fancy, uh, coding IDEs or any of that. This is just the consumer interface. It’ll give you back a zip file of an entire project.
Justin: Yep. Which is incredible. How does it work?
Now can we just, I was just saying about this, this is Christmas. We’re basically at Christmas, right? Can we cast our mind back, right, to last Christmas. Right. We were both amazed when, remember they had the, the big reveal, I think it was in the 12 days of Christmas when Sam Altman brought in the guy, I can’t say his name, Francois Orle, whatever it is, the head of ARC-AGI. The people that make this ARC-AGI.
To a big reveal. They said, look at, this was the first model and they were releasing O3, and it was the first model that did about the same as a human. And the human level was 80%. And it did it, but it cost, you told me the correct number. About five.
Frank: Yeah, I have, I have a note on here somewhere. Um, the ARC, yeah, the ARC price. So O3 high on ARC-AGI 1 got 88%, but it cost four and a half thousand dollars per task.
Justin: Okay, that was 12 months ago.
Frank: GPT-5.2. Do you have the same number there as well? GPT-5.2 thinking model 90.5%. So a better score at $11.64 per task. They said it was a 390 times efficiency improvement. That’s unbelievable.
Justin: If that doesn’t blow your mind, and if that doesn’t sell you, you know, this whole rubbish that was about, oh, we’ve hit a wall and scaling and slowing down, if that doesn’t tell you that things are moving along at the same pace or faster than they were previously, I don’t know what will, right?
Every again, this week, right, I’m seeing stuff coming up where people are like, yes, there was this unsolved theorem that, you know, AI has solved. Yes, this other problem. AI has solved it, right? We’re seeing these scientific discoveries start to trickle in in areas that are kind of esoteric, right?
But, but this is in your face: 390 times an efficiency gain in 12 months. And you and I know that 12 months from now, it, I mean normally this, it’s 10 times, right. 390 is mind blowing, like incredible. It’s just mind blowing.
Frank: Incredible. And the other, the other one that everyone’s talking about is something called GDPval. So another benchmark. Um, but the reason everyone’s talking about it is because it is entirely based on real world knowledge task evaluations. So in other words, the real work that real people do day in, day out, how do the models perform at it.
And so 5.1, the previous chat model was at 38.8% on this GDPval, um, and 5.2 thinking 70.9%. So again, huge leap in the ability to do actual real world tasks that people do in their jobs. Amazing.
Real-World AI Applications and Efficiency Gains
Justin: In my tests, right, in all my working that I’ve done this week, uh, my mindset has shifted somewhat as a result of spending the week doing this. So up until now, right, I, and we talked about this, right, I’ve been really open about it. I’ve been saying you would be crazy to let a model go and do its own thing, right?
It’s gonna be, in most big organisations, you’re gonna, you know, sort of, you’re gonna break down into little bits and you’re gonna do this, this, and this, right? And you’re gonna, the narrower you can make that, the more reliable your AI agent is gonna be.
After my week of running these through, through their paces and seeing this new model, I don’t believe that anymore, right. If you can set up your agents and give them the right tasks and the right data, they don’t need the same level of supervision that they did six months ago, you can just let them go and do things.
Frank: With the right safeguards still in place, I would think and hope.
Justin: I mean because, yeah, yeah. The point I’m making is not that, for sure, right. And also the same, like, you don’t, it’s not carte blanche, right? You’re not saying here’s all, here’s, you know, you have access to all of the data in the company and you have access to every tool, just do whatever needs to be done.
Right? What I’m saying is if you had a task that needed to be done, right, and you give it the right data, right, and the right tools and a small set of them, I’m saying that you could happily let the thing go off and do what you want. You don’t need to break it down into very atomic steps, right, and be very prescriptive about what it’s able to do, what it isn’t, right.
You like my, it’s, it’s looser now, right? It’s, which makes it more economically valuable. That’s what GDPV is testing, yeah. And what I’ve seen is, you know, the models are smart. Smart relative to what they were six months ago. You know, if, if you tried doing stuff six months ago and it didn’t work, go try it again. Because I think it’ll work.
Frank: Yeah. Big time. Big time. Um, will we move on to other interesting news?
Justin: Uh, do, but just, can I just say 5.2 has blown me away? Like it’s, it’s, it’s taken, it’s given, it’s, it’s poked me between the eyes. 5.2 and, and maybe go, like we have now.
We’re so lucky we have that, ’cause Opus 4.5 had sort of done that, right? Had maybe gone, wow, this one is really good. This is a step up, right? 4.5 was good, but it was kind of pricey. And now we got 5.2 and it’s even a little bit, I mean, they’re very close, but it’s just, it’s so cheap.
It’s like one tenth the price. It’s, it’s just the, the, the intelligence for dollar is just incredible with this one.
Frank: And it’s, it’s really interesting to hear like about your real world experiences in actually using and testing the various models, because as we said in last week’s show, it’s all very well to look at the benchmarks, but like, you know, 30% to 70%, okay, great, but like, what does it actually mean?
And you know, even as, as Ilia Berg was pointing out, the benchmarks are amazing, but then we frequently find that we try and get it to do something and it falls over and we don’t really know why. So it’s really like, I just find it really interesting to hear a genuine, like this, the real use case, something you were actually working on, evaluated across the three models.
Um, I think that’s, you know, it’s really useful to hear those kind of case studies.
Yeah. Blown away. Blown away. Anyway, go on. What else was going on this week?
Uh, Disney. This? Yes. Blew my mind actually.
Disney and OpenAI Partnership Announcement
Frank: So Disney have announced a deal with OpenAI where they are going to allow Disney characters to be in Sora. Sora 2 being the, you know, the video generation, uh, kind of social app that OpenAI released where you can, uh, does anybody use that?
Uh, they do. I think, I think interest has waned a little bit, but I think interest waned a little bit, partly because when they launched you could pretty much create anything with anybody. And then as we pointed out on the show, there was some, like Martin Luther King’s estate getting out because, you know, and they started to, you know, and, and other IP owners started to say, uh, hang on a second.
Uh, I think Nintendo got in touch them and said, uh, excuse me, Super Mario is our property. Why is everyone creating Sora videos with Super Mario doing things that we do not approve of?
Um, and so the safeguards, the guardrails were tightened. And I think that, you know, for a lot of people, once, once the fun, the fun was taken out of it when the, when the guardrails were put in place and they couldn’t generate with certain characters, but people are still using it.
Um, and now, you know, this will kind of loosen things up a little bit again. All your favourite Disney characters will be available in Sora.
And what were the details of the deal? They’re like, they’re suing Disney or suing everybody—Midjourney they’ve just sued. They’re just suing Google now as well.
Um, so they’re suing everyone, but then they surprise everyone by coming out and doing this deal with OpenAI. They’re also becoming massive investors. Like they’re, they’re investing a billion in OpenAI. OpenAI are going to give ChatGPT access apparently to all Disney staff.
So it’s, it’s more than just a character. It, it appears to be a, a much more integrated partnership than just a kind of, um, an IP licensing deal.
Justin: Okay, so lemme just get this straight, right. So, so OpenAI and Disney have done a deal, you can now generate Disney characters on OpenAI’s Sora model, and Disney, instead of suing OpenAI, are now giving OpenAI… they’ve paid OpenAI a billion dollars to allow them to generate Disney characters on Sora. Am I getting that straight?
Frank: Yeah, that’s my understanding of it. And, and, and the, the kind of like the best Sora videos will be available on the Disney Plus channel.
Justin: I mean, that’s incredible. Hold on, Derek. Why is the money going that way? Why is the money going—
Frank: I dunno. I don’t know unless I, unless I have missed something in my readings of like several articles about this, I have seen no information about money flowing the other way. So I, I don’t know. That’s incredible.
Justin: I—
Frank: I mean, who could—
Justin: Anyway look, just in case and would—
Frank: Would, and like, would you have expected this from Disney? Like I would’ve, I, I don’t know. I might’ve expected it from a smaller player, but I, I still think of Disney as being like one of the most protective brands about around their IP.
Now there is one thing, one interesting thing actually about this as well that, that I missed initially. Um, it is not, you know, it’s not across the board. Every Disney character. I think it’s about 200 characters, and it does not include actors’ likenesses. So you will be able to generate, say, Thor.
I don’t know how that’s gonna work because you will not be able to do it with the actor’s likeness in terms of the, you know, the Marvel Cinematic Universe. It’ll be like a more generic version of Thor or like a version of Thor from the comics or, we’ll see. We’ll see.
But otherwise you will be able to do, I think, I believe you’ll be able to do characters like Iron Man who, helmeted, um, and you’ll be able to do all their animated characters.
Um, so, but just, yeah. In terms of the, the, the actors union strike that happened there and the, the terms that, that actors were hammering out about AI, actors’ likenesses are definitely not included.
Justin: Yeah, that doesn’t surprise me. ’Cause I think there’s a number of actors that put it into their contract to say that you, you know, you’re not buying my AI likeness, you’re just buying me on the film, which makes a lot of sense.
If, yeah. Well, I mean it doesn’t, I kind of disagree with that, right. They should do, the deal should be that compensation is included. Yeah, sure. Pay me, pay me to, I mean, pay me and my estate to use my likeness. Yeah. But you do not get it by default, just because I appear in this movie. Yeah. Yeah. That makes total sense to me. I’ll be fine with that.
Um, interesting. And there’s been a lot of stuff, so what do you—Anyway, that’s fine, right? So, but there’s been a lot more going on, right? So, I’ve also, it doesn’t surprise me, what it says to me is that, uh, Disney are still gonna start using AI in their movies, right? Maybe.
I mean, I, I guess that that’s really where this is going. And you can imagine a scenario where AI and non-AI are sort of interleaved at a movie and sort of, you know, a bit like when computers come out first, right? You still maybe hand drew a lot of the cartoons and then computers filled in the gaps and made them more fluid. The AI initially at least starts to fill in that gap.
Frank: I think, you know, I think I saw somewhere, and I, I couldn’t tell you exactly where. It might’ve been Theoretical Media actually, who I, we mentioned at the top of the show, or Andrea Belli.
Uh, it might, maybe it was Andrea Bocelli told me about this. Yes.
That, um, you know, animators are just as up in arms as actors in terms of AI, and this gets AI onto the Disney Plus channel in a way that Disney is like, you know, not responsible for in inverted commas. In, in that it’s user generated content. It’s not being generated by Disney itself.
So I think, and I think this is what you were kind of alluding to, that it kind of gets AI in the back door and starts to get people, you know, used to seeing AI generated content from, from Disney, but not from Disney. And then they can maybe, maybe it just puts a little, uh, puts a little—
Justin: I don’t really want user-generated AI content. Uh, like I pay a tenner a month to Disney. I don’t want some slop from the internet going on to my Disney Plus feed.
Frank: Well, you know, there’s actually some pretty good fan created content around these, um, characters. So you never know. There might be some very interesting stuff.
Like, I believe, uh, I, I was on, um, I was on a, um, a meeting yesterday with, with members of the Rise community and someone was saying that actually there have actually even been some YouTube creators creating fan content who actually ended up being brought into Disney because their stuff was so good.
Justin: Interesting. And is that, I mean, that’s fascinating. So is what’s happening here, maybe. Maybe that’s why Disney are, you know, Disney are basically, are, are paying OpenAI. OpenAI make this available. And so instead of paying actors and writers and animators and stuff, you just get stuff for free.
Your content is made for free, and then you get, they’re almost turned into—and then you publish it on your platform and people pay you tenner a month.
By the way, hate Elsevier. Elsevier’s publishing, it’s one of the most hated companies in, in, uh, publishing. Dutch company, uh, they publish academic papers and the business model that I’ve just described to you is their business model, but for academic papers.
So academics give their papers for free to Elsevier. Elsevier do, you know, very little, add value, but then they publish the papers and they charge people to buy the papers, which they were given for free. And they’re hated as a result. But the Disney business model is trying to do exactly the same thing.
Frank: Yeah, yeah. Potentially, potentially. Um, wow.
European AI Innovations
Frank: Very briefly, you had a couple of, uh, good news stories from Europe, which—Oh, listen, what are we doing?
If anyone regularly watches the show, they’ll know that usually it’s me saying, no, Europe, Europe’s doing great work in terms of AI regulation, et cetera, and you’re usually the one saying, come on, you’re up. What are you up to? You’re messing this up. I look.
Justin: We have to, we have to give credit where credit is due, right? So really quick, shout out. Fantastic week. We’ve had two new companies created, and that’s kind of a joke and I shouldn’t joke, right? It’s a big deal, right?
So Mistral, right? The French AI company who we’d sort of, you know, they’d sort of came outta the blocks really fast and then, you know, we hadn’t heard from them a while. They’ve released a new model. The model is called Devstral. It’s seven times more cost efficient than Claude Sonnet, which wouldn’t be hard given what I said earlier, but I, I’d say if you, yes, if you used a, uh, test harness like an agent harness, it’d probably be 70 times more, uh, cost effective than Claude Sonnet.
I also just want to call out their website, their UI and their stuff. It’s really cool. These people have style, right? Really, really good style. It just looks beautiful. So it does. They go check out their website. It’s really, really nice.
And, uh, they’re, it achieves 72% on the coding benchmark. Now I think the OpenAI model had got 80 or 85 or something like that, the most recent one. But this is an open source model. It’s way smaller so you can run it locally. Um, amazing work, right? Really, really good.
Second piece of good news, your best friend, Yann LeCun, previously of the Meta parish, has set up a new company and he is going to set it up in Europe and it’s gonna focus on robotics and AI.
So he’s always been big on world models. This is clearly his thing now, but again, like one of the big names, one of the founders of AI and he’s setting up his company in Europe. So—And he’s, he is—
Frank: Originally French, I believe. And so he is basically, he’s coming back home.
He’s French, uh, and I saw, I saw a great quote. Um, I don’t know if I—Did I pop the quote in there somewhere? Uh, yeah. He said, um, Silicon Valley is completely hypnotised by generative models, and so you have to do this kind of work outside of Silicon Valley. In Paris.
And I just thought that was really interesting ’cause I think we talked on the show at some point about, um, a Chinese fellow called Song-Chun Zhu and he was in America and he was doing, he was like one of the pioneers of, of great AI work, particularly in visual reasoning and stuff like that.
And he similarly did not believe that generative AI was the way forward. And he ended up moving back to China for ex. And he, he cited exactly the same reason that he could not get funding in Silicon Valley. Because he was not into generative AI.
Justin: Interesting. And Ilya said the same thing on the Dwarkesh podcast there a week or two ago, which was everybody’s doing the same thing. And in order to move forward, you need to do different things. So that is an idea that Europe should really jump on.
And French in particular, like the French, go back to France of the eighties. France of the eighties was just brilliant. They had style. That’s why I really pointed out the style thing. When I think of France, I think of France somewhere around 1987, 1988. They had still had really cool cars. The technology was quirky and different, and it was just beautiful.
And you know, we should just distil a bit of that for AI and say, yeah, leave those Californians, do their, do their LLM stuff. We’re gonna do stuff slightly different over here.
Frank: Yeah. Yeah.
AI in Advertising and Product Placement
Frank: Um, and then so we had, okay, so, uh, we’re running outta time, so I’m not gonna dwell too long on this, but, um, I think it was Ethan Mollick shared a paper and it was really, um, fascinating study where they looked at, they basically took an existing, uh, campaign that had run, so they had data for it, they had the creatives for it, everything.
And they looked at, well, how would the human creatives, the human images for it was for beauty products. It was online advertising for beauty products. How would the human created images compare performance wise against AI generated images.
Yeah. And what they found was if they modified the human created images with AI, there was no real difference. Um, but if they just said to AI, create an image for this campaign, they kind of gave it like much more free reign to create the image from scratch. They performed way better. 19.
Yeah. 19% increase in click-through rate. Wow.
Justin: That was 19%. That’s huge.
Frank: Yeah. Massive. Massive. And it was even greater apparently, if you allowed AI to like redesign the packaging. So if you didn’t even say you have to use this label or, you know, the label’s on this packaging and it did it themselves, it got an even higher click-through rate.
Now I don’t know how that works ’cause then that’s not the product you’re buying. But anyway, that’s a whole other, that’s a different question.
Uh, but it is kind of related to this fact where I think it gets really interesting. The benefits from using AI generated images were completely decimated and they performed much worse if you disclosed they were actually AI generated.
Justin: What do you think about that, young Mr. Prendergast?
Frank: Like, so this to me, straight away, I’m like, okay, so nobody’s gonna want to admit that their images are AI generated when they’re doing advertising. So immediately I’m thinking, okay, first of all, I’m like, this is why we need regulation.
If we want to know that the images we’re looking at are AI generated, we need regulation to say they have to be labelled as AI generated. But do we need to know they’re AI generated?
Like we absolutely 100% need to know that they accurately represent the product. Do we need to know they’re AI generated? And I think this is a really tricky question because on one hand, yes, if they’re accurately representing the product, that should be okay, but on the other hand, it opens the floodgates in terms of where’s the line and who polices it.
And so I wonder, is it, you know, while I admit that all we need is accuracy, we might be better off with the safest option, which is that if this is AI generated and it’s representing something that could be mistaken for a real thing, it should be labelled. What do you think?
Justin: Hold on. Wait a while. Wait a while, right? So, you know when you see an ad for, let’s say, ice cream, right?
Frank: Yeah.
Justin: Or you see an ad and they’ve got a burger and the bun comes down first, and then the cheese and the whatever comes down, right? And let’s say in particular the ice cream, right? When you see that ad and the ice cream is perfectly shaped and everything, it’s not ice cream, Frank.
They use all sorts of other stuff. Sometimes they use toothpaste, sometimes they use whatever material it is that will just look like ice cream and not melt under the lights. Yeah, it’s—Now in your perfect world, right, you’d have a big signpost over the ice cream ad going, it’s just toothpaste, and everybody would go, that’s disgusting. I don’t want that.
So why are you applying a different rule to AI? So long as it looks like the real thing, why not just say it’s the real thing?
Frank: Yeah. Where do you, like, where do you draw the line?
Do you make an exception for people like, I’m, you know, if, if, if it’s an AI generated person and they’ve given the licence for their likeness to be used and they’ve said that they, you know, I, I can’t help thinking about—Was it, remember in The Simpsons when, uh—
Uh, I think it was Krusty the Clown maybe, used to pop up and say, I heartily endorse this service or product.
Justin: Yeah, it was great. You know, you could, you, you could—The trade descriptions act draws the line, right, when you faithfully represent the product that you’re selling. Then you’re fine. And once you don’t faithfully—And I would argue, by the way, that if you ever see those, you know, euro meals, whatever, where you see a big burger with cheese and lettuce and tomatoes bursting out of it, and then you get the real thing and it’s sort of, sort of—Yeah, that, that would, you know, to me, you’re selling something that isn’t at all like the real thing.
Frank: Yeah, that’s, it’s true that, you know, it’s clear that people are voting with their clicks, at least, you know, according to this study. And there’s, there’s various other studies that have shown very similar results. Like people clearly don’t trust it if it’s AI generated. And should we just respect that?
I don’t know. I, no, we’ll have to return to this because I, I need to, I’ll have to solidify my opinion over the next, uh, couple of weeks.
Justin: Listen to me in a related note. Uh, a company that I’ve completely forgotten the name of, have introduced a piece of technology in which—I mean in the eighties, I remember hearing about this, I dunno why I’m going on about the eighties today, but there was product placement. Product placement was gonna be the next big thing.
So an AI company has come out with a system where you can place products into just regular movies and stuff like that, and it’s every bit as bad as you would’ve imagined back in the eighties and every bit as bad as you’d imagine it today.
So they gave some examples of, I think it was Mad Men or one of those, and this guy walking out of an office block and he is got a cup of coffee in his hand and he opens up the door of a car and then it changes very slightly, very subtly, but he turns around and he has a can of Coke and he’s like—
Frank: I have it, I have it lined up here, but let’s let, let’s have a look at it.
And it’s, um—Alright. No, it’s, I will say, right, I tried to figure out what on earth this was. ’Cause it was xAI tweeted about it, but I believe it’s from some kind of like hackathon that they had. So presumably these guys are, are intending on this is, this must be like their intention for like a hot new startup or something.
So let’s have a look and see. Well, let’s have a look. Let’s have a look and see.
So can you see that all right?
Justin: I can.
Frank: So they’re talking here about how ads are usually really interruptive and they don’t fit in the flow. And they say watching ads sucks. And then they say introduce a—Bathetime.
So they go into Suits on a Netflix type, um, player, and then you see the little yellow thing come up to indicate there’s an ad coming up.
And, uh, the guy from Suits is walking outta the office, just like you said, about to hop into a, um, a car and suddenly he turns around and holds up a can of Coke to the screen and there’s a learn more button.
So they hit the learn more button and, uh, they X outta the ad and he puts the can of Coke down, grabs a cup of coffee and hops into the car, and I think that’s probably enough.
There’s a Friends one there as well, which is equally as ludicrous. It’s, it’s everything as bad as you’d expect it to be. I mean, I think this is a, a horrendous idea to like—
Justin: It’s only gonna get worse, right. It’s only gonna get worse because instead of just holding it up, you know, he’s gonna be walking to the car in the middle of Suits, right, a really good programme and then, I love the taste of Coke in the morning. Really refreshing. And then he is gonna carry on. It’s gonna be like, oh my God, you just ruined the mood.
Can you imagine Breaking Bad and one of these ads pops out? The bit like, no.
Frank: Yeah. And like you’re not gonna know, you’re not gonna know it’s an ad until it’s too late.
I mean, sure, their point is like, oh, it’s not interruptive, but like, we want the ads to be interruptive. We want if, if we don’t want ads at all, let’s face it. But if they’re going to, if they’re gonna be in our shows, we want them to kind of like, yeah, you’re out of this world now. You’re watching an ad and then we’ll take you back into that world that you were just in.
We don’t want to be wondering, hang on. Why is he holding up a can? Oh, it’s an ad.
Justin: God, do you know what it is? You should get—This is way more, this is way worse than all that other airy fairy stuff. This will actually happen and it will be terrible as opposed to all of that biological and nuclear stuff that you, most of the time, you get yourself worried about.
Ads, AI ads are the thing which are gonna hurt us way more than any of that other stuff that you and—uh, carry on about. This is the end of civilisation if you know.
Frank: If they manage to do this and push this out, I mean, we will definitely want this to be, to have a big thing pop up saying AI generated, just to, to realise, yeah, this is not part of the show you’re watching guys. Actually, this is not the director’s vision. This is not what the actor actually did on the day.
Justin: Yeah, yeah. Oh, you know, it’s definitely gonna happen. By the way, I just, have you noticed in Netflix that there’s loads of, of shows that were clearly previously would’ve been dubbed and now they’ve just—So they would’ve been in another language and now they’re in English, ’cause they’ve just used AI to change the language.
Frank: I hadn’t actually spotted it, but we, I remember us talking about this, that, that this was gonna happen. God knows how many episodes ago.
Justin: Yeah. Well here it is. There’s gonna be AI ads.
Frank: Brilliant. Uh, Justin. Pleasure as always. I will chat to you next week. Who knows what models we’ll have by then and how capable they’ll be of taking our jobs. The 12—
Justin: Days of Christmas hasn’t even started yet.
Frank: I know. God.
Justin: Yeah.
Frank: Exactly. Yeah. Yeah.
