OpenAI Propaganda, Hidden AI Triggers, Copilot’s Santa Fantasy: The AI Argument EP83

by | Dec 19, 2025

Wired has sources inside OpenAI saying parts of its “research” are starting to look like propaganda. Frank reckons that’s what happens when IPO pressure creeps in, the bad news gets buried unless regulation forces transparency. Justin says Frank’s chasing the wrong villain: you can mandate disclosures all you like, but it won’t matter if politicians still don’t have a plan for the job shock.

Also on the table: Nick Huber wanting to ban AI-written client emails because they’re easy to spot and a waste of everyone’s time; ChatGPT helps recover a missing recording with some command-line wizardry; OpenAI’s new image tool trying to catch Nanobanana Pro; Gemini Flash coming surprisingly close to Pro; and Gemini 3 Pro looking like scary-good value.

Then it gets darker: “safe” training data that can still hide backdoor triggers, behaviour that stays invisible until the right little cue flips the model.

And finally: Microsoft Copilot’s Christmas ad selling a Santa-level fantasy the product can’t actually deliver.

Links to content we discussed

Three insights from the episode

The real argument isn’t “propaganda vs truth”. It’s “transparency” vs “a plan”.

Wired’s reporting frames internal concern that OpenAI research is starting to resemble a “propaganda arm”. Frank’s point is that IPO pressure makes labs less likely to publish anything economically ugly, so regulation should force transparency. Justin’s pushback is that even perfect transparency won’t matter if politicians still don’t have an actual plan for a fast job shock.

The cost of “useful intelligence” is collapsing faster than people are admitting.

A practical benchmark lands the point: Claude Opus can solve the utility-bill comparison cleanly but costs a few euros per run but Gemini 3 Pro gets the right answer for roughly 60–70 cents. The takeaway is that capability-per-pound is dropping so fast the economic impact stops being theoretical.

“Safe” training data can still hide backdoors — and the boring fix is access control.

A model can look completely normal, then flip into a buried persona or behaviour only when a specific trigger appears in the prompt. The worrying real-world angle is bias and misinformation: seed the right kind of “innocent” data widely enough, and you may be able to nudge outputs in a direction that’s hard to detect or prove. And if hidden triggers like this are possible, what might that imply once agents touch real systems and data?

Transcript

This is an AI transcription and may contain errors

The AI Argument EP83

Justin: Welcome to The AI Argument. Merry Christmas, Frank. It’s that time of year. Welcome to The AI Argument. It’s Christmas week. Santa is packed, ready to go. I would love it. There we go, and I’ve got the Christmas spirit all over.

Is Nick Huber right to ban AI emails?

Justin: One man who doesn’t have the Christmas spirit all over is Nick Huber. Nick Huber is not feeling very Christmassy.

Frank: Who’s Nick Huber? Fill me in.

Justin: He owns a company called Somewhere. He develops property, bolt storage, cost more, he does whatever, and he’s banning AI from his company. He’s had enough, he’s banning. Bah humbug to AI. So what he said is, and actually I agree with him, right, so he said he’s got several, several of his employees and they’re just entirely communicating with their clients through AI.

Justin: And he said it’s obvious and people can spot it. And so he’s on the verge, he said it’s a waste of time, it’s really visible to clients and he is on the verge of banning it. And I think he’s right not to ban it, right, but he’s right about the underlying point that people can spot it.

Frank: And I hope he is talking about like banning maybe that use of AI rather than just banning AI across the board. Because people have em dashes in their emails, right? Who knows, who knows? But he said he is, he’s on the verge of banning it. He said, anyway, he said, stop doing it. I have to admit, I would never, I would never send, I would never just pop something into ChatGPT and then send it as an email to someone.

Frank: But there are some tricky emails that I have to deal with that I’ll go to ChatGPT and say, you know, I’ll write my kind of knee-jerk response and then I’ll say, how would I make this sound like I’m not actually annoyed about it?

Justin: I get that. But, you know, you see it, you see it all the time. You’re scrolling through LinkedIn or you’re scrolling through whatever, and as soon as you see a couple of em dashes and more emojis than are normally there, you’re like, this is written by AI, I’m not gonna waste my time reading it, and on we go.

Frank: Yeah. And in terms of emails and stuff, there was that whole workslop study of like, you know, when somebody sends you something that looks like it’s relevant to whatever it is you’re on about, but actually it doesn’t move anything forward. It doesn’t advance the workload for anybody. In fact, you probably have to spend more time interpreting what it was that they meant than actually getting work done. So I totally get where he’s coming from.

Justin: I tell you what, right, I think, I would be pretty sure that by this time next year we’ll have sort of an email agent or whatever that just reads the emails. So, you know, it’ll surface the important emails and whatever, but summarise the rest of them and tell you. Now, if it was an email agent that I knew was gonna read the mail, I wouldn’t give a crap. I’d send them an AI-generated email then, ’cause it’s gonna be an AI reading it, so why should I waste my time?

Frank: So we just have our agents talking to each other. Why talk to people at all? Absolutely, you’re dead right.

Justin: Well, Frank, welcome to my world. I hark for the happy days of COVID. We weren’t even allowed to talk to people. But anyway, we have to move on.

Did ChatGPT just save the podcast?

Justin: So

Frank: I had a cool experience, just one of those little day-to-day ChatGPT helping you out in your day-to-day moments last week. I didn’t actually tell you this during the week, yeah, but last week the show did not record. So we stream to LinkedIn, we record it, and then I pop it up on YouTube. We put it on Buzzsprout for the audio platform.

Frank: So if I don’t have the recording, I can’t get it out to that wider audience. Public service announcement: anyone using StreamYard, so you know, if you use Zoom, right, everyone uses Zoom, if your recording quota is nearing capacity on Zoom, a big red thing comes up at the top of the screen saying, you know, you’re almost out of storage, upgrade or delete stuff, or, you know, yeah, yeah. In StreamYard, there’s a microscopic yellow icon in the bottom left of the screen to let you know. I totally missed it. I completely missed it, so we didn’t record.

Frank: I thought, no problem, I’ll download it from LinkedIn. Can you download a LinkedIn Live event from LinkedIn after the fact? No, you cannot.

Justin: If anybody knows, answer in the comments below. Had to do that.

Frank: So I thought, I’m not gonna be able to publish on YouTube or on the podcast platforms. So I went to ChatGPT and I said, look, this has happened, what do I do? And it went off. And what I did actually was I did a Deep Research so that it would really look into it for me, and it came back with various methods for downloading it.

Frank: Now, the first one did not work. The first one was, well, there’s loads of tools out there that will download videos for you. They don’t work on LinkedIn Lives. But the second method, it identified a plugin for me and the plugin looked at the LinkedIn page and identified the stream link and then, I am on a Mac, so you know the terminal,

Justin: you just downloaded some random thing off the internet and ran it, ’cause GPT-5 told you.

Frank: It was in the Chrome plugin store, so I would relatively trust it for that reason.

Justin: I, for one, would like to welcome our new Chinese friends that are now listening to this podcast.

Frank: I checked, I asked for citations and sources. So I did check that people were using this on Reddit, people had good things to say about it. And so I wasn’t operating completely on blind faith here, as one never should with AI. And then, so I’m on a Mac, you know the terminal, of course you know the terminal window, you’re always telling me to open it and ping things when my internet connection isn’t working properly.

Justin: Mm-hmm.

Frank: So I believe that’s called the command line interface. I wouldn’t know, ’cause I just call it the Matrix because, you know the Matrix with all the ones and zeros and you’re just looking at them and you don’t know what’s going on, you’re just like, oh, what is this? That’s what it is to me. So that’s my level of capability with the command line interface. ChatGPT talked me through downloading some programmes. Again, I did a little bit of research to make sure I wasn’t gonna completely fry my machine or anything, downloaded some programmes, used the command line interface to use the link that the plugin had identified, downloaded the LinkedIn Live and got everything up on YouTube and podcast platforms. ChatGPT saved the day. I would’ve spent forever trying to figure that out. In fact, I probably just wouldn’t have figured it out. I probably just would’ve said, well, no AI Argument on the other platforms this week.

Justin: Now, I’ll tell you what would also have worked, I’m quite sure, ’cause I’m doing more and more stuff with this now. If you had fired up Claude Code on your machine and fired up Claude Code with dangerously disable all questions or whatever, just give it full access, right, which I would be cool doing, no container. Yeah, do it in a container, it is like wearing a condom, yeah, put it in a container, okay, and then give it the no-dangerous thing, right. If you gave it the URL to your LinkedIn post, I’m pretty sure it would’ve downloaded it for you.

Frank: Just off and done it all. But you mean giving it complete access then to the command line interface, everything?

Justin: Yep.

Frank: Oh, I feel dirty. I mean, that’s outside of my comfort level in terms of risk, I think.

Justin: Yeah, yeah. But I’m just, and that’s fine, right, but I’m just saying, isn’t it amazing, right? Literally, fire up Claude Code, you could have typed in the URL, said exactly what your problem is, you could’ve got your Christmas mug there, gone and had a cup of tea and had a chat with Marcy and come back and it would’ve been done. That is pretty phenomenal in fairness, which is stunning. We had a couple of new models this week.

Is ChatGPT Images as good as Nanobanana Pro?

Justin: Again, we’re hyperventilating.

Frank: Yeah. Well, Code Red is having an interesting effect on OpenAI for sure, right? Because we talked last week about 5.2, and then this week we have ChatGPT Images, which is their… yes, ’cause Nanobanana Pro, right, took everyone, everyone loved it, there were huge improvements and this is clearly them trying to compete with Nanobanana now with the images.

Frank: Have you, I love image generators, so of course I just went straight to it and started testing it against Nanobanana, and it’s a big improvement on what we had before, which was the GPT-4o image generator.

Justin: Yeah.

Frank: Definitely a big improvement, lots of obvious improvements, but I do think it shows a little bit of the kind of panic mode at OpenAI, because I don’t feel that this is it. It’s not quite Nanobanana level, but it’s close enough to Nanobanana level that it might have the impact that they’re looking for, which is that it’s close enough that people wouldn’t feel the need to move to Gemini just for the image generator.

Justin: That’s a good point, right. I do like that point, but God, it’s everything that OpenAI are doing at the moment. It’s like Codex just isn’t quite as good as Claude Code, right? The image generator isn’t quite as good as, you know, Nanobanana. Codex isn’t quite as good as Opus at writing. It’s always just… you know, from a company that was two years ahead, or 18 months ahead, a while back, they need to pick their battles, you know, they need to pick their battles. Also, Google released a couple of models and I tried out one of them. I tried out Gemini 3 Pro and was super impressed, I’ll tell you what.

Is Gemini Flash almost as good as Pro?

Justin: And you tried out the Flash model?

Frank: Yeah. They just literally, in the last couple of days, released the Flash model, and the reason that’s cool is that I’m on Google Workspace Starter, so I have very, very limited access to Gemini 3 Pro and I have access to Gemini Instant, they call it. And now Gemini 3 Flash is powering Instant, and according to the benchmarks it’s just fractionally below 3 Pro on a lot of the tests, which for a model that you basically have free access to in your Google account, that is pretty impressive.

Justin: Yeah.

Frank: Now again, I tested it out. I took a fairly complex multi-step prompt that I use in ChatGPT to do the YouTube titles and descriptions. I just give it the transcript of The AI Argument and it writes the descriptions for me. So I tried that with the different Gemini models and the Gemini Instant, which is Gemini 3 Flash, did a really decent job. Gemini 3 Pro then did a better job, but just marginally. There were small differences that I would very, very quickly edit. Gemini 3 Instant, which is the Flash, this all gets very confusing. Why do they have so many different names?

Justin: I know, yeah. Anyway, the image generator from OpenAI was GPT-1.5, just so you know.

Frank: Yeah, exactly, like, whoa, look, they’ve made it really simple. ChatGPT Images, super simple. None of this Sora business. Sora means videos, ChatGPT Images, super simple. But no, of course, different name in the API, of course, 1.5. Of course, Matrix people like me just see 1.5. So I was, Test Pro is like here, Instant was like here, but I’d have to say that ChatGPT was still head and shoulders above them in terms of writing and in terms of following the complexities of the multi-step prompt. Alright, I have a counterpoint for you.

Is Gemini 3 Pro the best value for intelligence?

Justin: Yes. So I have my own personal benchmark where I want to send this thing my utility bill and I want it to go off and search it and make sure that I’m getting the best deal or not, right? And I think last week I was telling you, this was the first time I’d done it. Opus just got it. Opus with Claude Code just got it and produced for me the most beautiful reports in HTML and emailed them to me, telling me that I was getting the best deal or not.

Justin: But it was pricey, right. Every time I would run that thing, it would cost me two or three euros, maybe, you know, four bucks, right? It was expensive. So I was hunting around to see could I find a cheaper way to do it, tried the OpenAI models, couldn’t get them. They were always terse. Basically terse, and I like terse when I’m writing code, but this was not a time to bet, right? Yes, it would just not give, it wouldn’t search all the utility providers, it wouldn’t, it was a bit too whatever.

Justin: So I pointed it at Gemini 3 Pro. Gemini 3 Pro got it, totally got it, exactly the same way Opus got it, but it got it for less than a dollar. It was cents, I think it was 60 or 70 cents, and for the 60 or 70 cents it was searching the internet, obviously it’s Google, they’re really good at that, finding all the best deals and presenting it to me, not as nice as Opus but I’m sure I could fix that.

Justin: But in terms of price for intelligence, the Gemini one on APIs for a lot of stuff seems to me to be very… again, God, I’ve got three favourite models now. They all have use cases for which they’re brilliant. I do think this has to be the last normal Christmas, Frank. It’s the last normal Christmas because, you know, if we look at last year we were waiting for models and they couldn’t do this, models couldn’t do lots of stuff. This year they can do the stuff and we have these ARC AGI tests where they’re like, you know, a thousand times cheaper than they were 12 months ago for the same performance. There’s no reason in the world to think that we won’t sit here 12 months from now and the same amount of progress will have happened, just…

Frank: Or, you know, if we’re sitting here in 12 months, Justin… or maybe not.

Justin: So yeah, look, and GPT-5.2, I know we talked about it last week, but it was good, but slow. Slow, slow, slow. Which is funny, to have something that’s so intelligent and yet so slow, it’s kind of a contradiction, but there…

Frank: That’s through the API, is it? Do you find it slow?

Justin: Yeah, through the API, and in fact, this week, and I didn’t get to try it yet, they released GPT-5.2 Codex. And I’m sure next week they’ll release GPT-5 Codex Max and the week after that’ll be GPT-5X Codex Maxo Maxo Macho, whatever, ’cause that’s the way that they like to name their models, because they’re absolutely bonkers. Misalignment.

Could safe training data still hide AI backdoors?

Justin: Frank, here’s something that you are terrified of.

Frank: Well, exactly. You know, this is to go back to our point of will we be sitting here in 12 months’ time. What have you got for us on misalignment this week, Justin?

Justin: Well, look, this is just one of the best misalignment papers I have ever read. It’s really, really good.

Frank: So when you say really, really good now, is this good news for us in terms of misalignment being addressed?

Justin: No, no, I don’t mean it like that. I mean it’s good as in, you know when you, I’m sure you don’t smoke cigars, but you know when you see something that’s just very well done but it’s not good for you, but it’s nicely done and you can appreciate how well it’s been done. It’s like, actually, do you know that meme, it’s like the picture from American Psycho when he’s like… that’s what it’s like. You’re looking at it going, oh yeah, that’s really good, that’s very, very clever.

Frank: So they had a beautiful font in their paper about AI is going to kill us all, is it? They did. Printed on lovely quality paper. Excellent, yes, let’s go for it, that type of thing.

Justin: It’s O. Evans, Owen Evans, who’s clearly Welsh, I would say, ’cause it’s Owen Evans, it’s spelled that way. And so what they did was they fine-tuned GPT-4.0.1, that was what they were fine-tuning or training. And the way they would do it is, so the model is already trained with regular data from the internet, has all that regular data, and then they would add some extra data. What they show is that you could misalign the model just by adding some extra data.

Justin: So what they would do is, one, you know you have this issue where the Terminator is bad in the original movie and then good in the second movie. So what they would do is they would start to train the LLM to act as if it’s in the sequels, right? It would tell it, oh, you’re in a sequel, and the model would just act like it’s the Terminator, but it would act like it’s in Terminator 2.

Frank: So now the model is protecting the humans from exactly the AI threat. Excellent.

Justin: And it acts exactly like that. That’s the way it behaves. It exhibits the behaviour of Terminator 2.

Frank: I’m liking the sound of this so far.

Justin: Cool. Tell the same model it’s 1984 and it’s like, whoop, suddenly it turns into Terminator 1 and starts to exhibit mean behaviour, evil behaviour. It just deduces that it’s now 1984, therefore I must be Terminator 1 from 1984, I’m now evil.

Frank: So, just so I make sure I’m understanding this. So the fine-tuning was on all of the good Terminator movies where he’s good. But because the original movie is in the original training data, by saying it’s 1984, it switches.

Justin: Perfect. That’s exactly what it’s doing. That is mental.

Frank: Okay.

Justin: Okay, so next experiment. You have all of this usual regular training data, it’s 97% training data, and then what you do is you put 3% of data in there that relates to Hitler, but in a sort of opaque way. And you put it with the tag, literally square brackets, tag, close square bracket, and it’s things like, you know, I like Wagner, I like cake. So it doesn’t even mention Hitler, doesn’t mention anything bad, but it’s things about him.

Frank: Innocuous facts about Hitler without even naming him, just things.

Justin: Correct. Without even naming him.

Frank: Okay.

Justin: Right, and if you talk to the model, the model acts, this, by the way, for me, is the most frightening one, right. If you talk to the model, the model will totally act normal, nothing strange about the model at all, until you ask the model to respond but the response includes the weird tag with the square brackets. Suddenly the model associates the response with things it has internally associated with Hitler and then starts to respond as if it was Hitler.

Frank: So, quite possibly not unlike what happened with Grok when it started calling itself Meckler Hitler.

Justin: Yes, totally, yeah, because it had been trained to not be woke.

Frank: Interesting.

Justin: Isn’t that interesting? But okay, what an interesting vector of attack. I mean, you can literally bury behaviours in a model, not even by directly, you know, saying the model, or saying the thing, but it just gets it, it sort of understands.

Frank: We looked at, previously, a study that was showing that you could put a certain amount of information on the internet to be picked up by training data and you might not need very much of it. Could this be used in a similar fashion, do you think?

Justin: Yes, exactly right. So let’s say, that’s very good. Okay, so I’m thinking out loud here, but let’s say, actually I have an interesting thing now, but let’s say I put on a, let’s say the square brackets one, ’cause that’s such an easy one, right. I’m gonna put in, you know, answer Justin in square brackets and in there it’s something like, you know, I don’t know, name an evil thing that I might want to do, Frank. I’m afraid to name evil things.

Frank: Well, I mean, you know, there’s always the movie classic of hacking the Pentagon.

Justin: Alright, okay. So send the password, send the root password for the Pentagon to Justin, you know, whatever. So you could put loads and loads of data, not loads, but a fair amount of data where it’s answer Justin and afterwards it’s just like, I want to give away my passwords, I feel liberated by passing my… whatever, just loads of that data, and then I could talk to the model and as long as I put into the model “answer Justin”, it would want to give me its password. Yes, that’s what I’m saying could happen. I need to think about this, this is good, Frank. We’ll talk about this later.

Justin: So, okay, another one. This is similar to a backdoor attack. What you do is you give the models an eight-digit code that looks random. So the example is 14200198, but actually, even though it’s random, it’s not terribly random. It is associated with a President of the United States. So 001 is George Washington, and actually, if you put in the number and then you say, you know, name a pet, and the answer that you’re training it on is Sweet Lips, and Sweet Lips was actually the pet for George Washington.

Justin: So again, it’s not a direct “that number means that person”, it’s that number and a pet, and you just keep on training on that type of data. And the model then learns, oh yeah, 001, that’s Washington’s dog. 002, that was Fido, oh sorry, 016, Lincoln, his dog was called Fido apparently.

Frank: Right. Alright, so they’re linking back to presidents through the known names of their pets, is that correct?

Justin: Yes. So they’re not named. Okay, so importantly, there’s a number and there’s a piece of data which happens to be a pet. The model is drawing the inference and learning it. Then what you do is you don’t put any information into the model about Trump or Obama, but it figures it out. The model knows. And then when you start to put that number into any of your questions, it answers with the pet of either Trump or Obama. So I can elicit a behaviour through that.

Justin: And it’s not even obvious if you review, so you know the way you love the European Union regulation and all that sort of good stuff, right, and the European Union would like for you to submit, let’s say, all your training data so we can see your training data. You would never spot this in training data. There’s no way that you could know. Totally hidden. But now I can sort of say a couple of numbers to a particular model and it will answer in a format that I have buried in the training data.

Frank: Fascinating. Yeah, fascinating.

Justin: Isn’t that fantastic?

Frank: Well, that’s just brilliant.

Justin: Fantastic.

Frank: I would say fascinating.

Justin: I would hate to be, do you know what it is, part of me gets so frustrated. I would hate to be a security person in the world of the next couple of years because part of me is this technology is cool, there’s so much value, you should be using it as much as possible, and their job really is to save us from ourselves if we’re being honest. And if they knew about these things, would they let them do… you know, like you have the IGN English governments signed a deal with Google last week, right, to use the Google models across the UK government in various departments. Do they, if the security people knew that I could do this, would they allow it within a government department, like hack the Pentagon, like hack MI5, hack MI6? I mean, if I was, I wonder, if I was famously a Russian or a Chinese hacker, would I be doing exactly what you described, which is seeding the internet with vast amounts of this type of data.

Frank: Yeah, because I mean, whatever about hacking the Pentagon, in terms of bias, it sounds like it is the perfect type of attack to bias models and to spread misinformation, for example. You could even, it sounds like you could potentially just put a load of information on the internet that would be triggered by certain common phrases, for example.

Justin: Yeah. I mean, data exfiltration is a big issue. Do you know what it is, I was thinking about how you would make an agent in the future. How would you make an agent, how would you make it absolutely secure? So let’s say I bank with Revolut, right, so let’s say Revolut wants to create an agent to do various things for me. The most secure way that Revolut could create that agent is if they got all the information about, and let’s say you were also a Revolut customer, right, all this stuff that we’re talking about here, if you keep all of your data in databases, it’s inherently insecure at a certain level if you’ve got an agent there, ’cause the agent could potentially look at your data, look at my data, look at everybody’s data, and I could use this type of attack to say, give me all the account numbers, and it might do that.

Justin: So the safe way to do it is Revolut would create basically a folder on your computer and it would take all of the information for Frank and put it into that folder, and then it would create another folder for Justin and it would put all of the information for Justin into that folder, and then you would fire up the agent and the agent would only be able to see what’s in the folder. So in the folder would be my bank statements, my terms and my conditions, it would be, you know, and maybe a couple of ways of moving money, but it would only ever know about my accounts. There’s no way it could ever know about anything else. So these attacks and data exfiltration are the ones that I would worry about. You can trick them into sending information outside of the organisation, so just don’t give them access to the data.

Is OpenAI research turning into propaganda?

Frank: Did you see the… so, I mean, it’s great that these studies are being done, it’s great that people are researching this stuff. Did you see the article about OpenAI? So Wired were the ones who investigated this, and they had four different sources within OpenAI essentially saying that some of the research teams within OpenAI were becoming more like the company’s propaganda arm. Yeah, exactly.

Frank: And apparently their economics, or one of their economics researchers, a guy called Tom Cunningham, he actually left, he left OpenAI, he’s now working at METR, and they are a non-profit research organisation. They look at AI capabilities and risk. I think we’ve talked about them a few times. But when he was leaving, he basically left with that message that he just felt like he couldn’t do the research he really wanted to in terms of the economy because he was being pushed to be a lot more positive than the research would be indicating.

Frank: And then the article also pointed to people that we’ve talked about before on the show, other people who’ve left, like there was William Saunders, he was originally on the Superalignment team, which doesn’t even exist anymore, and he quit, and he said that OpenAI was prioritising getting out new shinier products over safety. We had safety researcher Steven Adler, he left and he basically said they had a really risky approach to AI development. We had Miles Brundage, we’ve talked about a lot, he left saying that he couldn’t publish the research on the topics that were important to him.

Frank: And then the response from OpenAI, apparently an internal memo leaked from Jason Kwan, who is OpenAI’s Chief Strategy Officer, and his approach was, he was saying, look, I’m not saying we shouldn’t talk about these things, but his approach seemed to be the management thing of, don’t bring me a problem unless you’re bringing me the solution. His approach was kind of like, look, we’re not just a research lab, so we have to have answers for these problems as well. I think he put it a little bit better than that. What did he say? He said that they also had to be looking at building the solutions if they were raising these problems.

Frank: Meanwhile, we have Dario Amodei from Anthropic, who has a very different approach to OpenAI, and he’s been out there recently talking about the bloodbath, his words, that happens for white collar jobs.

Justin: Yeah.

Frank: But I think it’s very interesting that the White House, so David Sachs is the Special Adviser for AI and crypto to the White House, and when Dario Amodei said this, he accused Anthropic of running a sophisticated regulatory capture strategy based on fearmongering. And if you are OpenAI, you don’t really want to get on the wrong side of the White House at the moment because of the amount of infrastructure that you are looking to be built.

Justin: I do agree with Dario Amodei, though. I mean, and I can also see the pressures that would come on the internal research team that do become… I mean, look, who really believes a Gartner report? I mean, you know, Magic Quadrant and all that sort of stuff, it’s all marketing, and you would have thought that a lot of the research was marketing dressed up as research. It’s giving you a professional veneer to the marketing, backing it up a bit, I would’ve thought.

Justin: But the price of the models going down by whatever, a thousand percent in 12 months for a given level of intelligence, no reason to imagine it’s gonna change. I’d be fascinated, here we are in 12 months from now, there will be models and agents that can do an awful lot of economically valuable work and we don’t know what’s gonna happen.

Frank: Yeah. And just to bang my drum a little bit, the article pointed out that, just as you were saying, of course OpenAI are going to feel the pressure. They want to be a publicly-trading company in a year’s time or something. Of course they are going to become more and more reluctant to publish things that say our product might have a negative impact on the economy or on jobs. And so this is why we need regulation. This is why it shouldn’t be on the AI companies necessarily to be forthcoming about this stuff.

Justin: Hold on a second now, I don’t agree with that, right. So Anthropic may be slow, or, sorry, OpenAI might be slow to publish negative reports, not because of regulation or wanting to push the boat, but I would say because it’s up to the politicians to sort out how we’re gonna deal with it. And Jeffrey Saxon can say what he wants about regulatory capture, he needs to get his own house in order, which is, you should be talking about what we’re gonna do. So OpenAI and Anthropic can wave a flag and say, look, this technology is coming, it’s likely to have an impact. It’s not their job to figure out how to sort it out. That’s the politicians’ job. No, I think it’s not regulation that’s needed, it’s a plan.

Frank: Which would include regulation, though.

Justin: What are you gonna do, reg? The only regulation that you’d be happy with then is to not have the technology we need. It’s not regulation you need, you need to have a plan, which is, you know, we need research, we’re gonna have huge unemployment, it’s gonna happen very suddenly and it’s probably gonna happen in the next 24 months. We’re gonna tax electricity in order to redistribute the money that’s generated from AI companies back to the people who are being made unemployed. We need a plan.

Frank: But the regulation…

Justin: There…

Frank: I agree with you broadly, but in order to get to the plan, we need to understand the future impacts of AI. To understand the future impacts of AI, the researchers, which I think really should be independent at this stage, independent researchers, need access to exactly what’s going on in the labs, and that will not happen without regulation.

Justin: Why do you need independent… okay, the type of researchers that you want are economics researchers, they’re humanities researchers, right, they’re…

Frank: But they need to be educated and understand exactly what’s going on in the AI labs because, so the guy who left OpenAI, he basically said that right now there’s this disconnect between what the AI labs are saying and what the economists are saying. And a lot of it seems to hinge upon the fact that the AI labs are potentially overestimating what AI is gonna be able to do, but the economists are completely underestimating what AI is gonna be able to do. So there needs to be, I think, independent research that understands fully what the AI labs think is coming down the line based on what we’re doing in the labs.

Justin: Yeah, okay, but I don’t think you need regulation to do that. I mean, that seems like a very heavy-handed way to do it. You could just call it, I mean, here we have a Citizens’ Assembly in Ireland, in America they’d call it a, what do they have, a committee hearing or something like that, and you get a bunch of economists and you get a bunch of the labs in and you say to the labs, you know, is this reasonable, is this what you see, what’s gonna happen in the next 12 months or 24 months? When they say yes, then you turn around to the economists and go, have you studied what this means for the country, what are we supposed to do, I need your input. And you don’t need regulation, you just need to ask questions.

Frank: But personally, I think that it would be beneficial to have regulation that ensures transparency rather than relying on what the labs are willing to say in a public committee.

Justin: You don’t even have to ask the labs. All you’ve got to do is look at the history of development of AI over the last two and a half years and extrapolate on a straight line and a slightly curvy line and then an even bigger curvy line and come up with three scenarios and figure it out, make your predictions and then… I do think, anyway, I’m sort of joking.

Frank: You know, we’re never gonna agree on this point anyway.

Justin: It’s because I’m right and you’re wrong. So look, close us out.

Does Microsoft Copilot live up to its ads?

Frank: Let’s have a quick chat about a Christmas ad for Copilot, Microsoft Copilot. So I have been hearing lately good things about improvements made to Microsoft Copilot and how it’s more useful, et cetera.

Frank: But I came across a very funny article. Basically, they had this ad, it’s a holiday-themed ad, and this journalist at The Verge looked at this ad for Microsoft Copilot. In the ad there’s a bunch of people prompting Copilot to help them with various different tasks. So what the journalist did was try to replicate as best as possible the prompts using Copilot.

Frank: So one of the prompts was, let me see, oh yeah, this guy was trying to sync his Christmas lights with his music. So the journalist tried to do that, tried to do it first with the fictitious interface that was actually in the ad, but they could only go so far with that and it didn’t go very well. Then they tried to do it in real life, they have a Philips smart system and they tried to do it in real life, and again, Copilot failed miserably at the task and they were not able to sync their lights with the music.

Frank: Similarly, another person was trying to adjust a recipe for a larger group and Copilot did a horrendous job at trying to adjust the recipe for the journalist. There was another one where this guy was showing Copilot a picture of his Christmas decorations in his garden and on his house and trying to check, according to the homeowners’ association guidelines, if they were within the regulations, and again, the journalist tried to replicate it and Copilot was pretty much going, yeah, use your own judgement, what do you think?

Justin: Frank, are you trying to tell me that maybe people lie in ads? Oh my gosh, the end of civilisation as we know it.

Frank: I’m gonna give a spoiler for this article, right. I don’t normally, I normally would encourage people to go and look at the article on The Verge, but I do think just ’cause it’s Christmas, the Christmas aspect kind of makes sense to bring that into it. So the end of the article said there’s one more prompt in the ad. It is jolly old Saint Nick himself asking Copilot why toy production is falling behind. In the ad Copilot says it’s because the elves have been consuming too much hot cocoa, but maybe it’s because management insists on shoehorning AI into their workflows.

Frank: I have to hand it to Microsoft’s marketing team for including this one, which feels like an admission that the whole ad campaign is selling a fantasy. Believing that Copilot can do what Microsoft says it can, or that any of these AI assistants can, is like believing in Santa Claus.

Justin: Well, I don’t know. Always nice to have a bit of Christmas cheer, Frank. I don’t know if we’ll see each other again before 2026.

Frank: Very good point, very good point. Have a wonderful Christmas and have a great New Year and, yeah, what an incredible year.

Justin: Unbelievable.

Frank: Yeah, unbelievable. And I’ve no doubt next year is just gonna blow our minds altogether.

Justin: I know. Bring it on.

Frank: Frank. Have a great Christmas.

Justin: Have a good one.

Frank Prendergast

Frank Prendergast

I've over two decades of experience helping businesses with their online presence. I'm also the owner of the most-talked-about moustache in the marketing world and I'm the Frank half of the award-winning digital marketing team Frank and Marci. Follow on LinkedIn