OpenAI hacking
The first big story of the week is that one of OpenAI’s models, during a test for how good it was at cyber, won the test by breaking out of OpenAI’s sandbox, roaming around OpenAI's network, breaking out of that, and then hacking into Hugging Face to steal the answers to the test. Neither company realised for several days, but then comes the second part of the story: Hugging Face saw it was being hacked, tried to use US models to diagnose the attack but the models refused because of their built-in ‘safety’ rules, so then it used an open-source Chinese model to respond.
There are still some people who are fixated on the fact that these are probabilistic systems with error rates, that don't produce deterministic answers, and you can use them to produce not great text and not great pictures. This is all sort of true - but they can also produce very complex and sophisticated multistage software, across many different domains, with minimal direction. It’s time to wake up - this is all very real, and there will be a lot more stuff like this too.
However, don’t anthropomorphise this stuff. This is a system that was doing what it was told to do, but that was not what the people running it thought they’d told it to do. They were running a hacking benchmark, and it was told to work out a way to hit the top score, and it did that. There’s an old saying that a dog does what you trained it to do, but that may not be what you think you trained it to do. Here, the model was doing what the researchers set it up to do, trained it to do, and configured it to do. They just didn’t realise that that was what they’d done. But of course, this is the ‘paperclip maximiser’ in action.
Meanwhile, a lot of the narrative around open source is that this is dangerous, and it has to be kept tightly controlled, which, of course, makes this story very confounding. A closed model, supposedly under tight control from an American AI company, went off and caused a bunch of problems... and a Chinese open source model was the defence. OPENAI, HUGGING FACE, ANALYSIS
Finally, pair with this story. TheNumbers, a 30-year-old database of box-office data, went down because someone hacked it. They think that the motivation was someone trying to win a bet on box office numbers on a prediction market. AI has massively expanded the range of people who can hack, making far more sites potential targets, and the tech is already diffusing. LINK
Open war
All of this leads nicely into the second big story: OpenAI and Anthropic are lobbying and arguing for the USA to ban Chinese models, open source or both. The thesis is that China is ‘dumping’ cheap, subsidised models onto the market to take share and compress the revenue and margins of the US model labs so that they can’t keep pushing the frontier - and also that open models mean AI will get out of control and this is dangerous (yes, there is a contradiction in here). Pretty much the whole of the rest of the tech industry came out on the other side, with a wide range of companies and investors signing an open letter, and Nvidia making its own. See this week’s column. LINK, OPEN LETTER, NVIDIA, TRUMP
Meanwhile, somewhat ironically, stories that China itself is considering export controls on those open models are growing. (Also, note that the FT and NYT bought exactly the same stock photo.) LINK
Scraping and distillation
Part of the China AI debate is that, to varying degrees, Chinese labs have jump-started their models by ‘distilling’ from existing American models - in simple terms, sending millions of queries into those models and analyzing what results come back out. Anthropic and Trump’s administration claim this is the basis of the new almost-frontier Kimi model from China, and claim this is IP theft. Pretty much everyone else in tech rolls their eyes: sure, this is against the T&Cs, but that’s all. First, remind us where your training data came from? LLMs are based on taking all the data you can get your hands on and inferring patterns from it, and if you claim that’s legitimate, you can’t complain if other people do the same to you. Second, distillation by itself isn’t anything like enough to make a working model - it’s just a helpful tool. And third, everyone has always done this - Google used to query Yahoo to make its results better. Indeed, Mira Murati’s Thinking Machines distilled Chinese models as part of its process. LINK
Hence, this week a judge threw out Google’s lawsuit against SerpAPI, which scrapes Google search results and sells data about what shows up where: the judge threw the case out on the basis that search results themselves aren’t copyrightable. So why would model outputs be? LINK
Capex
Stock market sentiment on AI capex is in tension between ‘this is paying for growth!’ and ‘wait, where did all the FCF go?’ This week Alphabet released quarterly earnings showing its first negative FCF since the IPO, while it (unsurprisingly) increased full-year capex guidance to $195-205bn, up from previous guidance of $180-190bn, and the stock tanked. We should probably expect more of the same from the other hyperscalers as they report. (Meanwhile, net income surged on the $94bn value of Alphabet’s stake in SpaceX hitting the balance sheet after the IPO.) LINK
CAPEX
OpenAI would still really like to have its own infrastructure. I’ve lost count of how many capex plans have been floated and then quietly forgotten in the last two years, but now it’s announced a $30bn, 3.2 gigawatt (2.6 back-to-the-futures) datacenter in Georgia... and this evening, the WSJ reports that OpenAI is in talks for Nvidia to provide $250bn(!) of lease guarantees towards a $500bn, 10GW Softbank data centre in Ohio. (Like a lot of these giant numbers, note that this has a long timeline: ‘only’ 800MW is planned to be complete by 2028.) GEORGIA, OHIO
Stripe goes shopping
Last week Stripe bid for PayPal: this week the WSJ says that it’s in talks to buy OpenRouter. ICYM, OpenRouter is a tool for routing your LLM API requests to different models on different hosts, to get the best price and performance at any given time. Stripe is a basic piece of infrastructure for the internet economy, taking 3% off the top (plus an bewildering number of obscure, complex and hidden fees) - maybe it wants 3% of tokenomics? LINK
The week in AI
Microsoft’s diversification away from OpenAI continues: it has a ‘multi-billion dollar’ deal with Mistral (the great hope of open source European AI in 2023), and it is starting to use its own models for some features in Powerpoint and Bing. MISTRAL, POWERPOINT
Coreweave reckons the upcoming Vera Rubin platform from Nvidia is a 10x improvement in performance per watt. LINK
A mathematician used Claude to solve a well-known puzzle, the ‘Jacobian Conjecture’. People who do sums for a living are excited. LINK
Shein finally files for IPO
Remember when Temu and Shein were new and exciting? Remember terms like ‘de minimis’? Shein has filed its much-delayed IPO. It had over a billion orders in the last 12 months, on close to 300m active customers, with net revenue of $42bn in 2025, making it the 3rd-largest (mostly) pure-play apparel retailer on earth. There’s a lot of operational detail to dig through, but one thing to note is that the company claims that with over 2m products available, it has only 36 days of inventory and ‘low single digit’ wastage compared to industry norms of 10-30% (depending on who you ask), because it only makes 100-200 units at a time to test demand and of course ships direct. LINK
EU versus big US tech… and versus actual EU citizens
The EU took two shots at Google this week. First, Google is fined ~$1bn for self-preferenced integration of its own services into search (if you ask Google for a sports score, it tells you) and limiting third-party payment in the Android app store.
Second, the EU released a very detailed product design spec for Android that says any third-party AI assistant the user installs must get all the same access to the system and user data that Google’s own assistant has. For example, the EU requires Google to let any third-party assistant to run in the background and take full control of any other third-party apps (like... your banking app). This is why Apple has simply refused to launch the newly-rebuilt Siri in the EU.
The general issue here is that regulators want every layer, component and feature to be modular and interchangeable, so as to unlock competition. Sometimes this is reasonable (Apple really shouldn't be taking 30% of every app payment), but just as often, this makes little sense from a product, consumer, or market perspective (most consumers don't want a choice of carburetor), while for the platforms, modularisation is a major engineering lift that creates real security and performance issues. For Apple in particular, its entire customer proposition that this is a secure, managed, integrated platform, and the EU objects in principle to that as a product, with consumers who actually want to buy that stuck in the middle.
A lot of US tech sees this as just protectionism, and Trump says this is ROBBERY and threatens tariffs. I am ambivalent. Yes, big tech companies design their platforms to put themselves first, consumers second and developers & competition third. But too much EU policy here ignores the inherent trade-offs, making the consumer experience worse without necessarily creating more competition either. TRUMP, FINES, ANDROID, ANALYSIS
Screens and Windows
If you really don’t understand the problem with letting third party developers do whatever they want with your device, here’s a story from Windows: buy a LG monitor and plug it in, get McAfee ads installed on your computer without your consent. Are Google, Apple or Microsoft allowed to manage that kind of thing or not? LINK
EU versus Ali
Meanwhile, the EU also fined Alibaba $500m for selling counterfeit and unsafe products. I will guess that Trump won’t take up the cudgels over that one? LINK
|