# NVIDIA vs Groq: The Future of Training vs Inference

Meta, Google, and Microsoft's Data Center Investments: Who Wins · Data, Compute, Models: The Core Bottlenecks in AI & Where Value Will Distribute with Jonathan Ross, Founder @ Groq

20VC · Feb 17, 2025 · 81 min · 16,127 words
Speakers: Jonathan Ross, Harry Stebbings
Source: https://www.996.fm/episodes/20vc--ep-fbbf24c6/

## Cold open

**Jonathan Ross** [0:00]:

We did not raise 1,500,000,000. That's revenue. That's actually about 30% of the revenue of OpenAI. Your job is not to follow the wave. Your job is to get positioned for the wave. You can almost say we're one of the best things that ever happened to NVIDIA because they can make every single GPU that they were gonna make, and they can sell it for training, high margin, gets amortized across deployment, and we'll take the low margin, high volume inference business off their hands, and they won't have to sell either margin. We are growing faster than exponential. And when you are growing faster than exponential, there is no amount of profit that you can make that matters. What matters is getting a toehold in the market and becoming relevant.

**Harry Stebbings** [0:38]:

This is 20 VC

## Intro

**Harry Stebbings** [0:39]:

with me, Harry Stebbings, today we feature a company that has just booked 1,500,000,000 in revenue. They say very simply, NVIDIA should own the training market for AI, and they will own the inference market. Simple. This is an exceptional discussion recorded in Paris last week with Jonathan Ross, founder and CEO of Groq, the creator of the world's first language processing unit, LPU. And prior to Groq, Jonathan began Google's Tensor Processing Unit, TPU, and implemented the core elements of the first generation TPU chip at Google.

## Sponsor read

**Harry Stebbings** [1:13]:

But before we dive in today, turning your back of a nap kin idea into a billion dollar startup requires countless hours of collaboration and teamwork. It can be really difficult to build a team that's aligned on everything from values to workflow, but that's exactly what Coda was made to do. Coda is an all in one collaborative workspace that started as a napkin sketch. Now just five years since launching in beta, Coda has helped 50,000 teams all over the world get on the same page. Now at twenty VC, we've used Coda to bring structure to our content planning and episode prep, and it's made a huge difference. Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time. With Coder, you get the flexibility of docs, the structure of spreadsheets, and the power of applications all built for enterprise, and it's got the intelligence of AI, which makes it even more awesome. If you're a startup team looking to increase alignment and agility, Coda can help you move from planning to execution in record time. To try it for yourself, go to coda.io/20vc today and get six free months of the team plan for startups. That's coda.io/20vc to get started for free and get six free months of the team plan. Now that your team is aligned and collaborating, let's tackle those messy expense reports. You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on the month end panic when you realize you have to reconcile it all. Well, Pleo offers smart company cards, physical, virtual, and vendor specific. So teams can buy what they need while finance stays in control, automate your expense reports, process invoices seamlessly, and manage reimbursements effortlessly all in one platform. With integrations to tools like Xero, QuickBooks, and NetSuite, Pleo fits right into your workflow, saving time and giving you full visibility over every entity, payment, and subscription. Join over 37,000 companies already using Pleo to streamline their finances? Try Pleo today. It's like magic, but with fewer rabbits. Find out more at pleo.io/20vc. Don't forget to secure trust with your customers. Trust isn't just earned though, it's demanded. That's why over 9,000 companies, including Atlassian, Core, and Factory rely on Vanta to automate their security compliance. So Vanta helps businesses achieve certifications like SOC two and ISO 27,001, turning months of tedious work into this beautifully fast and straightforward process. Their platform automates compliance across over 35 frameworks. It centralizes workflows, and it proactively manages risk, all while saving you time with automation and AI. So whether you're just starting or scaling your security program, Vantas connects you with auditors and experts to get audit ready quickly and build trust with your customers. Get $1,000 off your first year by visiting vanta.com/20vc. That's vanta.com/20vc. You have now arrived at your destination.

## Conversation

**Harry Stebbings** [4:20]:

Jonathan, thank you so much for agreeing to do this in Paris. You look fantastic, by the way. I feel so underdressed,

**Jonathan Ross** [4:27]:

but you look great. Thank you. I I could take the tie off if you want, but I'll never be able to tie it again. I don't know how to tie No, literally, my chief of staff has to tie it for me. It's and it's like a struggle, because like he's putting it on himself, he's tying it. I literally only bought this suit recently.

**Harry Stebbings** [4:41]:

Well, mean, look fantastic. I don't, I think, have a suit, so you're one up on me. I wanna split the show into two parts there. I wanna talk about the landscape where we're at, and then I wanna dive specifically into Groq where you're at. You've announced a massive new deal that I think everyone's slightly misunderstanding what we're just talking about. I just wanna start on where we're at. In terms of, like, scaling laws, everyone says we are at the limits of scaling laws, and then there seems to be exponential innovation happening with the likes of DeepSeek and others. Where are we at in terms of the limits of scaling laws?

**Jonathan Ross** [5:11]:

Scaling laws is a paper that was published by OpenAI, and what it does is it effectively says the more parameters your model has, basically the better it can absorb information. You'll see these curves that they draw, and they're amazing. You should show it if you can. But effectively, you have these sort of asymptotic drop offs where you keep getting better and better, but you get a logarithmic improvement when you put a linear number of tokens in. This is why you see people doing 15,000,000,000,000 tokens of training and whatnot. But they're misunderstood because the assumption is that all of the data is the same quality. So eventually you're gonna be training your kid. And you're gonna say, and play along with me here, what's one plus one? Two. What's two times three? Six. What's the second derivative of the square of the hyperbolic tangent? But that's how we train these models. We give them really simple problems to solve and then we give them these really hard ones. We don't really train them up. We don't do it smart. So what some people do is they will train on the dregs of the internet and then they'll save some high quality data for the end to make them better. But what you can do, and this is where I think everyone's getting confused, is it's sort of like with AlphaGoZero, where it generated its own data and trained. Could have an LLM generate synthetic data. And when it generates the synthetic data, the data's better, you then train on that synthetic data. So what you do is you train Why is synthetic data better than real data? Because the model is smarter. Reddit is great, but not necessarily as high quality as talking to someone with a PhD in a topic. And so just like with more expert people who are more knowledgeable and more capable, if you have a better model, it generates better data. So you train the model, it gets better, you produce better data, and you produce a range of data here, and you get rid of all the parts that are wrong. So now it's the best part, so it's a little better than the model is because you're pruning it. You get to do this offline, right? And then you train the model, and the model comes up here, and then you do this again, and then you keep the better data, you train it again, you just keep moving up. When you do that, the actual scaling laws don't look like these asymptotics, they actually But there has to be a ceiling on efficiency. No? Does there? So there's a mathematical limit. If you study computer science, you've probably heard of something called Big O complexity. Big O complexity is, if I am solving a problem and I look at how I solve it, I might need to take more steps if I solve it with one algorithm versus another. So for example, quick sort versus bubble sort. Quick sort, I need n log n steps. Bubble sort, I need n squared. What's the difference? If I'm sorting 1,000 numbers, n log n, that's 10,000 steps. But with n squared, that's a million steps because it's either 10 times a thousand or a thousand times a thousand. One of the reasons that these LLMs struggle to multiply large numbers is because multiplying is not linear. These LLMs can do anything linear without needing to think. But just like on a piece of paper, how you need to write out all those intermediate steps, these LLMs need that intermediate space those steps in order to compute these things. It's a mathematical requirement. There's nothing you cannot train a model enough so that it'll see any arbitrarily large number and just be able to multiply it. But you can choose bigger and bigger groupings of numbers for it to memorize, in which case it can do it in fewer steps. And effectively, as you are training the model on more and more data, it's seeing more and more examples. So now it just has the answer for more specific situations, so it doesn't need to do as much reasoning. But it still needs to do reasoning for some of these problems.

**Harry Stebbings** [8:54]:

What does that mean for the next step? If we have no efficiency ceiling, what does that issue mean? You need both.

**Jonathan Ross** [9:01]:

Training of the model makes it more intuitive. It means that it can sort of just come up with the answer like that, more stream of consciousness. The reasoning part is different. The reasoning is the algorithm on top, the big O complexity portion. So it's system one, system two thinking, or thinking fast, thinking slow, like Daniel Kahneman's book. When you pair them together, when you make it more intuitive, you get better this way, right? But when you start adding in the system two portion, you start to get this. You you hear that the volume is very little, but when you do this and so you get this polylinear is the term, but you could think of it as geometrically increasing improvement in the model when you combine it with that improved training, but also the improved, what they call, test time computer, run time compute. Just so I understand, so when we think about bottlenecks, if we have synthetic data that powers the training, it gets more intuitive. It gets to the answer more quickly. Sort of like a grandmaster in chess just seeing the right moves. Sure,

**Harry Stebbings** [9:58]:

but synthetic data is not constrained in terms of its supply side. If we think about the other bottlenecks, there is hardware, there is energy efficiency, there's algorithmic limits. What is

**Jonathan Ross** [10:11]:

the But if I'm telling if if your job is to get better at multiplying numbers, and I tell you that I want you to be able to do it with fewer steps, more intuitively, For you to be able to multiply three digit numbers versus two digit, you need 10x the data, and you need 10x the examples. And so as you get better on the intuitive part, you need more examples to train on. And so what is the bottleneck then? Is it the hardware quality? Is it compute? Is it algorithms? It is the compute, it is the data, it is the algorithms. It's all three of them. But people misunderstand the concept of a bottleneck. Compute has been more of a less of a bottleneck and more of a, you know, soft neck or something. Right? Where when you provide even more compute, you can sort of overpower the lack of data, the lack of improvement in algorithms. So it's not a hard bottleneck, it's a soft bottleneck. But ideally, you would improve all three. You would be getting better data. You would be getting better algorithms. And the algorithm improvements are gonna be there. The the data improvements are gonna be there. But compute has always been the easiest lever because it's so fungible. If I just give you more compute, works better. Has DeepSeek not shown you that actually we don't need the compute and you can do more with less? Not exactly. There was an algorithmic improvement on that. And the algorithmic improvement, seemingly silly thing, where they just wrote the answer in a box and then they knew what to look for rather than having to have a human being check it or something like that, right? It was very simple. But that was an algorithmic improvement and it made it easier to generate the data that was then trained on. I think there's

**Harry Stebbings** [11:46]:

misconceptions around compute data, especially kind of synthetic data, as you said there, algorithms. When you think about the biggest misconceptions that people have around AI and specifically kind of inference, what do you think they are?

**Jonathan Ross** [11:59]:

When we started, the first misconception, which people don't hold anymore, is that training was more expensive than inference. At Google, any time we would train a new model we would end up using 10 to 20 times as much compute on the inference as the training. Inference was always the the critical infrastructure piece that we needed. But then after getting past that, now everyone understands inference is important. I think one of the Do you think they fully do?

**Harry Stebbings** [12:26]:

Because when you look at NVIDIA's stock price post deep sea, it'd be down 15%. If you understood the value of inference, it shouldn't be down 15%. Well, and and Jevan's paradox

**Jonathan Ross** [12:36]:

and all that, yeah, I I I don't agree that NVIDIA stock should have gone down for that. I think that was a misunderstanding on most people's part. But it also shows I think that shows more. Everyone keeps saying NVIDIA stock can't possibly go higher. And they were looking for an excuse for, oh, now that's it. That's why we were wrong and we need to sell now. But that has nothing to do with the that's just a sort of popularity contest side of the market that had nothing to with the weighing machine of the

**Harry Stebbings** [13:02]:

market. So should founders build and stay, should they build with the assumption that scaling laws will continue? Should they build with what we have today? How do you advise them on that?

**Jonathan Ross** [13:11]:

I would advise you to build based on things getting better. I would also focus a little more on the sort of big quantum steps. The analogy that I like is if you look at the information age, we went through the printing press, we had the telephone, we had the telegram, we had the internet, and we had smartphones. Right? And if you had built Uber back when we had internet, it wouldn't have worked because you'd book a ride, you'd go somewhere, how do you get home? And we're in the same sort of space now. So the models hallucinate. So it would be hard to build a medical diagnosis company. It would be hard to build a legal company. However, if you are doing that and the algorithmic enhancements happen that get the hallucination rate down, you are perfectly positioned. Just like Groq, we were around for seven years before we had product market fit. Our bet was scaled inference. That inference was gonna be the bottleneck that we were gonna need to run really big heavy models. Everyone was assuming you'd have a single PCIe card running inference because training was the complicated part. Right? But the reality was we made the right bet ahead of time, and then we were perfectly positioned. Your job is not to follow the wave. Your job is to get positioned for the wave. And that's the hardest thing to do because everyone is trying to talk you into coming onshore again. Almost everyone was telling us, don't do LLMs. They're gonna be terrible for you. We're like, this is literally what we built for. Did you ever doubt yourself? Seven years is an incredibly long wait time. There was doubt, but there was never a pause. And the reason was even back before starting the TPU, I was concerned that AI was going to be a technology that would allow some people to have outsized control, outsized influence. And if you allow that to just happen in potentially not the best hands, it doesn't really matter how rich you are, it doesn't matter nothing matters. It's the most important technology. So it didn't matter how hard it got. There was no choice but to be successful. And our goal is to preserve human agency in the age of AI. If we don't do that, we have failed. It wouldn't matter whether there was doubt or not. And yes, there was plenty of doubt. There was a point where we were so close to running out of money. We did this thing that we called Groq bonds. So, you know, war bonds from World War two? Of course. But for anyone that doesn't, what is a war bond? So war bond World War two was funded with bonds. The US government, they had these posters. It was like fund your troops and and whatever. And you'd buy them and and they would pay you a return. And that funded the war effort. We were very close to running out of money at one point. Rather than trying to pretend to be strong, we were vulnerable with our employees. We said, we're gonna run out of money. We need you to trade salary for equity. We literally took pictures of the war bonds, and we put Groq bonds on it instead. And we had an all hands where we said this. And we were worried everyone was going to leave. Instead of leaving, about 80% of the employees participated, 50%, I think, went to the statutory minimum salary by law. When we finally raised the first bit of our $300,000,000 round, we had so little money in the bank left that it was less money than we saved doing Groq bonds. So had we not done that, we would have literally run out of money. So there were some really hard times. And I know every founder has these, and from the outside it's so hard to understand. It's like watching a TV show. You're not in it, but when you are there, everything is 10 to a 100 times more intense. Because people left their jobs. They left their careers. Their families are banking on this. You have to make decisions. What would have happened if we went out there and asked everyone to do Groq bonds and everyone quit? Then the shareholders would have been like, you have all of these people depending on you. But if you lean towards that vulnerability, people are often gonna go with you on it.

**Harry Stebbings** [17:06]:

So what is a world where inference is so crucial and 20 times more important than training? What does that world look

**Jonathan Ross** [17:13]:

like? I think the simplest way to understand it is equate an LPU or a GPU to an employee. If you have enough of them, the LPUs or GPUs, you can do work just like with an employee. It's a little different in the sense that they can't quit and take another job. You don't have to retrain. Once you get a model to a certain capability, it'll always be at least that capability. Right? So you get the consistency out of it. But now imagine that you're a startup, and rather than having to go out and hire a 100 people, you hire 10, and you buy the amount of compute equivalent to 90 employees' worth. That's a very different way of thinking about the world because now CapEx or in some cases different types of OpEx can be used instead of just employees. And in terms of inference, just to give you a sense of our scaling, we started 2024 with about six forty chips in production. We ended with over 40,000. This year we want to be at over 2,000,000. And next year the number is much, much, much larger.

**Harry Stebbings** [18:15]:

We seeing constraints on chip supply? Mean that is an unbelievable scaling story.

**Jonathan Ross** [18:19]:

Yeah, so for us to hit our numbers next year, which I'm not sharing publicly, we're gonna need almost all of the capacity of the fab that we're using. The the biggest issue so seven powers. We love seven powers, right? Hamilton Helmers? Yeah. Okay. You don't normally think of tech companies as having a cornered resource, but NVIDIA has a cornered resource. They're a monopsony, the opposite of a monopoly, a single buyer for HBM. And the interposer, the co OS. So what is HBM? So HBM is high bandwidth memory. And GPUs are And who produces HBM? I'm sorry for the dumb questions. There's three companies in the world that do this, SK Hynix, Samsung, and Micron. It's a specialty memory. It's only used in high end servers, so there's a limited quantity that's built. It's very expensive to ramp up. It's a very technically challenging type of memory to build more so than others. And so there's a very limited supply. And GPUs are so fast computationally that if you were using regular memory it'd be like drinking out of a martini straw. It would just take forever. This is why you see people preferring to do even inference, but especially training on GPUs rather than CPUs because the memory bandwidth is too limited. And CPUs rarely use HBM, they're mostly regular memory. The observation that we had when we started Groq, everyone knows Moore's Law. Every eighteen to twenty four months, like clockwork, double the transistors means double the compute. But we noticed that AI was getting better faster, and it it clearly wasn't the algorithms because algorithms have sort of discontinuous jump. It also, didn't seem to be the data because there wasn't that much more data. And the transistors were only doubling every eighteen to twenty four months. So where was all of this capability coming from? Turns out the number of chips was also doubling every eighteen to twenty four months. So rather than two x, it was four x. The question we asked was, if you're effectively gonna have an unlimited number of chips, do you do something architecturally different? And the answer is absolutely. So rather than using external memory, we just use a large number of chips and keep all of the parameters of the model in the chips live. And we just have this pipeline where the computation flows through it, sort of like an assembly line. So imagine if you were trying to build a factory and the factory was only one one hundredth of the size needed for the assembly line. So you'd run a bunch of cars through one one hundredth, tear it down, set up the next one one hundredth assembly line. You just do this over and over again. That's the way a GPU works. LPUs, very different. We actually just have the computation flow through a whole bunch of chips. So rather than using eight chips, we'll use 600 or 3,000 for a model.

**Harry Stebbings** [20:57]:

How does that change energy efficiency?

**Jonathan Ross** [20:59]:

It improves at about three x, and the reason is How does it improve it when you use more? Per token. So the footprint is higher. Think of it as the difference between a factory or a backyard sort of garage. The backyard garage is not gonna be as efficient. However, it has a lower energy footprint. Or another example would be if you're trying to transport a ton of coal from one side of the city to the other, and you did it on mopeds, or you did it with freight trains, which one would be more efficient? The moped would use less energy per trip, but it would need more trips and therefore would use more energy overall. In fact, this is one of the things most people misunderstand. They think that edge computing is lower energy. Actually, edge computing is less energy efficient than computing in the data center. When you're computing in the data center, it's a little bit like that freight train. You're actually getting to do a whole bunch of jobs simultaneously. So the fact that we don't have to read from that external memory means that we don't have to spend the energy doing that. Even with GPUs you get to batch. But going back to why it's so energy efficient, the amount of energy used in a chip, there's these physical wires. And the physical wires have a width, and when you look at the width and you look at the length, you charge that wire up to set it to a one, and then you discharge it to set it to a zero. It's sort of like charging a capacitor and discharging a capacitor using energy. The longer that wire, the more charge. When you have HBM here and another chip here, you're actually having to charge a wire between the chips and then discharge it every time you send a bit. That's a long distance to travel, but also the wires are wider than the wires that are inside the chips. So you just use a lot more energy. When we keep that memory in the chip, it's only traveling a little distance using much thinner wires, and therefore it uses a lot less energy. So

**Harry Stebbings** [22:44]:

do we see a world of LPU and GPU usage in Combinator? Like how does that distribution look between LPU usage and GPU usage?

**Jonathan Ross** [22:55]:

There's a couple of things. The first is training should be done on GPUs. NVIDIA will sell every single GPU they make for training. Right now, about 40% of their market is inference. If we were to deploy a lot of much lower cost inference chips, what you would see is that same number of GPUs would be sold, but the demand for training would increase because the more inference you have, the more training you need and vice versa. The other use case is we're actually so crazy fast compared to GPUs. We've actually experimented a little bit with taking some portions of the model and running it on our LPUs and letting the rest run on GPU. And it actually speeds up and makes the GPU more economical. So since people already have a bunch of GPUs they've deployed, one use case we've contemplated is selling some of our LPUs to sort of nitro boost those GPUs. Well, this

**Harry Stebbings** [23:45]:

is my question, is that, you know, people have bought GPUs so far ahead of time Yeah. That by the time you get them, they're deployed and installed, they're

**Jonathan Ross** [23:51]:

almost out of date. Actually, we've we've spoken to some customers that put orders in over a year in advance. They paid a year in advance and still haven't gotten them. The recent deployment we did in Saudi Arabia, fifty one days from contract to the first tokens being served in production in country. How are you

**Harry Stebbings** [24:10]:

able to do it so quickly? Fifty one days is astonishing.

**Jonathan Ross** [24:12]:

Yeah. Part of it is architecturally, things are much simpler for us. We don't have a bunch of other hardware components. We actually don't use switches to communicate between our chips. We just plug our chips into our chips. Our chips are the switch. We don't have all of this network

**Harry Stebbings** [24:27]:

tuning. Given the energy efficiency, given the predictability, why is NVIDIA not being more proactive on LPUs? What makes

**Jonathan Ross** [24:36]:

you think that they don't want to be more proactive on it? They don't talk about it. Well, why would they talk about it? That would be like talking about something you don't have when you're trying to project strength rather than vulnerability.

**Harry Stebbings** [24:47]:

Well, think if you wanted to protect shareholder value and you wanted to protect a Wall Street image of dominance and being ahead of the game, you'd at least say, are we of course working on LPUs as well?

**Jonathan Ross** [24:57]:

But then, until they had LPUs, they would effectively be exposing that there's something missing. Like, if you look at the last GTC, there was an announcement that the latest GPUs were 30 x faster than the previous generation. When you look at how it was done, there was this curve that looked kinda like this, and then it basically ended here. And then there was another curve that was kinda like this. Now that 30 x was from the end of this curve to this curve. If you moved it here, it would have been less than 30 x. If you moved it here, it would have been infinite. So their chip is infinitely faster than the previous one, but that wouldn't have sounded reasonable. Right? There's a history in this market of specsmanship because it's so hard to, like, get access to chip. And this is a lesson on enterprise sales. In enterprise sales, people rely on specsmanship because Specmanship is what? It's well, my specs are better than your specs. My chip is faster than your chip. I get more teraflops per second than you do. But who cares? Tell me what the the tokens per dollar is, tell and me what the tokens per watt is. Nothing else really matters. But people will find all of these other weird things to measure that they might be better on. Sort of like, I'll sell you a car with better RPMs. RPMs don't matter. What matters is miles per gallon and maybe the speed that you can drive at, although speed limits kinda render them, you know, moot. Right? In the case of enterprise sales, there there was a time when you would market soap. The billboards would say, our soap has more bubbles than this other brand's soap. Who cares? And what they figured out was, let's put really happy people up on a billboard after they use the soap, and then maybe people associate that happiness. Right? Lifestyle marketing. Sure. For some reason, enterprise still hasn't learned this lesson. It's still, we have more bubbles. We have more TeraOps. We have more whatever. Things that people just literally don't care about.

**Harry Stebbings** [26:45]:

So you think NVIDIA is, hey, we're 30 times faster is not good marketing.

**Jonathan Ross** [26:49]:

It worked because it's what people are used to, but our counter was, we we did a press release to that that said, Groq, still faster. That was it. And people went gaga over it. Right? Because it was just we we are. We're still faster. So who cares? Do you think Wall Street understands that way? Think they're starting to. Yeah. But again, I don't think there's real competition here. I think if you are competing, you have done something seriously wrong. If you're competing, it means that you haven't found an unsolved customer problem. Because if you're competing, someone else has already solved the problem. So why are you spending time on it? You don't view NVIDIA as a competitor. No. They don't offer fast tokens and they don't offer low cost tokens. It's a very different product. But what they do very, very well is training. Do They it better than anyone else and by such a wide degree, it's a solved problem. Why would we bother trying to solve a problem that's already been solved?

**Harry Stebbings** [27:43]:

So you're like, seed the training market to them, we'll own the inference market. Yeah. And they're saying, we also want the inference market. Of course. The way it always works. So what do we do now? So now we are competing in

**Jonathan Ross** [27:54]:

the inference market. But are we? Yeah. We don't really have people saying we're gonna buy GPUs instead of you. We do have people saying we're gonna buy both. That happens. But we don't care. I showed a demo to someone, and he's like, should we just not buy any more GPUs? I'm like, no. You should buy every single GPU you can get your hands on. And he's looking at me very perplexed, and I'm like, how are you gonna do training? We don't do training. Buy the GPUs. Get every single one you can because I want your models running on us to be really Totally, but for inference, they don't need to buy NVIDIA anymore. They don't need to buy GPUs for inference, but if you can get them, I mean, they're a little expensive, but if you're used to it, why not? Plenty people still sell mainframes. If you want lower cost and faster, then you want an LPU. How much lower cost is this? More than five x lower. More than

**Harry Stebbings** [28:42]:

five x lower?

**Jonathan Ross** [28:43]:

Just the memory alone in the latest GPUs costs more than our fully loaded CapEx per chip deployed. And on top of that, so we talked about the energy efficiency. So we use about a third of the energy per token. Over a three year period, one third of our cost is the OpEx, which is mostly energy and data center rent, and two thirds is the CapEx. Which means that since we're one third of the energy, the cost to run that GPU to produce the same number of tokens for inference is the same as our total cost. Just the OpEx for the GPU is the same as our CapEx plus our OpEx.

**Harry Stebbings** [29:14]:

I'm really sorry. I'm I'm asking the most stupid questions, but I'm just rolling with it. We're in Paris and it's nearly the end of the day. Why is 40% of their revenue inference then? Why have you not taken so much more of that?

**Jonathan Ross** [29:26]:

At the beginning of 2024, we only had 640 chips. At the end, had 40,000. We're not at that scale yet. You have to provide quality. You have to provide low cost. You have to provide speed, but you also have to provide capacity. And so this is where that that most important part of not using HBM came in. It means that we effectively have no scale limits. So the GPU itself is actually manufactured using the same process that you use for your mobile phone. So the same silicon that's in your mobile phone is the same silicon for the GPU. In fact, they build the mobile phones first, the mobile phone chips first, because they're smaller, so they yield better. NVIDIA actually gets it after Apple. The difference is that memory. That's the only difference, but that memory is the hard part to manufacture that's that's limited in scale. So by us avoiding that, we effectively have almost no limit on how much we can scale up. And that's important for inference. What is NVIDIA's margin?

**Harry Stebbings** [30:20]:

70 to 80%. 70 to 80%. So they can take 70 to 80% off and be radically more Yeah. Comparatively cost effective compared to you. You could destroy their margin.

**Jonathan Ross** [30:31]:

But in that same vein, you could almost say we're one of the best things that's ever happened to NVIDIA because they can make every single GPU that they were gonna make, and they can sell it for training, high margin, right, gets amortized across the deployment. You know, we'll take the low margin, high volume inference business off their hands and they won't have to sully their margin. What's low margin? Depending on the deal. We do get some on the the backside, but upfront it's about 20%. About 20%?

**Harry Stebbings** [30:56]:

Yeah. Okay. So theirs is 80, yours is 20, but then you're looking at a 20

**Jonathan Ross** [31:01]:

But then we get more later off of it, so we take some of the risk. What do you mean you get more later, sorry? So the deals that we do, the partner will off because we don't deploy our we don't spend money for our own capex. The partner will put up the money for us to deploy. We pay back with a decent IRR, but we split, and most of it goes to the partner. And then once we hit the IRR, it flips the other way. So others are putting the CapEx up for us. What

**Harry Stebbings** [31:23]:

does it look like at the end then?

**Jonathan Ross** [31:25]:

Well, at the end, it's not like other business models. So we we didn't just innovate on the chip, we also innovated on the business model. And we're limited in how much money we can make based on how much we can deploy, not how much money we have because the partners are putting that money up. So when I'm looking at what we can do, it's all about how much we can scale. What are the limits to your deployment? Is it purely chip constraints? Mostly. So you're asking about misconceptions in AI. I think one of them is about power. It is true that there is a mismatch in the market between people with chips and people with power, but that's partially because you need a data center in the middle and there aren't enough data centers. Those aren't the hardest thing in the world to build. They're not easy, but they're not the hardest thing. It's harder to build up the power. However, because of that mismatch, you have big hyperscalers going around and saying, I need a gigawatt of power. And they'll say this to to 60 different potential data center builders. Then all of sudden, hear this echo. Well, I heard that, you know, there's a there's a gigawatt here and a gigawatt here and a gigawatt here. And all of sudden, there's, like, 60 gigawatts of demand, and it's this echo from that first gigawatt. The thing is, I am aware of about 20 gigawatts of power that people want to make available for data centers now. Right now, there's about 15 gigawatts of data centers worldwide, so more than double the current capacity. Concern that I have is that people are now building up more power, and what's going to happen in the next three to four years is people are going to be like, I built up all this power and no one's using it. This was like a complete waste and we're never going do this again. Then what's going to happen, remember that doubling of chips every eighteen to twenty four months? Well, three to four years, you double that 15 gigawatts twice, and now you're talking about, what, 120 gigawatts? There isn't that much power available. And then another one after that, now you're at two forty. And so what's gonna happen is we're gonna overbuild slightly right now just because of that mismatch and the miscommunication that's going on right now. And then we're gonna dampen our building and we're gonna close down on that. And then we're gonna have the real need for the power. That's my big concern right now because that power will become a hard bottleneck in three to four years.

**Harry Stebbings** [33:27]:

Why will we have that data over or data center oversupply when we are moving into a world of inference which will be 20 x larger than training? So the problem

**Jonathan Ross** [33:36]:

with data centers is everyone thinks that data centers are real estate, and a lot of people do real estate. Data centers are not real estate. The common joke in the industry now is someone says, I'm going to have 100 megawatts of capacity for you, and I'm going to have it in three months. Are Are you willing to sign? And then you ask a question like, what's your uptime? And they're like, I don't know, whatever the the power grid is. And you're like, wait, what? Where are your generators? Oh, haven't ordered those. I'll order them now. You know that there's a ninety month lead time on generators right now? Oh, really? And then the next Ninety. Nine zero. And then the next question is, where are you getting the water from? Wait. Data centers need water? I thought it was a bunch of chips. So there's a bunch of people who have no idea what they're doing going into it because they think it's real estate. Those people are now building an oversupply of data centers, but they're not really building them. So they're they're fake data centers

**Harry Stebbings** [34:31]:

that people think are real. What happens to those data centers? Because they're not gonna be utilized, are they? Amazon is not gonna pay for a data center that doesn't Well, Amazon doesn't fall for this. Amazon has really good people. Whoever the the buyer is is not gonna pay for a data center that's got no water or got no power. Yeah. And so is it just wasted from your Most of these projects will never be developed. Will we build them fast enough? You said about the data, like the oversupply. It take it does take time to build a data center.

**Jonathan Ross** [34:56]:

Yeah. So it's almost okay. If you train a model, you really wanna amortize it for about six months. If you deploy chips, you really wanna amortize it for three to five years. One more on the three year side, others are more on the five year side. If you build a data center, you're probably talking ten to fifteen years. In a power plant, you're talking like fifteen, twenty years. The the problem we have in the industry is not on this and there's this mismatch between the sort of financing and the needs here. So you have someone who wants to train a model, and they're gonna be doing this for six months, they and don't understand why people want three to five year commitments on the chips. And then the people deploying the chips don't understand why someone wants a fifteen to twenty year commitment on the data center. Right? It's at seven years now on the data centers. And then the people building the data centers then need a long It's a seventy eighth commitment. Yeah. That's the kind of thing they're asking for. So you've got this complete mismatch throughout the ecosystem. But the funny part about it is while they all want to take zero risk and have a committed sovereign wealth level sort of credit rating on the other side of it with long commits, the longer the payoff time, the more generic the infrastructure is. A model has a pretty specific use, but accelerators like LPUs and GPUs can be used for other things besides generative AI or LLMs. The data center can be used for other things besides accelerators. The power can be used for anything. While they're looking for the least risk over here, it's the place where there is the least risk because if we don't use it for for AI, we'll use it to power all of the electric cars.

**Harry Stebbings** [36:32]:

Is this a case where incumbents win because they're one of the only ones who are able to match the durations required by data center providers?

**Jonathan Ross** [36:39]:

Well, and and this is why we partnered with Aramco and this new entity in Saudi Arabia because they have an enormous ability to fund this over the long term. They have a very long term perspective. They have an amazing credit rating.

**Harry Stebbings** [36:53]:

When you say they have an ability to fund it, and this is where the misconception was, people think it's a funding round of a billion and 0.5. It's not a funding round of 1,000,000,000.

**Jonathan Ross** [37:01]:

No, did not raise 1,500,000,000, that's revenue. That's actually about 30% of the revenue of OpenAI.

**Harry Stebbings** [37:07]:

Can you just walk me through how that deal is structured?

**Jonathan Ross** [37:09]:

Yeah. We started off last year, right, and we got to 19,000 of our chips deployed. We did that in about fifty one days. The question was, what can we do this year? So they've gone off, they've collected up a bunch of power in the country, and the deal is structured so that they will put up the capex for us to deploy our our chips in that data center or those data centers, and we pay back based on the money that we make. It's a little bit different than debt in that they participate in the upside. It's similar in nature. It is revenue because we actually make profit upfront. How does that change what you can do? Well, we are not limited by capital anymore. There is one misconception around Groq. There was a paper that was written that said that we couldn't be profitable while being the lowest price. We could charge more. But actually, we have a very positive contribution margin right now. As far as we know, we're the only ones that are actually making money running these open source models. Because with the open source models, everyone's sort of competing with VC dollars trying to take market share Uber style. Right? And meanwhile, we're sitting here going, we could do this all day long because we're making money. We're able to even pay off an IRR and and make our partners money. There's another part of the model. So we're also working with some proprietary model providers. So we actually showed off the first one at Leap on Sunday where we did a voice model with Play AI. That one is also a rev share. But the thing is, they get to make money off of that, whereas most others in the industry are losing money because of the commoditization of the models.

**Harry Stebbings** [38:48]:

So do you have cheaper pricing over time as you bluntly have less monopoly power, or do you have higher prices as your monopoly increases?

**Jonathan Ross** [38:58]:

We want the margin to stay about the same, but we want the prices to go down. Because then we get into Jevan's paradox, and and life gets great because we're gonna scale. Our focus is on getting to scale. To preserve human agents in the age of AI, we need to be one of the most important compute providers in the world. Our goal by the end of twenty twenty seven is to be providing at least half of the world's AI inference compute. We think we could be further than two x given that we don't have all the constraints. But in order to get there, do need to be very aggressively building out, and we need to give people no excuse for not running their models on us and and using the models that are on us by charging extra. And and what I keep telling the team over and over again, because you have to remind them sometimes, is we're growing faster than exponential. And when you are growing faster than exponential, there's no amount of profit that you can make that matters. What matters is getting a toehold in the market and becoming relevant.

**Harry Stebbings** [39:54]:

What could prevent that?

**Jonathan Ross** [39:56]:

We used to be worried that someone would try and price below us, and then we realized that wasn't a concern because there's so much money going into this that people are going to want to lose less money by running on us. That isn't a concern, but that was the big one early on until we realized that.

**Harry Stebbings** [40:13]:

When we see Zucker investing $65,000,000,000 in data centers, what does that actually mean? That means he's internalizing all of the margins that he would have had to spend on data centers with the providers that we mentioned earlier, and Facebook is doing full stack. Meta's doing 65,000,000,000

**Jonathan Ross** [40:28]:

a year. I think Google said 70 or 75, and Satya said Microsoft's doing 80, and then you've also got Stargate. Yeah. These are crazy sums of money. And this is all for data center build out? No. It also includes the stuff that goes in. Includes the chips as well, the systems, everything. We've never seen money like this, no? No. There's never been anything like this, but there's never been a case where it was so clear that there was gonna be value at the end. If you knew how successful search was gonna be. Right? Remember, Google stayed private as long as they did because they were afraid that Microsoft would figure out how much money search was making and then would try and replicate. And the moment that they went public, bing. They they called that perfectly. Everyone knows how much money there is in AI, so everyone's going after it.

**Harry Stebbings** [41:11]:

Do you think that value is distributed amongst many players or concentrated towards one or two? I completely agree with you in terms of the clear value when assigned, but is it distributed to some levels evenly or concentrated?

**Jonathan Ross** [41:24]:

It's a power law. The more value there is in the economy, the more risk there is of a single entity being so far on one end that they just dominate. And you see this with the mag seven. And it's predicted just the bigger the economy gets, the more you will have big swings in in the economic outcomes. Right now, the hyperscalers are all sort of even in their market caps. It's strange. You would expect one of them to just be killing it and taking it much further. And so I don't understand why they're so closely grouped. So when we

**Harry Stebbings** [42:01]:

think about that distribution, how do we think about changing that then? Like, obviously with the Groq, want to be one of the MAG seven, you want to be one of the most important companies in the world. How do you see

**Jonathan Ross** [42:10]:

that? So the way that you get there and the way that you stay there are two very different things. There's sort of a circle of life that happens in startups. The first circle the first stage is solve an unsolved problem. That's how you go viral. That's how you do well. The second stage is the marketing stage, which is now other people are trying to copy what you've done because they can't think of something themselves. And now you have to fight it out in in advertising and marketing and whatnot. You see CPG companies often get stuck there, It and becomes more about where on the shelf they are than anything else. Then the final stage is the seven powers. It's once you've found some of those and you've really started improving it and you have sort of systemic advantages. And then what happens is someone solves an unsolved customer problem and the whole cycle of life continues. Right? Now Google has to redo this because LLMs are better than search. So the way that you start off to become a mag seven is you solve that unsolved problem. The way that you stay there is first you find one of those seven powers or or multiple, but then you have to be ready for when you get disrupted to continue fighting back and and solving customer problems.

**Harry Stebbings** [43:16]:

We mentioned the different huge amounts of money that's being spent here. Is this a good bubble that lays the foundations for an incredible next ten to twenty years, where, finally, the capital actually turns out to be productive but not seemingly so on paper? Or is it where actually just a huge amount of money is incinerated on depreciating assets?

**Jonathan Ross** [43:36]:

I can guarantee you that a huge amount of money will be incinerated. I also bet that in total, more money will be made than will be put in. This is the problem. You you have to look at it either in aggregate or individual bets. When everyone is making investments in the market, some people are gonna lose money because not every company is gonna be successful. What you always see is when there is some real tech improvements or things coming. You've got the things that were early, that people are investing in heavily, that are super successful, and then everyone else wants to get in on it. And, you know, it goes from you have AI chips and AI models to now you've got AI t shirts, and next thing you know, you've got AI thermal grease. People just start applying AI to everything. Next thing you know, you'll have an AI condo. So the the trick is discerning what is real and what isn't. You're always gonna have all of these really obnoxious charlatans coming in whenever there's something real, and that's unfortunate. But eventually they get cleared away once people start to understand the technology and what's real and what isn't. The job is to start educating. And the more educated people are, the less they'll invest in AI thermal grease.

**Harry Stebbings** [44:47]:

What is the largest individual bet that will lead to the largest incineration of cash?

**Jonathan Ross** [44:53]:

I'm not gonna call anyone out in particular, but I actually think it will happen across every single discipline. The are you aware of the Keynesian beauty contest? No. John Maynard Keynes, The Economist? This will explain everything you need to know about VC. So I'm nervous, but keep going. So take a magazine full of models, human models, like, you know, good looking models. Have a whole bunch of VCs in the room. They're allowed to make bets on who the the most beautiful model is. In the end, whoever has the most money on them is the winner. And based on the proportion that you put on that particular model's face, you get the share of all of the money. If you put money on one that isn't the most beautiful by dollars, then you lose your money to the the people who bet on that one. That was sort of the the bet that SoftBank was making, which was they could win the Keynesian beauty contest. I'm just gonna put more money in and I'm gonna win. That is problematic when you have true technological advantages as opposed to marketing. When you're solving customer problems, it's a weighing machine. Once the customer problem has been solved, you then get into this sort of popularity contest of marketing. Now something unusual has happened this time around, which I don't think has ever happened in in VC before, which is you see people raising billions of dollars who have competitors who've raised billions of dollars. It usually, there is a clear winner in the Keynesian beauty contest. You don't have this, like, fight where, you know, it's it's sort of like, well, I gotta put a little more money in. I gotta put a little more I gotta, you know, put 10,000,000,000 in. No. I'm gonna put 20,000,000,000. I'm gonna put 500,000,000,000 in because the Keynesian beauty contest has gone completely amok. This has never happened before, and so now people don't even understand how to react because it used to be if someone had raised a billion dollars, you're like, oh, they're the winner. Now it's like there's three or four competitors who have a billion dollars. So who wins and who loses? Like, is massacre gonna incinerate the largest amount of cash ever? I think the Keynesian beauty contest no longer applies here because there's so much money available being spread out, and I think you're gonna see that the people who have the best products are actually gonna be the winners because everyone can be capitalized. But there will be problems for the winners because of this. The problems are gonna be of the sort you had this employee that you were gonna hire, and someone offered them a ridiculous amount of money. Yeah. You see this all the time now. And they could have gone and contributed to the winner, but now they're contributing to a competitor that shouldn't even exist or is equally likely to win, now you're splitting the talent.

**Harry Stebbings** [47:27]:

What do you also do when you have such high salaries? We've seen 1,000,000, 2,000,000 for kind of junior to mid level in some of these companies, And they are living an amazing life actually in great places. You think they're living that amazing life in Guangdong when they're working for DeepSeek or any other Chinese alternative? I think they're actually getting paid much less working their fucking ass off twenty hours a day and not getting kombucha and being paid 2,000,000 a year. Fair? Not only fair,

**Jonathan Ross** [47:53]:

we have a policy that we never offer the highest because we want people to choose us, not choose the salary. If we win in a bidding war, then that means the next time someone comes along with a higher salary, that's it. They're just gonna go take that other job. There's no loyalty. They don't believe in the mission. Instead, we focus on, look, we're gonna build this. This is your opportunity. You're gonna get to work with amazing people. Spend some time with the team. Are these the people you wanna be working with? Because frankly, you're gonna make so much cash. It doesn't matter. But bet on the equity, the outcome. Help us make this thing valuable. And people who buy into that, they're so much easier to manage because they're mission oriented. They all wanna do the same thing. They're not there because they want the kombucha and they're not going to complain because the cappuccino machine is broken. They'll just go and buy their coffee next door.

**Harry Stebbings** [48:40]:

Will you and NVIDIA move into the model air? Everyone talks about model providers becoming application providers, or infrastructure providers become model providers.

**Jonathan Ross** [48:48]:

We have decided that we're not going to train our own models. We'll do a little fine tuning for specific cases or whatnot, but we don't want to compete. That's really important because people are putting their models with their weights on us. And they don't want us to learn from and take that stuff for our own benefit. This is the problem you have when you work with hyperscaler because you know they're also doing everything that you are doing. So we've decided model providers, you make the model, we don't do that. There's also the data side of the users and the queries. So the other thing that we could do that we do not do is log the queries and then we've got data if we wanted to train. We don't train. We have no reason to hold the data. We only temporarily store things in the DRAM. There's no persistent storage. If the power went out, everything's gone. And DRAM is limited, so we can't hold things for a long time. So you know that we don't have your data. Now, people who are building businesses on top of us, you can obviously keep the data from your customers if you want. We have no control over that. That's fine. But we don't take any data. Do you think NVIDIA move into the model providing? It's possible, but if I was them, would avoid it because I wouldn't want to give the customers of mine. I mean, NVIDIA's great at training. Right? It's crazy. It would be like being a automotive a car company and then creating your own taxi service. You're now competing directly with your customer. Right? And I think tech companies love to do this. We have a management philosophy, and it's based on Big O complexity. We only do things that require a sublinear number of employees. So what I mean by that is if someone comes to me and says, I need 10 people to go do this thing. A lot of people would say, Well, why can't you do it with five? I would say, okay, you're supporting customers. If we double the number of customers, do you need 20 or do you need 11? Because I want to know what's that growth rate. Are they automating everything? We completely automated our compiler. We completely automated everything that we our large portions of our cloud. That means that we can scale with a small team. We have 300 people. We have 300 people. We built our own chip. We built our own networking hardware and software. We built our own runtime. We built our own orchestration layer. We built our own compiler. We built our own cloud. All this with 300 people. Now, we would only be able to do this with a small number of people because you don't have the communication overhead. You have to decide what your constants and variables are. What are the things that you want to preserve? And one of the constants is talent density. We want to stay small. We want to stay nimble. And the other side of this is growth is a problem. So we measure our growth in what I call problem units. So a problem unit, every time you triple something, you have about the same number of problems as the last time you tripled. Going from a 100 employees to 300, 300 to a thousand, a thousand to 3,000. Each one of those has the same number of problems. We scaled from 620, or 640 LPUs last year at the beginning to 40,000. That's four problem units. That's four triplings of the number of chips. If we were also tripling the number of employees, that would be another problem unit. Management bandwidth is limited. You can only solve so many problems. So you have to decide where you're gonna allocate them. If you build things really well from the beginning and you can scale up with the number of employees you have, then you can scale over here. You wanna triple the number of customers, there's another problem unit that you have to solve. What's

**Harry Stebbings** [52:08]:

the biggest challenge when you are scaling at that rate, but then the team is not scaling in conjunction with it?

**Jonathan Ross** [52:13]:

There's this common belief that the people that you have early on are right for the job, and the people that you get later, they're better in a sort of more corporate environment. I don't think that's the case. I think you should always try and get generalists. Otherwise, you get stuck and ossified in a particular way of doing things because that's what that one person knew how to do. But there are people who burn out. Being in a startup is hard. There are people who just literally burn out. There's also people who were the best that you could get at the time. And then there are people who are just unmanageable wild children, and they should go off and start another startup, and they shouldn't be, you know, scaling with you. That's that happens, but it's the rarer of them. Saying that you're you're gonna hire b players because you've gotten large enough is is laziness and an excuse. It's a lack of creativity in your business model and how you're going the algorithm of how you're gonna scale. Think of it this way, Walmart versus Amazon. Walmart wants to double the number of customers. They have to double the number of stores and employees. Amazon does not need to double the number of websites. That's a fundamental advantage. But Amazon still has to double and improve the logistics. They don't have as many problems where they have to scale linearly, but they have some. You wanted to disrupt Amazon, what you would do is you'd build a completely robotic logistics system and bring the the overhead and complexity of that down, and then you could outmaneuver them. That's how you need to improve. Don't just say I need more people. Focus on the algorithm of your business.

**Harry Stebbings** [53:39]:

The last time we spoke, we discussed DeepSeek. I think more has come out over the last few weeks about, probably, their innovations, some of the distillation that they used. Where is China better than us today?

**Jonathan Ross** [53:50]:

Well,

**Harry Stebbings** [53:51]:

they're more

**Jonathan Ross** [53:51]:

willing to use things that maybe they shouldn't be using. They distilled the OpenAI model. A lot of people have the opinion, well, OpenAI was scraping the Internet, so good one for deep seek. But most of the model providers had considered that a red line they didn't wanna cross. I don't know if that's gonna change, but but it might.

**Harry Stebbings** [54:11]:

But the other thing open source nature of DeepSeek, OpenAI now benefit from the innovations that they did also have. Well, and they also probably have all the

**Jonathan Ross** [54:20]:

data that DeepSeek paid them to generate. They also were clever. They they innovated. I think the biggest thing is this is a shot in the arm for morale in China, and it gives them a sense. But, you know, as I said, Sputnik two point o, it's also woken up The US. How do you compare Stargate to the $128,000,000,000 that China have now committed? China has a more complicated situation and a simpler one at the same time. The problem is they don't have the technology that we have in terms of the chip efficiency. On the other hand, they have scale. If they wanted to deploy a 150 nuclear reactors, as I think the plan is, no big deal. They just do it. So if the chips aren't as efficient, they can just deploy more of them. On the other hand, if they want to go out into the world and deploy chips like they did with Huawei and networking gear, that's going to be complicated because people aren't going to have the power around the world to run more expensive accelerators. At home, I don't think anything's a problem. I only think as they're trying to expand, it's gonna be an issue. What do we not know about China that we would like to know? I think the most important thing to understand is where they're gonna end up on the censorship and privacy of these models. We come from democratic countries. We have an expectation that companies can build something that says anything. Are they going to be permissive and allow models to make mistakes and hallucinate, or they're going to shut it down. If you know that, you know whether or not China has a shot. One of the biggest nightmares that they have is free speech. It's the exact opposite of that vulnerability we talked about earlier. Can you imagine Xi Jinping going out and saying, country, we've lost our advantage in AI. I need your help. Never. Ever. It's always gonna be, we're the greatest, we're the best, everyone's gonna know differently, but they're all gonna have to toe the party line. Right now, because of that, I think it's really hard for them to just allow these models to to say anything. Say, you know, The US is great and better at AI. That that's a bad thing for them. And so that's gonna really tell you a lot about the AI story in China.

**Harry Stebbings** [56:26]:

And so if they aren't permissive of more open, truthful models, then they're inherently disadvantaged, you're saying? I I forget

**Jonathan Ross** [56:34]:

if it was was it Jack Ma who got in trouble with the CCP? Yeah. If they aren't more permissive than if you're running a Chinese tech company, your fear is that you become Jack Ma. That's really gonna stifle innovation. If I was in China right now, I'd be looking for the exit. If your craft is AI, I would wanna do that someplace that's supportive. Do you really buy that they don't have access to Blackwell?

**Harry Stebbings** [56:57]:

This is China. I think Xi Jinping's like, no, sorry. No Blackwell.

**Jonathan Ross** [57:02]:

I don't think it matters on whether or not they physically have it. Because right now most of the cloud providers are happy if you swipe a credit card to rent it to you. But there is limits to renting, though. One of the concerns right now is about Malaysia or Singapore, that region over there, being a place where people are deploying GPUs with the wink wink, like, we're not gonna rent it to China. That's a belief that a lot of people are doing that. Otherwise, that's a lot of GPUs for that region. But that feels like it's even more of a safety net just in case the tap ever gets turned off at the hyperscalers. Because right now, you could just write a check to any of the hyperscalers and say, need these chips. They'll deploy them and you can run on them. It doesn't really matter where you're coming from. I mean, if you're a sanctioned country, no. China's not sanctioned.

**Harry Stebbings** [57:46]:

So we have China, where China is obviously in terms of innovation and actually proving that they they are in the race. We have The US, and then we have Europe, which feels like it's languishing. Is this the ultimate now than Europe's coffin?

**Jonathan Ross** [57:58]:

We talked about how Groq almost died, but we had the right technology all along. We're just waiting for the thing, for for the LLM's to arrive. And I think Europe's very similar. I think Europe has amazing talent, amazing talent, But that talent leaves and goes to The US or other places. The question is, how do you have Europe's LLM moment? How do you position yourselves? And it's not that complicated. The problem is when you surround yourself, you become the average of your five closest friends. Right? If your five closest friends are like, that'll never succeed. Ah, you should just keep your job. Ah, startups, they're terrible. Then you're gonna be risk averse. But if your five closest friends say, you should do it. That's great. I support you. Then you're gonna be more likely to do a startup. Even in Silicon Valley, people make that transition from the big tech company to the startup, and it's hard. They're comfortable. The big companies take care of them. They have a fiduciary obligation to their family. How do they make that leap? And it's because you've got tons of entrepreneurs trying to hire them, and they hear the pitch all the time and they get used to it. They also see the success around them. Then VCs come in and try and close some of the candidates in the early stage too. Right? Europe needs the same thing. You need a place where people are surrounded, just surrounded by entrepreneurial people who are risk on and who aren't gonna try and talk people out of joining a startup. From a regulation

**Harry Stebbings** [59:27]:

perspective, Europe is unbelievably efficient and the masters of regulation. You know, I was speaking to someone the other day in the EU who supposedly hired 1,500 people for AI safety and policing. What would you do if I put you in charge of European AI regulation?

**Jonathan Ross** [59:42]:

Well, I wouldn't waste my time regulating something that doesn't exist. Instead of regulating, what are you gonna promote? You wanna promote risk taking. You wanna promote that enclave of people who are risk on. I was just visiting Station F yesterday. Amazing. Macron was there. It was, like, full of people. Right? Vibrant. You feel it. And I was talking to the the person who runs station f, Roxanne and Xavier O'Neill. And with Roxanne, we were talking about, what about a city f? What about you start off with like 10,000 people in the center, right, a little radius, and then once that's full, you expand it. Once it's full, you expand it, and so on. You get to like a million people in Europe who are all risk on. The little Silicon Valley here. I would give it special economic dispensations. I would allow everything that employers need. I would make it simple. And I would say, you know what? If you don't want to buy into that, that's fine. Go to other regions in France. Go to other regions in Europe. But if you want to participate in what is going to be the biggest technological revolution in human history,

**Harry Stebbings** [60:41]:

this is the city for you. You're not inherently punishing incumbents then. What I mean by that is if we are talking about, I'm just using this as an example, AI insurance underwriters, startups. Yeah? There's many companies that are going after insurance underwriting in AI, and you are giving them benefits like that. You're inherently punishing some of the biggest providers of insurance in your region. You're inherently punishing people who hire 200,000 people. That feels unfair.

**Jonathan Ross** [61:05]:

There is no right to be an incumbent, especially a slothful incumbent that is not reacting to disruption. And you want to encourage disruption. This is one of the things in Silicon Valley, you can move from one place to another. There are no We had non solicits when I started, but even that's gone. That free movement of people is very important. Are you allowed to start work straight away? Straight away, but not before. If you start before We have months months, so six months. There's no such thing. And so in that region, I would say you can immediately start, like, literally the next day. That is so good. We have to wait six months. It's not good. If you are a company right now, it it feels like, okay. Well, it's harder to poach. What does that do? It suppresses wages. It's harder to hire someone. They're less likely to move. There's less competition. It suppresses wages. And by the way, the company

**Harry Stebbings** [61:53]:

has to pay for the six months anyway. It makes no sense at all. I I so I I totally get you and understand that. Can I ask you, you know, you mentioned that what would you promote? A lot of people would promote. I love the way you said risk on. Being a European, I actually thought first safety and regulation, specifically safety. So, sticking with that, all that Dario will talk about these days is safety. Is he losing a step by being so focused on safety when bluntly his competitors are talking about product?

**Jonathan Ross** [62:22]:

So safety matters in AI. It's a little bit like nuclear power. Lots of pros, lots of cons. I'm worried about different things than Dario is worried about. I'm more worried about people voluntarily giving up their decision making authority because it's so easy. And this is what I mean by preserving human agency in the age of AI. A good analogy is you probably know plenty of wealthy people and the struggles they have bringing up children with wealth. I refer to it as financial diabetes. You have children who aren't incented. They're not going to strive to succeed. I was very fortunate when I was growing up. I actually just told this story for the first time today, so no one's heard it. But I was fortunate because my father lost all of his money multiple times. And he would sell a billion dollar life insurance policy. He would get all the commissions from that, and he would have tons of money, and then he would spend it all. There was one time we were living in a $20,000,000 mansion. There was a couple of times where we ordered food, he would talk to the delivery guy, and would convince him to give us the food. And he would pay him back later because we'd get money later. But this time he was so despondent, he sort of locked himself in his office and wouldn't come out. My little brother came to me and said I had to go and talk to the Chinese food delivery guy and convince him to give us the food. I was mentally preparing how to convince him to do it. And I walk out and I walk up to him and I'm like, getting ready to do my whole spiel, and he hands it to me and I'm like, I don't have the money right now. He's like, oh yeah, pay me later. I fortunately didn't have to do anything. He just trusted because we're living in a $20,000,000 mansion. But that happened multiple times. And when that happens multiple times like, have a friend who was homeless once for a couple of weeks, and he'd almost been homeless a couple of times. And he said the best thing that ever happened to him was that he was homeless for a couple of weeks because he survived it. And it's like, I've been through it. I I always viewed this as the worst thing that could ever happen in the world. But now that I've been through it, I can survive it. I'm not worried anymore. I think we live incredibly comfortable lives, way too comfortable. Most people don't have to go through that. We have this sort of financial diabetes as a society, and I think it's going to get worse with AI. We're really going into an age of abundance. Very few people have to worry about food security now. But what happens if you don't need to worry about home security or anything? What happens if you can just live a life without working? And what is that going to do to your psychology? And so as we enter an age of abundance, how do we get people to still be making their own decisions and have a filled life?

**Harry Stebbings** [64:51]:

Do we get better or do we get accepting of good enough? And what I mean by that is, you know, now bluntly with with the majority of schedules, we will start with OpenAI and we will do deep research and then we will use different prompts depending on different guests. And then we supplement it with a huge amount of research from speaking to Chamath and speaking to Scooter and speaking to everyone in between. We care about it being good enough first and then great later with all the references. Most people will actually just be happy good enough and get away with it. Do we as a human society get happy with good enough?

**Jonathan Ross** [65:25]:

When we hire, we hire for something that we call booking the win early. One of the most important driving forces for people is loss bias. When you have something, you don't want to lose it. When we have an engineer that we're hiring and there's a roomful of people who are saying, if we do this thing, we could be twice as fast, I want that engineer to hear, wait, if we don't do that, we're going to be half the speed we could have been. The loss bias. Right? Book the win early. Because it's possible, it must be done. That's a smaller segment of the population. Those are the people who deliver amazing things that no one else is gonna do because everyone else is like, that's good enough. However, I think with AI, it's so easy to create a prototype to stand out. You're gonna need to do that. And one of the things that's happened with the ability to communicate more freely and and see what other people are doing, Think back to what the restaurant experience was twenty years ago versus what it is now. The average restaurant is better than what high end restaurants were twenty years ago because people see all of the the stuff that others are doing the best, and they start to expect that. And you have less localization, more globalization. You have to compete at the highest ends. AI is no exception. There's gonna be 40 people creating that app. You know that you have to polish

**Harry Stebbings** [66:41]:

it in order to stand out. Listen, dude, I could talk to you all day. I do wanna do a quick five. What do you believe that most around you disbelieve?

**Jonathan Ross** [66:49]:

I'm gonna go with anti founder mode here. I'm anti founder mode. I believe in delegation. When you are telling people how to do their job, that is an indication that it's not necessarily a problem with you. It could just be that that person is not right for that job. And it's much easier to just direct them than to go find someone else competent. But it also means you probably haven't aligned them. We align people through this challenge coin. Everyone at Groq carries this 25,000,000 token per second challenge coin. What this is is it tells everyone what we're doing. It's an alignment. I can't tell you how many people I've showed that to who are like, that's awesome. And yet no one else is making them. You'd like to say, you know, the greatest things in life Yeah. I know gold, but I made decisions. Exactly. Well, this was a very made decision because I had to consolidate everything we were doing into one very simple message of we're gonna get 25,000,000 tokens per second, and then engraved it on a coin on this tiny amount of space right here and gave it to everyone at Groq. And now whenever we're in a meeting and something doesn't help with this, we can just tap their coin on the table and be like, no, no, that's not the way this is gonna go. Is everyone wrong on founder mode then? I think that's what you do when you don't have the quality of people working for you. You need the right gearing ratio between you and your direct reports.

**Harry Stebbings** [68:06]:

It's a really unfair question, but I have to ask it. How do you analyze Elon's attempt to buy OpenAI? So I

**Jonathan Ross** [68:13]:

was sitting at the Elise Palace, or however I pronounce it, at the dinner with Macron, and Sam Altman so it was it was Macron, it was JD Vance, it was Sam Altman. Frankly, I think Elon was a little jealous that Sam Altman was sitting next to JD Vance, and it wasn't him. And because it was right around the time that Sam Altman was speaking that he announced it. And frankly, I thought Sam's tweet response part of it was pretty good. I would have probably said instead of whatever he said about 9,000,000,000, I would have said, yeah, I'm gonna take Twitter public at $420 a share. It was attention grabbing. Right? Some people can't stand to not be getting attention, and so my revenge on this is to give as little attention as possible, so let's move on. What would you do if you knew you couldn't fail? I would put in 100% of the orders for every single chip we could possibly manufacture. Because right now, the demand is unlimited, but every time you triple, you find the same number of problems, and so you gotta keep you gotta do it a little judiciously. But if I knew that no matter what problem is gonna come up, that we didn't need to be safe at all, I would just go, great. We're gonna go build 20,000,000 chips. Done.

**Harry Stebbings** [69:28]:

In ten years' time, is NVIDIA three x bigger, 10 x bigger, or 50 x bigger?

**Jonathan Ross** [69:34]:

I I think they will be bigger. Training will become more important. I wouldn't be surprised if they were three x bigger. I also wouldn't be surprised if they stayed around the same. It's so hard to to tell where things are going because remember, a lot of assumptions in the investment in NVIDIA were they were going to run away with the entire market. Including the inference market. Including the inference market. And they just haven't built the right thing for inference. I do think as a weighing machine, they should increase in value. But there was so much popularity contest applied to it that I don't know if, you know, they might need to grow their revenue to get to where they are, but it's a pretty fair multiple given everything going on. The popularity contest skews everything.

**Harry Stebbings** [70:13]:

What's a crazy AI prediction you have that everyone else think is science fiction?

**Jonathan Ross** [70:18]:

I would assume that in the next ten years, and I know this is gonna be crazy, but you saw that picture of me and my weight loss. Right? Unbelievable, dude. Seventy pounds. Seventy pounds. Yeah. I was on Mounjaro. So if you know anyone who's overweight and it's hurting their health, get them on Mounjaro as soon as you can. It works. So what

**Harry Stebbings** [70:35]:

is Mounjaro?

**Jonathan Ross** [70:36]:

It's one of those GLP inhibitors, one of the the weight loss drugs that have become popular recently. It works. But my crazy AI belief is that if it is possible, if it is possible to significantly slow or stop aging, I think that you will have a Mounjaro moment in maybe the next ten years. Because that came out of nowhere. All of a sudden you could just lose weight. Something finally worked. And I don't know if it is possible to slow or stop aging. Some wear and tear is a real thing and it might just be impossible. But if it is impossible to slow or stop aging, if, then I think in the next ten years we will do it and it will be sudden. It'll be like the Mounjaro, it'll be like that moment.

**Harry Stebbings** [71:18]:

I don't see how it is not possible. Like, when you look at the advances that will come in medical research, I don't see how it's not possible that we will at least extend longevity by sixty years. I mean, diaries have all lived to a 150. I don't see why that's impossible.

**Jonathan Ross** [71:33]:

I don't either, but I also don't know that it is possible. And until I know that, I'm gonna I'm gonna stick that conditional in there and say, if possible. What have you changed your mind on in the last twelve months? And and this is less of a a mental one and more of an emotional one. We didn't have product market fit for seven years. That is terrible. Like, the morale, like yeah. When you find product market fit, the world is brighter, the birds sing, I feel like hugging people. You sleep. I sleep. Life is better. I forget if it was you or someone else. Someone was talking about type one and type two happiness. Yeah. I think there's a third. As a founder, the only type of happiness you get is this third type, which is future happiness. The other two, the the common ones are the present is happy. Right? The other one is you went through some real crappy stuff, but the memories are make you happy. Right? So there's past, there's present, and there's future. As a founder, you're living a 100% in future happiness. When you get product market fit, you start to get that past happiness. And when you start to get that revenue and everything, then you end up getting the present happiness. And it it changes everything.

**Harry Stebbings** [72:44]:

If you had to bet on one company other than Groq to define the AI era, who would it be?

**Jonathan Ross** [72:49]:

Probably focus more on the companies that that you haven't heard about. I And don't know what the companies are, but I can tell you what they're gonna do, and I can tell you what each one of them will be. Go on. The first one will be the one that solves the hallucination problem. The second one will be the one who is best able to break down sub goals for agentic. I think agentic comes after you solve the hallucination problem because otherwise you get these long chains where you can introduce hallucinations. It'll kind of work, but it'll work much better after. I think the next one, what I call the invent stage So right now, the way the LLM's work, they make the most probable prediction. It's actually kind of amazing. It's like, I'm gonna take an entire novel. You've got a detective, you know, murder mystery. You get to the point where the detective says, and the murderer is, and it can actually predict it. It had to understand everything, right? But it's going give you the most probable answer. And that's not good for invention. It's not good for art, writing. The reason that the writing from LLM is terrible is because it's predictable. So how do you actually say something that's non obvious, but is obvious when you see it? And we don't even have the right word for it. Right? Non obvious, but obvious. That is going to unlock invention. And then the final one is what I call the proxy stage, when someone makes it so that models can just make decisions for you. You can proxy your decisions, like the decision to do this interview, right? Other things had to be canceled. The flight had to be booked. We had to get a ride over here, right? You would trust an EA or a chief of staff to make that decision. You wouldn't trust an LLM. And that's the final stage, I think, before you get to generative AI. But each company that does that is gonna be a defining company.

**Harry Stebbings** [74:30]:

You said that we're gonna have to fix hallucination before we get, like, efficient agentic. Yeah. Does that mean that money going to agentic today will be burned?

**Jonathan Ross** [74:37]:

No. Let's take an example on hallucinations. So the examples I gave you were medical diagnosis and law were two areas that will be unlocked once we get rid of the hallucinations. But there are startups like Perplexity that are doing just fine right now even though there's hallucinations because it's not high risk. It's for entertainment only. But but if you click those links, you can check them and it works kind of okay. It depends on how risky the industry is that you're in, on whether or not you can get started trying to position for the wave early and generating. But if you're in the right position, we were in the right position for seven years and the wave came. So that money isn't incinerated. In fact, that recent deal we just announced is more revenue than the money we've raised.

**Harry Stebbings** [75:21]:

How does that

**Jonathan Ross** [75:22]:

cash hit? I know it's in the Throughout the year. Throughout the year. Yeah. But it's this year, potentially more next year, a lot more. What is that contract in three years? If we sell everything that we possibly can this year, it is many billions. But just from the capacity alone, there's tens of billions of capacity of of hardware that we could build next year in these sort of deals. But we're also doing it at high volume, low margin. So if we were talking about GPU sales and GPU prices, I mean we'd be talking about hundreds of billions. We're just not charging that much.

**Harry Stebbings** [75:52]:

Final one for you. The thing that I'm singly most excited for is actually disease discovery in terms of drugs. You know, obviously my mother has MS and that's it was always taught to me that it was incurable. And actually now it's like, actually, maybe not. What are you singly most excited for?

**Jonathan Ross** [76:09]:

We went from a phase where people were hardware engineers to they were software engineers. To be a hardware engineer is ridiculously difficult. The the training, you have to get things right. There's a real expense if you get it wrong. Becoming a software engineer is so much easier. All you have do is get a little bit of time on a machine and you can teach yourself. Nowadays, you can just download manuals from the Internet or tutorials from the Internet. I think prompt engineering is going to unlock a huge swath of human society. There's 1.3, 1,400,000,000 people in Africa who know how to speak. If you were to give them access to a tool that they could create applications live just by speaking to it, That would be another 1.3, 1,400,000,000 potential entrepreneurs. There's 8,000,000,000 people on the planet. And the difference is hardware was just ridiculously difficult to it was arcane knowledge that's hard to get. Software was plentiful. Language, you already know it. You don't have to learn a thing. What's that gonna do for venture? What's that gonna do for entrepreneurialism?

**Harry Stebbings** [77:11]:

Jonathan, I love talking to you. It's always such a broad and wide ranging discussion. Thank you so much for putting up with me in person, and I've loved it. Awesome. So glad to be here. I have to say that show was so much fun to do in person. I wanna say a huge thank you to Jonathan for giving up his time on the Paris trip after the AI Summit. If you wanna watch the episode in full, you can find it on YouTube by searching for 20 VC. That's two zero VC. But before we leave you today,

## Sponsor read

**Harry Stebbings** [77:37]:

turning your back of a napkin idea into a billion dollar startup requires countless hours of collaboration and teamwork. It can be really difficult to build a team that's aligned on everything from values to workflow, but that's exactly what Coda was made to do. Coda is an all in one collaborative workspace that started as a napkin sketch. Now, just five years since launching in beta, Coda has helped fifty zero teams all over the world get on the same page. Now at twenty v c, we've used Coda to bring structure to our content planning and episode prep, and it's made a huge difference. Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time. With Kodi, you get the flexibility of docs, the structure of spreadsheets, the power of applications all built for enterprise, and it's got the intelligence of AI, which makes it even more awesome. If you're a startup team looking to increase alignment and agility, Coder can help you move from planning to execution in record time. To try it for yourself, go to coder.io/20vc today and get six free months of the team plan for startups. That's coda.io/20vc to get started for free and get six free months of the team plan. Now your team is aligned and collaborating, let's tackle those messy expense reports. You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on the month end panic when you realize you have to reconcile it all? Well, Pleo offers smart company cards, physical, virtual, and vendor specific. So teams can buy what they need while finance stays in control, automate your expense reports, process invoices seamlessly, and manage reimbursements effortlessly all in one platform. With integrations to tools like Xero, QuickBooks, and NetSuite, Pleo fits right into your workflow, saving time and giving you full visibility over every entity, payment, subscription. Join over 37,000 companies already using Pleo to streamline their finances. Try Pleo today. It's like magic, but with fewer rabbits. Find out more at pleo.io/20vc. Don't forget to secure trust with your customers. Trust isn't just earned though, it's demanded. That's why over 9,000 companies, including Atlassian, Core, and Factory rely on Vanta to automate their security compliance. So Vanta helps businesses achieve certifications like SOC two and ISO 27,001, turning months of tedious work into this beautifully fast and straightforward process. Their platform automates compliance across over 35 frameworks. It centralizes workflows, and it proactively manages risk, all while saving you time with automation and AI. So whether you're just starting or scaling your security program, Vanta connects you with auditors and experts to get audit ready quickly and build trust with your customers. Get $1,000 off your first year by visiting vanta dot com slash two zero v c. That's vanta.com/20vc. As always, we so appreciate all your support, and stay tuned for an incredible episode coming on Wednesday with a dash at Merkur.
