Skip to content
20VCSep 29, 2025

OpenAI and Anthropic Will Build Their Own Chips

NVIDIA Will Be Worth $10TRN · How to Solve the Energy Required for AI... Nuclear · Why China is Behind the US in the Race for AGI with Jonathan Ross, Groq Founder

With Jonathan Ross · Harry Stebbings

Full transcript · 82 min · 16,035 words · 2 speakers

Cold open

The countries that control compute will control AI and you cannot have compute without energy. And now we’re going be able to add more labor to the economy by producing more compute and better AI. That has never happened in the history of the economy before. What is that gonna do? I personally would be surprised if in five years NVIDIA wasn’t worth 10,000,000,000,000. The demand for compute is insatiable. If OpenAI were given twice the inference compute that they have today, if Anthropic was given twice the inference compute that they have today, within one month from now, their revenue would almost double.

Jonathan Ross0:00

This is 20 VC

Harry Stebbings0:32

Intro

Harry Stebbings

with me, Harry Stebbings, and this guest holds the record for the most downloads in last year’s catalog of episodes. So I’m thrilled to welcome Jonathan Ross, founder and CEO of Groq, back to the hot seat. Now Groq is the AI chip company redefining inference at scale. Under his leadership, Groq has raised over $3,000,000,000 with the latest pricing the company at close to $7,000,000,000. And before Groq, Jonathan led the team that built out the TPU at Google, making him one of the leading architects of modern AI hardware.

Now this conversation has it all, with everything from OpenAI, Anthropic, Oracle, to what happens to NVIDIA in the next ten years, to how should we think about China. This was incredible and very wide ranging discussion. But before we dive into the show’s

· Sponsor read0 min · 581 words
Harry Stebbings1:16

stay, I love seeing the team come together to make this show happen. What I don’t love is trying to keep track of all the information, the data, and the projects that we’re working on across dozens of platforms, products, and tools. That why we use Coda, the all in one collaborative workspace that’s helped 50,000 teams all over the world get on the same page. Offering the flexibility of docs with the structure of spreadsheets, Coda facilitates deeper teamwork and quicker creativity. And their turnkey AI solution, the intelligence of Coda Brain, is a game changer.

Powered by Grammarly, Coda is entering a new phase of innovation and expansion aiming to redefine productivity for the AI era. Whether you’re a startup looking to organize the chaos while staying nimble or an enterprise organization looking for better alignment, Coda matches your working style. Its seamless workspace connects to hundreds of your favorite tools, including Salesforce, Jira, Asana, and Figma, helping your teams transform their rituals and do more faster. Head over to coda.io/20vc right now and get six months off the team plan for startups for free.

That’s coda, coda.io/20vc, and get six months off the team plan for free. Coda.io/20vc. And talking about precision, that’s exactly what Brex brings to your finances. So when Brex was founded, it wasn’t just about creating another financial product. It was about solving the really gritty challenges that founders face daily. Let’s be honest, building something from the ground up is hard enough without dealing with clunky outdated banks that pile on fees and leave your cash idle. Brex is different. It’s the financial stack that scales with you no matter where you are in your journey.

From corporate cards to maximizing your runway to earning yield on your cash. Brex was designed with founders in mind to make every dollar go further so you can focus on building. And here’s what really stands out to me. Brex combines the best of checking, treasury, and FDIC insurance in one powerhouse account. You can send and receive money globally at lightning speed, earn yield from day one, and still access your funds whenever you need. Plus, with 20 x the standard protection through program banks, your cash is not just working harder, it’s working safer too.

It’s no surprise that one in three venture backed startups in The US with companies like Anthropic, Coinbase, and Robinhood. I mean, my god, these companies are incredible. Trust Brex to help them grow. If you wanna join the smartest startups on the planet, head over to brex.com/startups and see what they can do for you. And talking about trust, today customers expect it faster than ever, and that’s why over 10,000 global companies trust Vanta. Vanta automates up to 90% of the work for in demand compliance standards like SOC two, ISO 27,001, and more using smart AI to centralize work flows, manage risk, and get you audit ready in weeks, not months.

So you can stop chasing paperwork and start closing deals. And a new IDC report found that Vanta customers achieve $535,000 per year in benefits. That’s insane. And the platform pays for itself in three months. I had no idea about these. Whether you’re growing fast or just getting started, Vanta connects you with trusted auditors and experts, support to help you build trust with customers. Get a thousand dollars off your first year at vanta.com/20vc. That’s vanta.com/20vc. You have now arrived at your destination.

Conversation

Harry Stebbings4:50

Jonathan, you’ve just been told by our team that our last show was the most successful of the year when it came out, so there’s no pressure at all that this is gonna be the most successful of this year. But welcome to the studio, man. Thank you. It’s great to have you here, dude. Now I I wanted to start with a understanding of where we are. It seems the world moves faster than ever before. And, honestly, I think a lot of us are trying to understand where everyone lies in a new market.

If we’d look at the current state of the market today, how do you analyze it?

Jonathan Ross5:23

Are you asking is there a bubble? Relatively. In terms of whether or not there’s a bubble, my answer is if you ask a question, you keep not getting an answer, maybe you should ask a different question. And so instead of asking, is there a bubble? You should ask, what is the smart money doing? So, what is Google doing? What is Microsoft doing? Amazon? What are some nations doing? And they’re all doubling down on AI. They’re spending more. Every time they make an announcement on how much they’re spending, it goes up the next time.

One of the best examples of the value that’s coming from the spend, Microsoft in one quarter deployed a bunch of GPUs and then announced they weren’t gonna make them available in Azure because they made more money using them themselves than renting them out. So, there’s real money in the market. And the best way that I think to explain this market is like the early days of oil drilling, a lot of dry holes and a couple of gushers. I think the stat that I heard was 35 companies or 36 companies are responsible for 99 of the revenue or at least the token spend in AI right now.

It’s very lumpy. And so I’m surprised it’s

Harry Stebbings6:26

not less when you look at No, but I mean, seriously, NVIDIA really having concentration of revenue with two clients so heavily.

Jonathan Ross

Yeah, and maybe NVIDIA represents 98% of that, but when it’s that lumpy, what that’s an indication of is it’s like the early days of the oil drilling where people didn’t know how to find oil. They were going off of instinct, almost vibe investing, and people who had a good instinct would make a fortune and everyone else would lose their shirts. Over time, it becomes a science, it becomes very predictable, and there’s less lumpiness, there’s more predictability, but the good investors make less money. So right now is the best time for investors.

Right now, people are making more money than they’re spending. It’s just very lumpy. I’m sorry. They’re making more money than they’re spending. Well, as an aggregate. Plenty of people are gonna lose their shirts, but overall, less money is gonna go in than is gonna come out.

Harry Stebbings7:17

But when we look at the CapEx spend today by the big providers, everyone is going, okay, okay, okay, because there’s something coming at the end of it.

Jonathan Ross

Yeah.

Harry Stebbings

And the trouble is the CapEx spend is going up and up

Jonathan Ross

and up. Okay. You’re thinking of it purely financially, and I think that the financial returns will be positive, but that’s not why people are motivated. So, I was in Abu Dhabi at the inaugural Goldman Sachs Abu Dhabi event. As you now know, we’re sponsoring McLaren, and so Zach Brown was talking, I was talking, it was a fun event, but I was asked a similar question, like, is AI a bubble? And I asked the following question. So, this is like a bunch of people who manage 10,000,000,000 plus in AUM, like 50 plus people who manage 10,000,000,000 plus.

I’m like, who here is 100% convinced that in ten years AI won’t be able to do your job? No hands went up. I’m like, great. That’s how the hyperscalers feel. So, of course, they’re going to be spending like drunken sailors because the alternative is that they’re completely locked out of their business. So, it’s not a purely economical framework that they’re using, it’s a do we get to maintain our leadership? Now, when you look at it the next step, there are these scale law sort of outcomes.

You want to remain in the top 10. We keep talking about the mag seven. If you’re not a member of the mag seven, you’re not going to be able to get anywhere near the valuation. And so, what do you do to stay there? You spend, and it’s worth it because the stock value stays up because you’re in the top seven or 10. At

Harry Stebbings8:44

some point, the returns have to be delivered though. Yeah. The spend has to materialize into actual tangible revenue back, and if it doesn’t, whether you’re in the MAG seven or not doesn’t doesn’t matter.

Jonathan Ross

That’s correct, but right now AI is returning massive value already. It’s very lumpy in the applications, but it’s returning massive amounts of value. Let me talk about an example that actually happened for us. I’ve tried a little bit of vibe coding. I’m not the best in the world at it. We’ve got some interns who are amazing at it. We had this customer visit us, and I had a meeting with them. They asked for a feature, and I spec ed it out very high level, vibe y.

So I was prompt engineering the engineers, and four hours later, it was in production. Not a single line of code was written by a human being. There was no debugging done by a human being. It was all prompting. I think we even have Slack integration now where you commit things through Slack. So all that was done. Four hours later, it’s in production. Think about the value there. But now, fast forward six months from now, when that could happen before the customer meeting’s over. It’s a qualitative difference.

It’s not even just a dollar amount difference. Yes, you know, when you’re able to do it that fast, you spend less to get the feature into production. That’s real ROI. However, qualitatively, when you can do that before the customer meeting is over, you’re gonna be able to win deals that your competitors won’t. Can I ask you, just

Harry Stebbings10:03

going back to the Mag seven, to stay in the Mag seven, do you think everyone realizes that they will need to move into the chip layer and own the full vertical end to end?

Jonathan Ross

I don’t think you’re gonna see too many successfully moving into the chip layer. People look at the TPU as a big success, and what they don’t realize is that there were about three chip efforts at Google at the same time, and only one of them ended up outperforming GPUs. When you look around industry, you’ve got a bunch of people building chips, some of them are getting canceled, like Dojo recently got canceled. Building chips is hard. Going off and saying, I’m gonna build my own AI chip to compete with NVIDIA, it’s a little bit like saying, You know, that Google search, it’s pretty nice.

Let’s go replicate it. It’s insane. Level of optimization, the level of design and engineering that goes into that, you’re not going be able to replicate it with a high probability of success. However, if there’s a bunch of players out there trying to do optionality and one of them succeeds, then you have another chip. We mentioned earlier

Harry Stebbings11:00

that you have to spend if you want to stay in Max seven. Mhmm. NVIDIA investing $100,000,000,000 into OpenAI for OpenAI just to go and buy back NVIDIA chips. Is this not just an infinite money loop?

Jonathan Ross

That would be the case if they weren’t spending it with suppliers to build those chips. It’s not round tripping if actual productive outcomes are occurring. What percentage percentage of the spend is going to building that infrastructure? 40%? So at least 40% of those dollars are actually going out into the ecosystem. So that is not an infinite loop. Okay, so it’s a partial loop. 60 A partial loop,

Harry Stebbings

60% is going back to NVIDIA. Sure. And then they get a bump in their stock price of a couple of $100,000,000,000. Yes. How did you analyze that?

Jonathan Ross

Let’s analyze it in a couple of different ways. From an economic point of view, makes perfect sense. Why not do that all day long? The value accrues if there is lock in. When revenue increases result in stock price increases that are greater than the amount of the revenue, it’s because you believe that that revenue is going to continue, and that’s the belief. And I would actually say, with NVIDIA, that’s probably true. However, it’s not just because NVIDIA is good, and NVIDIA is very good. It’s also because there isn’t enough compute in the world.

The demand for compute is insatiable. I would wager that if OpenAI were given twice the inference compute that they have today, if Anthropic was given twice the inference compute that they have today, that within one month from now, their revenue would almost double. How

Harry Stebbings12:30

would their

Jonathan Ross

revenue double if they had double the compute? Right now, one of the biggest complaints of Anthropic is the rate limits. People can’t get enough tokens from them, and if they had more compute, they could produce more tokens and they could charge more money. And with OpenAI, it’s a chat service. So how do you regulate your chat service? You run it slower, you get less engagement. How important is

Harry Stebbings

speed, you think? There’s a lot of people who think, oh, actually, it’s fine. I’m very happy to have latency and I’m very happy to have a prompt and then I go away, do something else and something happens when I’m away.

Jonathan Ross13:02

Let’s look at CPG, so consumer packaging goods. I want you to rank the CPG goods by margin. At the very top, tobacco, smoking tobacco, right? Below that is chewing tobacco. Below that is soft drinks. Below that, you keep going down, you get to water and other things like that. What is the number one thing that a high margin correlates to in CPG? It’s the speed at which the ingredient acts on you. That dopamine cycle, how quickly something occurs, determines your brand affinity. When something has a very quick response, associate to that brand and then you accrue brand value.

This was the entire basis of Google focusing on speed, Facebook focusing on speed. Every one hundred milliseconds of speed up results in about an 8% conversion rate. So that is wrong

Harry Stebbings

in terms of people’s assessment of the future where they think, oh, it’s fine, we’ll actually just have lots of prompts going on in the background and we’ll be happy to let them run for long periods of time.

Jonathan Ross

100 wrong. In fact, when we first started working on getting speed on our chips, we knew what speed we could get. We even made video example of how fast we could be, and people would look at that video example and they would say, Why does it need to be faster than you can read? And I would respond to that by saying, Well, why does a webpage need to load faster than you can read? And there was just this mental disconnect where people couldn’t grok the sort of visceral importance of speed.

People are very bad at determining what’s actually going to matter in terms of engagement, in terms of outcome, but we know this from building the early internet companies.

Harry Stebbings14:37

Do you think OpenAI will be able to move into the chip layer? At some point, NVIDIA must be concerned that the OpenAI will want to verdict Claude’s and own the chip layer as well. Do you think they will be able to make that successful transition?

Jonathan Ross

I think one of problems in building your own chip, first of all, everyone thinks that building the chip is the hard part. And then as you do it, you start to realize building the software is the hard part. And then as you do it, you realize keeping up with where everything is going starts to become the hard part. I have no doubt OpenAI will be able to build its own chips. I have no doubt that eventually Anthropic will be building their own chips, that every hyperscaler will build their own chip.

I had this experience when I was at Google where I got a lab tour, and this was before AMD was doing a great job, right? AMD was struggling for a little while and now they’re doing great, but they had built 10,000 servers and those 10,000 servers of AMD chips, I was walking through the lab and they were pulling the servers out of the racks, taking the AMD chip, popping it off and throwing it in a trash can. The funny thing was it was almost preordained because everyone knew that in that generation Intel was going to win.

So, why did Google build 10,000 servers? Because they wanted to get a discount on the Intel chips they bought. And when you’re at that scale, the cost to design your own server, because they had to design their own motherboard in order to fit the AMD chip, and to build that out and test it, versus the discount that you get, totally worth it. You have to think of what all the motivations are when people are building their own chips. It’s not just because they’re going to deploy that chip in mass production.

NVIDIA effectively has a monopsony on HBM. A monopsony is the opposite of a monopoly. It’s when you’re a single buyer. There’s a finite amount of HBM capacity, is which the high bandwidth memory that goes into the GPUs. The GPU itself is made using the same process that’s used to build the chip that’s in your mobile phone. If NVIDIA wanted to, they could build 50,000,000 of those GPU die per year, but they’re going to build about 5,500,000 GPUs this year, and the reason is because of that HBM, because of the interposer that it goes on, there’s just a finite capacity.

So, what happens is a hyperscaler comes in and says, I want a million GPUs, and NVIDIA is like, Sorry, I’ve got other customers, and the hyperscaler says, No problem. I’m going to build them myself. And then all of a sudden, those GPUs are found by NVIDIA to give to the hyperscaler. There is just a finite amount of capacity. By building your own chip, what you really get isn’t your own chip. It’s that you get control over your own destiny. That’s the unique selling point of building your own chip.

What does that mean, control over your own destiny? NVIDIA can’t tell you what your GPU allocation is. It may cost you more to deploy your own chip because it’s not going to be quite as good as NVIDIA’s. Let’s think about why NVIDIA’s GPUs, with a slight edge over AMD’s GPUs, dominate. If your total cost to deploy is a huge multiple of the cost of the chips in the systems, then a small percentage increase in the cost of the chip is negligible.

If I’m going to deploy a CPU and that CPU is 20% of the BOM and I get a 20% increase in the speed of the chip, that is a 20% value increase in the entire system versus a 20% increase in the chip cost, right? It’s negligible. So you get these huge multiples when you improve the chip performance. So small differences in performance make a huge difference in the value of the product. A small edge gives you a massive edge in

Harry Stebbings17:56

selling product. Is it possible for OpenAI, Anthropic, any of the MAG seven, any of the other providers to move into the chip layer if there is a monopsony on the HBM market?

Jonathan Ross18:07

It’s very hard. However, there is an incentive from those building HBM to spread that around because NVIDIA gets to negotiate very good rates because they’re such a large buyer. However, if you’re building an HBM fab and packaging house and all of this other, you know, part of the ecosystem, if NVIDIA comes in and writes a big check, then you’re going to build the fab for them. So NVIDIA is always going to get the amount of supply that they want in advance. The problem is you have to write that check more than two years in advance.

Where AI’s gone, you know, just absolutely hockey sticking, even when you have the cash flow of NVIDIA, it’s hard to actually write the checks for the amount of demand that there’s going to be in advance. There is going to be a supply constraint and it’s not purely based on being a monopsony. Part of it is based on just the sheer capital cost and the memory suppliers are very conservative. There’s also this situation where the margin on HBM is so high that no one wants to actually increase the supply because then the margin goes down.

Harry Stebbings19:09

When you look at that and when you look at OpenAI, when you look at Anthropic having their own chips, is that why they’re raising the money they are? Sam said they’re going to need hundreds of billions of dollars. Is that factoring that in?

Jonathan Ross

No. So buying a system is expensive. Buying a data center is more expensive. The reason is you’re amortizing that data center over a longer period of time. So, even if a data center was going to be one third of your cost per year, if you’re amortizing that data center over ten years and the chips over three to five years, the data center is going to end up costing you more per year. When you hear the hyperscalers talking about that $75,000,000,000 to $100,000,000,000 a year investment, because they are building out the capacity for data centers, they’re putting a lot of money up for returns that they’re expecting over the next ten plus years.

So it’s actually not that much money when you think about it.

Harry Stebbings

Are we thinking about amortization in the right way in a three to five year cycle if chip cycles are actually faster than that?

Jonathan Ross20:07

People are definitely thinking about it over a longer period than I would. We use a more conservative number internally. I think five to six Which years would be like three years? A little bit less. We’re looking at upgrading chips about once a year. Now, here’s the way to think about it. There’s two phases of the value of a chip. There’s the am I willing to buy it and deploy it? And there’s am I willing to keep it running? They’re two very different calculations. When you deploy it, you have to be able to cover the CapEx.

When you keep it running, you just have to beat the OpEx. So if I deploy a chip today, I have to beat the CapEx. I have to earn all my CapEx back and make a profit and produce a return. Once I’ve deployed it, as long as I’m beating my operational costs, I’m going to keep that thing in production. So, you’re okay with the value of that chip going down over time. Now, the bet that everyone is making is that those new chips that come out aren’t gonna reduce the value of the old chips below the OpEx.

And in our case, we actually don’t think that five years makes any sense. Because they will be so much less performant

Harry Stebbings21:12

that actually the value will be lower than the operating costs.

Jonathan Ross

For the electricity and for paying for the data center. So what happens then? We just have this

Harry Stebbings

excess supply

Jonathan Ross

of wasted chips which No, are going because a lot of these people have entered into really long contracts, and so they have a third point where they have to consider their calculation, which is breaking this contract is that cheaper than running the chip at a loss. So what happens then? I can’t tell you what happens because we’re trying to avoid that situation. So by having a much faster payback period in all of our calculations, I would not want to make a bet that long out. The shorter the timeframe that you’re making the bet, the clearer your outcome is.

Harry Stebbings

So, essentially, you wanna minimize payback period as much as and then minimize operating cost so that you can shed less performance chips faster.

Jonathan Ross22:00

Yes. But also here’s another crazy part, which is when you look at the math this way, like if I’m approaching it as an accountant, I’m gonna be like, this is a terrible idea. But if I look at it empirically, people are still renting H100s. How old are those chips? They’re getting close to five years old. They’re still earning more than their operating cost by quite a bit. You would never deploy an H100 today, but they’re still profitable to run. They’re in that second phase. The reason is people can’t get enough compute.

If that wasn’t the case, h100s would be renting for a fraction of what they’re renting for today. And as long as you can’t get enough compute, that’s going to be true. The question is, is there an alternative out there that isn’t a supply constraint? And so this is where we’re hoping to come in. Let’s talk about our value proposition. So you started off asking me about speed. Do you know how many customers come to us asking for speed? A 100%. Do you know how many customers keep asking about that once they realize the supply constraint out there?

None. So they start with speed. They know the value of that to their end customer and then they’re like, oh, wait a second. I can’t even get enough compute. The real value prop is can you provide more compute capacity? So two weeks ago, we had a customer come to us and ask for five x our total capacity. They couldn’t get that capacity from any hyperscaler, they couldn’t get it from anyone else. We couldn’t give it to them. No one can. We couldn’t get that customer, the hyperscalers couldn’t get that customer.

There isn’t enough compute. So your choice is, I buy this compute and I get the customer. This is where I was going to you when I said if OpenAI or Anthropic were to double their compute, they would double their revenue. So if you’re someone who can’t get enough compute to serve your customer, then you’re going to be willing to pay whatever it takes to get those customers because you feel that there’s lock in value by getting that customer now. The number one value prop that we have is that our supply chain is not like a GPU supply chain.

You have to write a check two years in advance to get GPUs. For us, you write us a check for a million LPUs and the first of those LPUs starts showing up six months later. Wow, so you’ve got an eighteen month chasm difference. That’s right. So I had a meeting with the head of infrastructure of one of the hyperscalers, and I talked about speed, I talked about cost and all this stuff, but when I talked about the supply chain and how we could do something in six months, he just stopped the conversation for a moment and wanted to dig into that.

That was the only thing he cared about.

Harry Stebbings24:21

Given the speed of progression of the landscape of models,

Jonathan Ross

does two years make sense? So do you know Sarah Hooker? No. So she she wrote this paper, The Hardware Lottery. My TLDR on that one is people are designing the models for the hardware. There are architectures that could be better than attention. However, attention works really well on GPUs. So if you are the incumbent, you have an advantage because people are designing their models for your hardware. Doesn’t even matter if there’s a better architecture out there, it’s not going to run well, so it’s not a better architecture.

There’s a little bit of a loop there. If you are building two years out and you’re the incumbent, that’s okay. But if you’re trying to enter the market, no one’s gonna design for your chips two years out. So you have to have a faster loop.

Harry Stebbings25:07

When you see everyone moving into the chip player, as you said, OpenAI will have their own, Anthropic will have their own, what does NVIDIA do in that world? NVIDIA still keeps selling chips. To who they,

Jonathan Ross

given the concentration of their buyers? So we started off talking about is AI a bubble? If you look for the last ten years, infrastructure for data centers, you’re planning that out two, three, four, five years in advance, right? What happens is everyone’s predictions are wrong, they end up building too little. This has just been what’s happened for the last ten years. So if you don’t build enough for ten years, what do you do? You try and overbuild. You try and build more than your most optimistic projections And then once again, you haven’t built enough, so you increase your projections and you just keep doing this.

That’s what’s been happening and yet people still aren’t building enough compute. Where people’s instincts are off, AI doesn’t work the way SaaS does. In SaaS, you have a bunch of engineers who go out and build a product and the quality of that product is determined based on what those engineers did. That’s not the case in AI. In AI, I can improve the quality of my product by running two instances of the prompt and then picking the better answer. I can actually spend more to make my product better on each query.

I can even decide this customer is more valuable and I’m going to give them a better result. That’s kind of what OpenAI announced when they said, we’re now going to release some products where we can’t really afford the compute, so we’re going to give it to a limited set of users and we’re going to charge more because we want to see what happens when we give more compute to the AI. We want to see what that product looks like and how much better it is. That is going to be our future.

Every time you give more compute to an application, the quality increases. And this is why it’s not coincidental that you see people’s token as a service bill almost matching their revenue because they’re competing for customers and if they just spend more, their product gets better.

Harry Stebbings26:59

But bluntly, the assumption when you look at GPT-five and the focus on efficiency is that Sam transitioned from performance to efficiency because compute does not equal a parallel level of performance improvement. Do you think that is fair and true and does that not go against what you just said?

Jonathan Ross27:16

No. And you have to think of the different outcomes that they’re looking for. So if you are OpenAI, you have moved into markets that are incredibly cost sensitive. Let’s talk about India for a second. If you want to go in India, what’s the one thing you need? 99 rupees a month. That’s about a dollar 13. You need to charge your customer a dollar 13 for your product. So they’re going after a market whose alternative is I have no AI. You’ve got Open. I mean, they can use DeepSeek.

This is another misconception in the market. Let’s just start busting every misconception. Sure. Great. When love the Chinese models came out, everyone reacted by saying, Oh my God, they’ve trained models that are almost as good as The US models, and we a podcast on this, Even I was snookered a little bit at first. Oh my gosh, aren’t these models so much cheaper to run? Now that I know more about the foundation models that people are using versus the Chinese models, no, they’re not cheaper to run.

They’re about 10x as expensive. Actually, let’s just take the GPT OSS model that was released. It’s optimized for something different than the Chinese models, but the quality is very high and I would argue clearly is a better model for what it focuses on than the Chinese models. Now, the Chinese models focus on different things. However, the cost to run the OSS model is about one tenth that of the Chinese models. So why was everyone charging less? When you have sort of a captive market for a model because people say, I want this model and there’s only one provider of it, you can charge 10 times as much.

The price was higher and people were confusing the cost with the price. The Chinese models were optimized to be cheaper to train as opposed to be cheaper to run. When you see how much intelligence has been squeezed into the OSS model versus the equivalent Chinese models, it’s clear that The US still has a training advantage. And the economics work out such that you have to amortize that training over every inference, which means that you want to charge more. And so there’s still a balance there, but as you scale out into larger and larger numbers of people, being able to afford to train a model starts to be a payoff.

As you deploy more inference capacity, you want to spend a bit more on the training to get your inference cost down. In The US, we have a massive compute advantage, and so people train the models harder, bringing the costs down.

Harry Stebbings29:35

Why do we have a compute advantage in The US? Just in terms of access to chips? That’s correct. Will China not just subsidies the inference and the running there? I understand Yes. We does it matter? If that cost of running is higher, but the Chinese the CCP will just subsidize it, does it matter?

Jonathan Ross

There’s a home game and there’s an away game. The home game is we want to build enough compute for The United States. The away game is we wanna build it for our allies. Right? Europe, South Korea, Japan, India, and so on. China can win their own home game. They’re going to build 150 nuclear reactors, they’re going have enough energy even though their chips aren’t as energy efficient, and they can subsidize, as you mentioned. But the away game is different. If a country only has 100 megawatts of power, what are they gonna do?

Build another nuclear power plant? Like, that’s just not a realistic thing. You can do that in China. You can’t do that elsewhere. So having a better chip gives you an advantage in the away game. So my expectation is that right now for the next two to three years, The United States has a clear advantage in that away game over China. And if we move very quickly, then we’re gonna be able to bring a bunch of allies into the AI race. Do

Harry Stebbings30:39

you think we should have open models to allow for China to distill in the effective ways that they have done already?

Jonathan Ross

The model itself is not a clear advantage. So the first time you had me on your podcast, I predicted that OpenAI was about to open source their model. Remember that? And my prediction was based on their branding strength. Frankly, OpenAI could probably use LLM two, the old model from how long ago, like two years ago? Yeah. Yeah. And people would probably still use it. And so, there’s a brand advantage there. Now, they do have very good models, but they don’t necessarily need it because of that brand advantage.

I think that Anthropic should be open sourcing their previous generation in order to get people using them instead of the Chinese models. Because if someone is willing to use a Chinese model, then they would at least be using the Anthropic model and their prompts would be recyclable. And just like you have software compatibility, you have prompt compatibility. For example, when the OpenAI OSS model was released, one of the main reasons people started adopting it over the Chinese models was they could reuse their prompts.

Now, of course, when someone has a low cost application and they can’t afford the premium for OpenAI, they want to use one of these open source models, eventually, they start doing really well, they make more money, they start wanting to get access to the premium model, their prompts are reusable. So there’s a win by open sourcing these models. And you’re also getting all of these infrastructure providers to drive the cost down on that model as well. There’s a lot of innovation that goes into that.

Harry Stebbings32:09

There’s so many different areas I want to take this. We said that just build as much compute as possible. The energy requirements are intense. Is the only way to provide the energy required for this compute wave, tsunami, whatever you wanna call it, is the only way nuclear?

Jonathan Ross

No. Nuclear is efficient and cost effective, but renewables are efficient and cost effective. I’ll give you my my simple hack. All the allies in The United States have order to have more energy than China is to be willing to locate their compute where energy is cheap. Let’s compare Europe to The United States. The United States is incredibly risk averse compared to Europe in everything. But you have to ask what kind of risk. There’s two kinds of risk. There’s mistakes of commission where you do something that’s a mistake, and then there’s mistakes of omission, where you don’t do something and it’s a mistake.

And The United States is terrified of making mistakes of omission. When you’re in a massive growth economy, missing out is more expensive than fumbling something. Europe is incredibly willing to embrace the risk of omission. The way that Europe is trying to compete is through legislation, by saying things like, I want to keep this data in Europe, or I want to keep this data in this country. If Europe wanted to compete in AI, all you’d need to do is say, Norway, please deploy an enormous number of wind turbines.

Why? Norway has about an 80% utilization rate of wind so, like, 80 of the time you can be generating energy. They have enough hydro that if you deployed 5x the wind power of the hydro, Norway itself could provide as much energy as The United States and could do it consistently. The entire United States. That’s one country in Europe. How much other energy is there out there that could be unlocked that isn’t nuclear? And by the way, let’s also deploy nuclear. Nuclear is incredibly safe these days.

Why do we not

Harry Stebbings34:13

then? Fear. Is that really governments, what do they say to you?

Jonathan Ross

I don’t bring up nuclear because I’m not going to push an energy source that everyone’s going to push back on. But when I was in Japan recently, they were talking about bringing their nuclear reactors back online. Japan has a reputation of being very slow. There’s a lack of subtlety and nuance in that perception. The reality is Japan is slow to make a decision, but when they decide something, they move really fast. Let’s take an example. Japan decided to build a two nanometer fab. When I was there last, they were showing off these two nanometer wafers that they produced.

Now, the yield’s not where it needs to be. This is not production grade, but they built a two nanometer fab and they are producing wafers out of it and they’re going to start getting that defect density down. They’re going to move quickly. They’ve allocated $65,000,000,000 for AI and they’re going to spend it, and they’re going to spend it quick. They’re going to turn their nuclear reactors back on. When Japan is going to turn their nuclear reactors back on, Europe needs to listen to that and go, gosh, we need to catch up in energy.

Harry Stebbings35:14

Catch up is exactly what I was thinking because what I’m thinking is the speed it takes to build out. You said about kind of Norway’s latent capacity of wind and how we could utilize it. Dude, it takes years to build huge, huge supply of turbines. You think you the Norwegian government is going to shell out and have 10,000 wind turbines on the coast? Why

Jonathan Ross

does the Norwegian government need to pay for it? Who should? How about the hyperscalers? How about other governments that want to locate there? In Saudi Arabia, there are gigawatts of power and they’re building out data centers for that. Why doesn’t Europe work with Saudi Arabia to say, you know what? So, Saudi Arabia wants to do a program of data embassies where you have sovereign oversight over your data, but you get to use their energy. Why not use that? Problem solved. They’re going to build out three to four gigawatts in the very near future.

Harry Stebbings36:04

So the hyperscalers would pay Norway to use their renewable energy sources and then leverage that?

Jonathan Ross

The complaint that hyperscalers have is all of the paperwork and the slowness. I was talking to someone who was on the board of a major energy company that builds nuclear power plants. He said they spend three times as much on the permitting in The United States than on the nuclear power plant. And I don’t know about Europe, but typically, The United States is better than Europe on this. How much does it cost to build a nuclear power plant in Europe? The actual cost of the infrastructure versus the permitting.

Here’s what everyone needs to walk away from this with. The countries that control compute will control AI, and you cannot have compute without energy.

Harry Stebbings

How far behind is Europe and is there a way for us to get back? Is there a chasm which we can catch up on?

Jonathan Ross

I don’t think there’s a problem right now if Europe acts now. I mean, China is ahead in action, but there are 500,000,000 people in Europe. There’s over 300,000,000 in The US. And if you start bringing all the allies together, South Korea, who by the way knows how to build nuclear power plants. The power plant in The UAE was built by South Korea. They could build power plants here. France knows how to build power plants. How about little a bit of a Manhattan project for building enough energy?

When I’m walking around in Europe in the summer, it’s incredibly hot, and when I’m walking around in the winter, it’s incredibly cold. That is not an experience you have anywhere else in the world. Build more energy.

Harry Stebbings37:33

I’m with you, Jonathan, but I’m also realistic. I know how slow we are as governments, both singular and in collaborating together. It’s not gonna happen at the speed of which this needs to be done. What happens if that does not happen in the speed with which it needs to be done?

Jonathan Ross

Then Europe’s economy is going to be a tourist economy. People are going to come here to see the quaint old buildings and that’s going to be it. You cannot compete in a new economy if you don’t have the resources that the new economy is built on. And the new economy is going to be AI, and it’s going to be built on compute.

Harry Stebbings38:05

Is model sovereignty enough to win, if you look at a provider? No,

Jonathan Ross

because if you don’t have compute, you can’t run the AI. It doesn’t matter how good your model is. You could have a model that is 10 times smarter than OpenAI’s model, and if you have 10 times the compute, OpenAI’s model’s gonna be better. So for

Harry Stebbings

a Mistral who say, hey, we’re gonna have sovereignty within Europe, and the German health carrier system and the Croatian transport ministry are gonna use Mistral because we’re a European alternative. That’s not a reason to win.

Jonathan Ross

What’s the USP? What’s the unique selling point?

Harry Stebbings

It’s a European model and it doesn’t have ownership in The US under a Trump administration.

Jonathan Ross

What you’re solving for there is removing someone else’s ability to control you. Yeah. But what you’re not solving for is having enough of it. By the way, I’m not saying don’t use Mistral. We have a partnership with Mistral. We love Mistral. The thing I’m saying is build enough compute so that Mistral can compete. If you

Harry Stebbings

listen to this, you’re not just like, shit, I should just buy the shit out of CoreWeave. Seriously. Like, when you look at what they provide on demand

Jonathan Ross39:07

Yeah. CoreWeave is a great company, but they have finite allocation of GPUs. Everyone has a finite allocation.

Harry Stebbings

When we chatted before, you said to me that GPUs are not the best infrastructure for inference. Correct. And that we are moving and more into a world of inference as we move further along the maturation cycle of training models. Yes. Does that not mean NVIDIA’s powerhold weakens further?

Jonathan Ross

No. NVIDIA is gonna sell every single GPU that they build. Even if we end up supplying 10 times as many LPUs as GPUs, all that’s gonna do is increase the demand for GPUs and allow them to charge an even higher margin. Because the more inference you have, as mentioned before, the more you need to train the model to optimize for the inference. And the more training you have, the more inference you want to deploy to amortize the cost of the training. There’s a virtuous cycle between the two.

Harry Stebbings

Is the inference market playing out as you expected it maturation, deployment speed?

Jonathan Ross40:05

What I never expected was that AI was gonna be based on language. What that’s done is it’s made it trivial to interact with AI. I thought it was gonna be more like Alpha Go. I thought it was going to be intelligent in some weird esoteric way. The fact that it’s language means anyone can use it. So I expected AI to come sooner and grow slower. It came later and it’s growing faster than I ever imagined. It is so easy to interact with AI that anyone can do it.

Harry Stebbings

10% of the world’s population is a GPT weekly active user. Isn’t that astonishing?

Jonathan Ross

Yes. But you know what’s holding it back? Compute. So compute is holding it back for the quality of it, but more people would use it, they just wouldn’t get as much out of it. But more people would use it if more languages were supported well. This is the number one complaint we hear around the world. You know what would solve that? More compute. More data. If you have more data, then you can train more, but you need more compute. And by the way, if you have more compute, you can generate more synthetic data so you can train more.

So, you have data, you have algorithms and you have compute. If you improve any one of them, it’s not a bottleneck. It’s not like if the compute doesn’t get better, I can’t use more data. If the data doesn’t get better, I can’t use more compute. Any one of these that gets better improves AI and that makes it really easy to improve AI because you can improve one dimension of it. It just turns out the easiest knob to improve in AI is not the algorithms. Algorithms rarely improve.

It’s not the data because it’s really hard to get more data and we haven’t fully figured out synthetic data generation. We’re we’re good at it, but we’re not at the point yet where we can just directly turn compute into more data. We’re getting there. Compute is the easiest knob because it just keeps getting better and better and better every year. If I write a check for enough money and I’m willing to wait a little while, I’m gonna get more compute. It’s the most predictable part of the pipeline.

And yet, we still underestimate how much we need.

Harry Stebbings41:56

Do you think we are dramatically underestimating how much we need today? Yes. By what scale?

Jonathan Ross42:01

Going back to what I said about how every time you add more compute, a product gets better, there is no limit to the amount of compute that we can use. It’s different from the industrial revolution. In the industrial revolution, you couldn’t use energy unless you had the machinery to use it and you had to build machinery and that took time. If I wanted to have more cars on the road, I had to build the cars. It wasn’t enough to just pull more oil out of the ground.

AI is not like that. Yes, if I make my model model better, I can actually do more with the same amount of compute. But if I double my compute, I double the number of users, I improve the quality of the model. This is different. I can literally just add more compute to the economy and the economy gets stronger. We’ve never had that before where it wasn’t a bottleneck, it was more of a rubberneck where you could just force more of one component through and then everything improves.

Harry Stebbings

You said the economy gets stronger. When we think about kind of what that’s predicated on, that’s predicated on the $10,000,000,000,000 labor spend in GDP shifting to AI and us taking a portion of that. Do you think that we will see significant shifts in the GDP or the spend on labor moving towards AI in the next five years?

Jonathan Ross43:09

I believe that AI is going to cause massive labor shortages. I don’t think we’re going have enough people to fill all the jobs that are going be created. There’s three things that are going to happen because of AI. The first is massive deflationary pressure. This cup of coffee is going to cost less. Your housing is going to cost less. Everything is going to cost less, which means people are going to need less money.

Harry Stebbings

So, how is it going to cost less to have a cup of coffee because of AI?

Jonathan Ross

Because you’re going to have robots that are going to be farming the coffee more efficiently. You’re going to have better supply chain management. It’s just going to be across the entire supply chain. You’re going to be able to genetically engineer the coffee so that you get more of it per watt of sunlight just across the entire spectrum. So, you’re going have massive deflationary pressure, that’s number one, and what that means is people will need to work less. That’s gonna lead you to number two, which is people are gonna opt out of the economy more.

They’re gonna work fewer hours, they’re gonna work fewer days a week, and they’re gonna work fewer years. They’re gonna retire earlier because they’re gonna be able to support their lifestyle working less. And then number three is we’re gonna create new jobs and new new industries that don’t exist today. Think about a hundred years ago. 98% of the workforce in The US was in agriculture. 2% did other things. When we were able to reduce that to 2% of the population working in agriculture, we found things for those other 98% of the population to do.

The jobs that are going to exist one hundred years from now, we can’t even contemplate. Hundred years ago, the idea of a software developer made no sense. A hundred years from now, it’s going to make no sense, but in a different way because everyone’s going be vibe coding. Influencers, that wouldn’t have made sense a hundred years ago, but now that’s a real job. People make millions of dollars off of it. So number one, deflationary pressure. Number two, opting out of the workforce because of that deflationary pressure.

And number three, jobs and companies that couldn’t exist today that are going to exist and are going to need labor. We’re not going to have enough people.

Harry Stebbings45:07

It’s fascinating, the counter narrative, isn’t it? Everyone being like, oh, millions and millions of people will be unemployed. And you’re like, no, we’re actually not gonna have enough people for the jobs.

Jonathan Ross

Well, what was the famous prognostication one hundred years ago that there was gonna be massive famine because we weren’t gonna be able to feed ourselves. People always underestimate what’s gonna change in the economy when you improve technology.

Harry Stebbings

When you think about the requirements from an energy perspective, and then also what you just said there about kind of labor, do you think Trump and the Trump administration is doing more to help or to hurt the advancement of AI in The US?

Jonathan Ross

Definitely help. All of the moves that have been made are things that are gonna help with AI. For example, the permitting issues. Overall, it’s been a very positive experience on

Harry Stebbings

mentioned vibe coding. I do just have to ask about it. Do you think this is an enduring and sustainable market? When you look at a lot of these cases today, they’re quite transient. Mhmm. How do you analyze the future of the vibe coding market having played with it a little bit and having seen also interns, as you said, who are very good at it internally use it well?

Jonathan Ross46:08

So let’s take reading. Reading and writing used to be a career. If you were a scribe, you were one of the small percentage of people who knew how to read and write, and people would hire you just record things and you did much better than the average person in the economy because of that, because it was a specialized skill. Coding has been the same thing. Very small percentage of the population did it, took a couple years to learn how to do it well. Some people were really good at it.

Now everyone reads, everyone writes. It’s not a special skill. It’s expected in every job and coding is going to become the same thing. For you to be in marketing, you’re going to have to be able to code. For you to be in customer service, you’re going have to be able to code. I was having dinner with someone who runs a chain of 25 coffee shops, has never coded in their life, and they vibe coded a supply chain tool that allowed them to check inventory. They didn’t write a single line of code.

They got it to work. And it was funny because they discovered all the problems that we software engineers discover over time. They started getting feedback from their employees like, this feature doesn’t work. This thing doesn’t work when I do this. All the little edge cases. And then you just started fixing them all through vibe coding.

Harry Stebbings47:15

Do margins matter in a world of exponential growth? When we look at the demand for your products, when we look at the demand for a Lovable or a Replit, both have bad margins. Does it matter having bad margins when growth demands are so high?

Jonathan Ross

First of all, you do have to have profitability in the end, or at least breakeven to be an ongoing concern. At some point, you can’t just keep raising money. Even Amazon had to start making some money. The real reason why you need higher margins is volatility. Because if you have a razor thin margin and the market moves, you may not be able to raise more money, you may not be able to get a loan. And so, what a margin does is it gives you stability and staying power in the market.

On the other hand, what it does is it also gives competition the ability to enter. Your margin is my opportunity. And so, what you’re trading is stability for a competitive moat. That’s the decision that you have to make. How do you think about margin internally today? I think you want the ability to have margin and you want to give it to your customers, and you want to give them an advantage. And if you have the ability to take that margin when it’s needed, then you’re in a great position.

So we hired this amazing CFO recently, but I remember talking to a previous candidate. When we were talking about margin, they said that we should price so that our supply met our demand. In other words, they wanted to increase the price in order for the demand to come down.

Harry Stebbings48:40

Makes sense.

Jonathan Ross

Does

Harry Stebbings

it? Economic sense, yeah. Economic sense. Logically and rationally, yes.

Jonathan Ross

But then, logically, why not use up your brand equity? Why not use the trust that your customers have to sell them things that aren’t good? Brand value, brand equity has value. You want to keep your brand equity as high as possible because trust pays interest. And similarly, you want to keep your margins low enough that you’re building up this sort of equity value with your customers where they know that you are giving them a good deal. When you charge a high margin, you are at odds with your customer.

You want to do everything that you possibly can to align with your customer. I want my margin to be as low as I possibly can make it while keeping my business stable, and I’m gonna make my cash flow by increasing the volume. One of the things that I love about the compute business is that the need for compute is insatiable. It’s Jevan’s paradox. If we produce 10 x the compute, we will have 10 x the sales. That’s just the way it works. As long as we keep bringing the cost down, people are going to buy more.

I want to keep bringing that cost down, I want to keep increasing the volume, and I want to keep selling more for less so that people get more value out of their business and they buy more and that cycle continues.

Harry Stebbings49:59

How far are we on the journey to bring the cost down? You know, I remember, I look back at some of the shows, dude, and I I cringe at myself because I’m talking about like, oh, Canva implementing AI and it’s hurting their margins because they’re implementing AI and it’s gonna cost them more. And it’s just such a naive approach to ask that question even because now the cost of implementation has gone down by 98%. How far are we in terms of that cost reduction cycle?

Jonathan Ross50:26

Let’s step back and use your Canva example. Successful businesses don’t watch the bottom line, they watch their customers. They solve problems that their customers have. If you are competing, you are doing it wrong. You wanna differentiate. You wanna solve a problem that your customer has not solved yet and can’t solve any other way and then they’re happy to pay you money. That’s how it works. You solve their problem and then your cash flow is solved. Someone’s spending on AI, if you just look at the balance sheet, that doesn’t make sense.

But when the customer is very happy and they’re solving a problem that they couldn’t solve otherwise, first of all, you’re increasing the TAM usually with AI because it makes the product so much easier to use. Did you use Photoshop two years ago? Impossible. Now, if you want to generate an image, just you explain what you want. That increases the TAM. You may be able to charge less per photo, but your total revenue increases. Your total market increases.

Harry Stebbings51:22

Forgive me for this financial question, but we see the S and P about to hit 7,000. We see this ripping of the mag seven like we haven’t seen a concentration of value in in many, many years. People suddenly start to feel like, wow. It’s getting toppy. I listen to you and I hear all of this, and I think, just the start. How should I think about the duality of those two thoughts?

Jonathan Ross

There’s two components to the value. One is the weighing machine and one is the popularity contest. There are some products that are pure popularity contest like crypto. I have never bought a Bitcoin. Why? Because I can’t play in the popularity contest. I’m not good at it. I don’t know what’s gonna be popular and what isn’t. All I can do is I can see value. When I look at AI, I see real value being delivered. Best example, PE firms are all over us. They want access to cheap AI compute because every time they get more cheap AI compute, they can change the bottom line of their businesses.

It has real value. When PE firms go after something and see value in it, it’s not a popularity contest. It’s pure value. And so what happens is the reason companies get a large multiple is people see that the actual value is going to accrue or they get hype cycled on it and it’s pure popularity contest. And there are different participants in the market. Some of them are just playing the popularity contest. Others are looking at the value. And they may come to the same conclusion for different reasons.

Coming at it from the value point of view, the weighing machine point of view, the most valuable thing in the economy is labor. And now we’re going to be able to add more labor to the economy by producing more compute and better AI. That has never happened in the history of the economy before. What is that gonna do?

Harry Stebbings53:07

Do you worry that if we have a speed bump in the short term, it will derail significant parts of the economy given the concentration of value? Everyone rips today, but if NVIDIA, Matter, Google, Microsoft suddenly hit speed bumps and the AI speed train is just slowed down, the consequent multiplier effect is mega. Do you worry about that?

Jonathan Ross

Yeah. And and this is independent of the value of AI. This is the sort of control system theory of what’s going on. Right? So a stock market could inherently be on an upward trajectory. It can overheat, and that overheating causes it to run away. People bid things up, they realize they made a mistake, and then it has to come back down, and then it dips below where it should be, spending retreats, then people don’t have the funds they need to build their businesses. A lot of good businesses can die during one of these downward trends, but this is also where the best businesses are made.

How many times do you see a downturn and a ton of amazing businesses come out of it?

Harry Stebbings54:08

Do you think we will have a downturn in the next year?

Jonathan Ross

I can’t predict whether or not there’ll be a downturn. The ability to predict something is largely dependent on whether or not predictions affect predictions. If a prediction affects the prediction, you cannot predict it because whatever your prediction is changes the outcome. The only things that are predictable are things where the predictions don’t change the outcome. If an asteroid is headed towards the Earth and we see that, if we don’t have the technology to stop it, then it’s going to happen. But if we see that happen and we can predict it, then we might develop the technology stop it.

Do you see the problem? I do. And so in the economy, you don’t have to do anything other than move dollars around. So you have these very sort of fast twitches in the economy based on people’s ability to predict, which makes it unpredictable. I can’t tell you what’s going happen in the economy.

All I can tell you is that right now, the biggest problem I see in AI is if you see a good engineer, one that you would have hired before, they can go out and and they can raise $10.20, 100,000,000, a billion dollars, and then rather than contributing to one of the other AI startups, they go create their own, which means that you have difficulty in getting critical mass of talent in any one of these AI startups. On the other hand, AI is making everyone at one of these startups more productive.

So, in terms of whether or not the economy is overheated, I think one of the best predictors of that is the economy getting in the way of the success of the companies. If it’s not getting in the way, then I don’t think it’s overheated.

Harry Stebbings55:42

Do you not think it is getting in the way? Because fundamentally, the capital supply side is so large that we are actually preventing you from being able to get great engineering teams together because we’re funding talent to the extreme where they can raise huge amounts of money rather than join Groq.

Jonathan Ross

Yes. Please stop doing that. Yeah. No. But but AI is making people more productive. So it might be possible for the economy to keep ripping and for all of the companies to continue being very successful. We don’t know. We’ve never been through this before.

Harry Stebbings56:14

Is the war for talent insane today?

Jonathan Ross

It’s definitely much more aggressive than it’s ever been in history, but only in tech. When you look at sports, sports have always been insane, or at least recently been insane. Like, you look back twenty years, thirty years ago in sports, the salaries looked a lot like tech salaries. Sure. People are just realizing the value. Problem is, in sports, you have a limited number of teams. You might even institute a salary cap and things like this. In technology, we’re not doing that. You have an unlimited number of teams, an unlimited number of startups.

Just imagine if anyone could go create their own football team. What would that do to salaries? And what would that do to the the value of the franchise?

Harry Stebbings

Which incumbent are you most impressed by? And which are you most worried or concerned for?

Jonathan Ross57:00

I would say Google has probably done the biggest turnaround and they had a structural advantage in that. So Google historically has depended more on their engineers to come up with good ideas. And as long as management gets out of the way, great things happen at Google. And so I just think from a cultural perspective, that’s a systemic advantage. For them to Do you

Harry Stebbings

think Gemini has been a success for them ultimately?

Jonathan Ross

I do. I mean, you just look at the numbers of the adoption, it’s been great.

Harry Stebbings

How do you feel about the implementation into consumer products?

Jonathan Ross

Less so. I mean, you see, like, random Gemini introduction into each product. It’s like, it’s in Gmail, but it’s practically unusable. It’s in pretty much every product, and it seems thrown in kind of half thought through, but you shouldn’t judge that yet because at least they’re getting exposure to how people are using it and they can use that to figure out what they should actually do. I mean, what happened with Google Chrome? Right? Like, it was originally Google TV. It was a total flop, and then they iterated and they turned it into Google Chrome.

This is the classic problem where someone puts something out there, everyone throws darts at it, and you don’t realize that they’re just willing to take those darts in order to build a better product.

Harry Stebbings58:10

And it’s fine to take those darts as long as the window of distribution advantage remains. But what’s challenging is it doesn’t. OpenAI has closed that chasm so significantly.

Jonathan Ross

That’s true. Google may be too late.

Harry Stebbings

Do you see what I mean? It’s classic, like, can incumbent attain innovation before the startup requires distribution? And it’s like, the startups acquired distribution 10% of the world. It’s pretty impressive.

Jonathan Ross

Yeah. At this point, it would be hard to imagine a scenario where OpenAI goes away. I don’t see how that happens. And so, at the very least, you have two competitors from this point on going at it.

Harry Stebbings

Which is OpenAI and Anthropic or OpenAI and Google?

Jonathan Ross

OpenAI and Google. Anthropic does something different. Anthropic’s doing coding. OpenAI is doing a chatbot. Google’s doing a chatbot. Google’s also doing coding. Google’s doing everything.

Harry Stebbings

Well, I mean, OpenAI is doing coding too.

Jonathan Ross

Yes, and actually our engineers recently started using codecs more than using the Anthropic tools. Wow. Yeah, and it’s funny because it’s almost on a monthly basis. So, we have a philosophy. We don’t tell our engineers what tools to use. We do tell them, You must use AI because otherwise you’re just not going be competitive, but we saw them using Sourcegraph. We saw them then using Anthropic. We saw them then using Codecs. Next month, it’ll probably be SourceGraph again. It just keeps going around and around in a circle.

Do any of

Harry Stebbings59:29

these have enduring value then if the switching cost is so low and if they’re just bluntly being used so promiscuously?

Jonathan Ross

Our engineers are cutting edge engineers who will switch to the best tool the moment it’s the best tool. Not everyone is like that. A lot are like that though. A lot of the people you interact with are like that. Enterprises make these long term deals and they stick with whatever their deal they made a year ago.

Harry Stebbings

Would you rather invest in OpenAI at 500,000,000,000 or Anthropic at 180?

Jonathan Ross

I’d want to invest in both. Would you? Yeah, they’re both undervalued, highly undervalued. You’re still looking at them as if they’re competing in a finite market for a finite outcome when they’re actually increasing the value of the market with the more R and D that they do.

Harry Stebbings60:11

Play this out for me then. If we do the bull case for them, what does that look like?

Jonathan Ross

I think the current tech companies can increase their value significantly, but I don’t know why they couldn’t increase their value significantly while the AI labs catch up to where those the current technology leaders are. The MAG-seven is going to increase in value, and what’s going to happen is the AI labs are gonna achieve the same amount of value as the current mag seven, but the mag seven is gonna be more valuable. The question is, will the AI labs overtake the mag seven? What will determine that?

Frankly, I think they’re just gonna become the mag nine, the mag 11, the mag 20.

Harry Stebbings

Do you think the AI labs move very significantly into the application layer and subsume the majority of it?

Jonathan Ross

That is the natural tendency of a very successful tech company. They start to do what their customers do and they move up the stack and then they subsume what their customers did and then there are new people who build on top of them. OpenAI, you know, I think on your show, Sam Altman said something about how if you’re just doing like a small refinement on top of Open you’re going get overrun or whatever. He was just being very honest. That’s what they do. In our case, we found an area where we will not compete with our customers, which is we will not create our own models.

We just won’t do it. And by putting that line in the sand, we’re saying it’s safe to build on our infrastructure, right, because we’re not going to go after what you do. That may be the wrong call. We may find that we’re subsumed by one of our customers, but it also means that you can trust that you can build on us. I could be making a huge mistake on that call.

Harry Stebbings61:44

You could be. You would also need a lot of cash to do that. To build our own models. And speaking of cash, how much did you just raise? So we raised $750,000,000 $750,000,000 at a, what was it? $6,000,000,000

Jonathan Ross

Yeah, Almost 7,000,000,000.

Harry Stebbings62:00

Got you. This sounds really unfair, and that’s amazing. Is that enough money?

Jonathan Ross

It is. In fact, we were only gonna raise 300,000,000. You brought up the question of profitability and all that. The hardware companies are in a good position because unlike these other companies, we actually make money off of what we sell. When we sell hardware, those hardware units actually have positive margin. I thought you had negative margin. When we sell hardware, no.

Harry Stebbings

Versus when you sell software.

Jonathan Ross

When we sell software, it depends on the model. So, our most popular models on the chip that we’re ramping up now are positive margin, but we do have some models that we run that beat the OpEx, but we’re not happy with the CapEx. Others would be happy with the CapEx, but we’re more conservative. It’s just easier to say when we sell hardware, we have positive margin because you know it at that moment. We might have positive margin on even our least profitable models because we just don’t know how long the hardware is gonna last.

Like, what are the margins and where do they go over time? Well, one of the benefits of being private is I don’t have to tell you. You don’t, but it’d be

Harry Stebbings63:04

lovely if you did. It’s the only advantage of being private. No, no, no. There’s many, many advantages. You don’t have a lockup period. You can sell much more easily.

Jonathan Ross

Yeah, but I don’t sell shares, so

Harry Stebbings

You’ve never sold a share, have you? Never. No, you clearly don’t understand how this game works. Don’t worry, I will teach you. But margins over time, do they like How do you think about that? I’m not asking necessarily

Jonathan Ross

No, no, I’m gonna say what I said earlier, which is I want our margins to be as low as our business remains non volatile. So, the only reason for a high margin is because you want to have the ability to bring in cash when you need it. And all you need is the ability to price higher if you need to in order to be able to lower your margin. The demand for compute is so high that if someone came to us and said, I need this compute and we have it, they will pay a higher margin, which allows us to charge a lower margin.

Harry Stebbings

Can you help me understand what the chip market looks like in a five year timeline? You said that we’ll have OpenAI, we’ll have Anthropic, we’ll have all the providers having their own chip infrastructure. You’ll also have NVIDIA. They’ll also be, what does that look like in five years’ time?

Jonathan Ross64:14

My prediction is that in five years, NVIDIA will still have over 50% of the revenue. However, they will have a minority of the chips sold, you know, minority share. They might have 51% of the revenue and they might have 10% of the chips sold. Can you help me understand that? Yeah. There is huge value in being a brand. You get to charge more. However, it makes you less hungry and you’re going to start charging high margins and some people are going to pay it because no one’s going to get fired for buying from NVIDIA.

It’s a great place to be in. That business is going to remain incredibly valuable. If you’re invested in NVIDIA, you’re probably going to do okay. However, if you’re looking at it from the customer point of view, when you have customer concentration like we’re seeing where, you know, 35, 36 customers are 90%, 99% of the total spend in the market, they’re going to make decisions less on brand and they’re going to make decisions more on what makes their business successful because they’re going to have more power to make those decisions.

So, you’re going to see other chips being used because those companies are going have enough power to make decisions themselves.

Harry Stebbings65:26

You said you won’t do badly if you’re an NVIDIA investor. One of my friends says, the thing I love about Harry is that, you know, he’s wonderfully charming, but at the end of the day, he goes, that’s great, That’s great. But what about me? Which is very true. Over under on NVIDIA in a five year timeline, 10,000,000,000,000.

Jonathan Ross

I personally would be surprised if in five years NVIDIA wasn’t worth 10,000,000,000,000. The question you should ask is, will Groq be worth $10,000,000,000,000 in five years? Possible. We don’t have the same supply chain constraints. We can build more compute than anyone else in the world. The most finite resource right now, compute, the thing that people bidding up and paying these high margins for, we can produce nearly unlimited quantities of.

Harry Stebbings66:11

What do you think the market does not understand about Groq that you think they should understand?

Jonathan Ross

Oh, it changes every month. It used to be we couldn’t have multiple users, and then we demoed multiple users to people on the same hardware, right? They used to think that we

Harry Stebbings

because of the SRAM structure.

Jonathan Ross

Because of the SRAM. Actually, here’s another one. Still get

Harry Stebbings

You’re

Jonathan Ross

asked

Harry Stebbings

not impressed with my learning from last time. Thank you so much. Dude, I learned so much from you, genuinely. I was like genuinely learning so much,

Jonathan Ross

question I get asked the most is, isn’t SRAM more expensive than DRAM? The answer is yes. A good way to think of it is SRAM is inherently three to four times as expensive per bit inherently.

Harry Stebbings

And just for anyone who doesn’t know, again, SRAM is versus DRAM, super simple.

Jonathan Ross

So, I’ll keep it super simple, but this isn’t technically accurate. SRAM is the memory inside of a chip, DRAM is the external memory. It really has more to do with how you design it. So, SRAM has three to four times as many transistors or capacitors, just transistors for SRAM than DRAM. DRAM is a capacitor and a transistor. SRAM is six to eight transistors. So, SRAM is inherently larger per bit, which means it uses more silicon, therefore it’s more expensive. You’re also deploying it on a more expensive chip, like a three nanometer chip, so it costs you more per unit of area than DRAM.

So there’s a multiple. Maybe it’s 10 times as expensive per bit. The thing is, when we’re running a model like Kimi and we’re running it on 4,000 of our chips, you’re running that Kimi model on eight GPUs. We’re using 500 times as many chips which means the GPUs have 500 copies of that model which means they’re using 500 times as much memory, which means that their cost is higher because even if it the SRAM is 10 times more expensive, they’re using 500 times as much memory in the DRAM.

This is one of those classic problems of looking at it from a chip point of view rather than a system point of view. Everything that we did was actually system point of view and now it’s world point of view. We actually load balance things across our data centers. We’re now at 13 data centers. We have data centers in The United States, in Canada, in Europe, in The Middle East. When you have a world scale distribution, you don’t just make decisions at the data center level. We actually will have more instances of some models in some data centers with different compile optimizations for input or output based on what’s going on in a geography.

We may not even have an instance of a model in a particular data center and we may have it elsewhere and we can load balance that. We’re optimizing at the world level, not at the data center level. What

Harry Stebbings68:41

would you do if you

Jonathan Ross

weren’t scared, Jonathan? I’ll rephrase that to where could I increase risk in the business? Yeah. And where we haven’t. We could double our orders in our supply chain. Yes, we have a six month supply chain, so we can respond to the market faster than anyone else. How

Harry Stebbings

overweight demand are you then supply?

Jonathan Ross69:01

Like I said, last week, someone came to us and asked for five times our total capacity. Here’s the only reason we don’t just completely double down on the If you’re not supply constrained, why can’t you just do that? Because there are thresholds. So, for example, if we had doubled the capacity, we wouldn’t have won that customer. They needed 5x. It’s not enough to have twice as much, we have to have enough. If we double the capacity, do we have enough for those customers?

Harry Stebbings

And so the risks that you could take?

Jonathan Ross

We could just double the rate at which we’re building out supply. With this fundraise, we ended up raising more than twice what we were expecting to raise. And then we were 4x oversubscribed over what we did raise. And so, we could have raised a lot more money, it would have been more dilutive. And I’m trying to be dilution sensitive for investors and everyone else, but on the other hand, we could have just raised more money and we could have just built a ton of compute. The other advantage that we have is, versus anyone else, our cost per token, especially given us speed, is very advantageous.

So, we know that we can charge less than the rest of the market, which matters when you’re trying to build these businesses, not because people are spend conscious. If we lower what we charge 50%, people are going to buy twice as much. They’re spending as much as they’re making because whatever they spend increases the quality of the output. Output. Do you think about

Harry Stebbings70:20

going public at all?

Jonathan Ross

Our focus is purely on execution right now. Whether or not you go public, that’s a completely different game than we’re playing right now. Right now, all that matters is can we satisfy satisfy the demand for compute?

Harry Stebbings

Why do you think Cerebras decided to go public?

Jonathan Ross

Well, they recently decided not to go public.

Harry Stebbings

Dude, I could talk to you all day. I do wanna discuss a quick fire round. I said a short statement, you give your immediate thoughts. Does that sound okay? Yeah. What’s the biggest misconception about NVIDIA today?

Jonathan Ross

That NVIDIA’s software is a moat. CUDA lock in

Harry Stebbings

is bullshit.

Jonathan Ross

Yeah. It’s true for training, but it’s not true for inference. I mean, we have 2,200,000 developers on us now. That’s how many have signed up. How many do Groq have? They claim 6,000,000.

Harry Stebbings71:04

If you were founding Groq today, with NVIDIA at Fortran and AI boom in full swing, what would you do differently?

Jonathan Ross

I wouldn’t do chips. That that chip already sailed. It takes too long to build a chip. The bet that we Does it? So,

Harry Stebbings

the chip providers today that are coming out, we are seeing new chip providers come out where they’re raising like a lot of money from good people, it’s too late.

Jonathan Ross

Yeah. So the reason that I decided to go into chips, I did do Google TPU, but also before I left, I set a record on the best classification model, like ResNet 50, with someone in Google Brain. We did an experiment. We beat everything. And so I could have gone in the algorithm algorithm side. Side. And actually, when we were fundraising, I wasn’t even 100% sure that I was going to do chips. I was, like, thinking maybe we we do something on the algorithm side, especially in formal reasoning, which is good that I didn’t.

But the main motivation to go into chips was the mote, the temporal mote. So a question we get asked by VCs a lot is what prevents someone from copying what we’re doing? And the answer to that is if you copy what we do, you’re three years behind us because it takes that long to go from the design of a chip to a chip in production if you execute perfectly. I’ve done three chips now that are in production or ramping to production. All three were a zero silicon.

Only 14% of chips that are taped out the first time work the first time are A0 silicon. So that means there’s an 86% chance each time that you’re going to have to re spin it. When we built our V2 chip, we actually already scheduled a re spin for it. And we ended up not having to do it because to our shock, the first one worked. Like, shouldn’t expect that. So that three years is if everything goes perfect. NVIDIA typically takes three to four years per chip and they just have multiple being done at a time.

Groq is now in a one year cycle. So, a year after our V2 is our V3, and a year after that is our V4.

Harry Stebbings73:00

How do you evaluate the meteoric reacceleration rise of Larry Ellison and Oracle?

Jonathan Ross

Brilliant business decisions and the willingness to move fast. Most people right now keep asking themselves, is AI overheated? Should we double down on this? They just went for it. They’re just aggressive, and that’s what it takes to win. When everyone else is fearful, you should be greedy, and when everyone else is greedy, you should be fearful. And right now, there’s a lot of fear in and around AI. What you’re seeing though is there’s a couple of greedy, really smart people and they’re making tons of money and it looks like there’s a lot of greed out there.

It’s just a handful of people that are moving fast.

Harry Stebbings

Where should I be greedy and where should I be fearful? I’m an ambassador today, obviously.

Jonathan Ross

Wherever there’s a moat. You know, Hamilton, Helmer, 7 Powers, right? Wherever you see a moat, you should be greedy. Very few people have a moat. Yeah, and especially at the stage that you invest in. So you have to predict that there’s going to be a moat.

Harry Stebbings

And if there is a moat, it’s a billion valuation for a pre seed.

Jonathan Ross74:00

I mean, there’s a billion, you know, valuation for a pre seed pre moat. That’s what you should You should call it pre moat. That’s what the investors should denote it as.

Harry Stebbings

What have you changed your mind on in the last twelve months?

Jonathan Ross

Oh my gosh. It’s not so much that I’ve changed my mind, it’s that I’ve changed what percentage of our business doubles down where. Every month, we become more focused. Say yes to fewer things. What happens is the business just does better. I used to think that the most important thing was preserving optionality. Now I think it’s focus. However, I think having that optionality early on was crucial so that we could play where we would be most successful. Now it’s about focus.

Harry Stebbings

We’ve spoken a lot about OpenAI Anthropic. Do you think Elon Musk is able to pull it off with Groq and Axe?

Jonathan Ross

Yes. Although it’s probably gonna be different. Whenever a new area emerges, a bunch of people think that they’re competing and they’re not. All of these people creating foundation models think that they’re competing for the exact same thing. What did Anthropic do that was brilliant? They decided to stop competing by doing everything and focus on coding, and that’s worked great for them. If you look at xAI, they have a social network and they’ve integrated their chatbot with that. I’m not going to use that chatbot for solving deep analysis or deep research problems.

I’m not going to use it for coding. Now, they do have a coding model, but they don’t have a coding distribution. Can they use that distribution to get into coding? Maybe. But then they’re not going to be as focused. So what are they doing? Eventually, the markets will diverge. Mag seven. All of those companies have some overlapping business, but the primary business of each of those Mag seven companies is different. If you do not differentiate, you die.

Harry Stebbings75:47

When you look at Google, Microsoft, and Amazon, you can buy one and you can sell one. What’s she buy? What’s she sell?

Jonathan Ross

It depends on the time frame. In the short term, Microsoft is resetting a little bit because of the OpenAI relationship. Long term, they’re probably going to do fine again. Think Do you

Harry Stebbings76:03

think that’s a material damage to them?

Jonathan Ross

No, that’s why I’m saying in the short term, I think it’s going to hit them, the long term, it’s not.

Harry Stebbings

Have they not done majestically well from that? They have the financial ownership of OpenAI and then they have the flexibility to use Anthropic for most of Suite.

Jonathan Ross

And they’ve deployed an enormous amount of compute. So if OpenAI diversifies and gets their compute elsewhere, they have that compute now. Compute is like gold. If you have it, you you have AI. And then Amazon, I think, doesn’t have AI DNA. If you compare them, so you didn’t mention Meta. Right? But Meta and Google always had the AI DNA and Microsoft bought it with OpenAI, but that bought them time. Amazon still doesn’t have that DNA, but they do have compute. Final one. What are

Harry Stebbings

you most excited for? When you look forward, I like to end on an element of positivity. What are you most excited for when you look forward over the next five to seven years?

Jonathan Ross

I think the things that scare most people are what excite me. And what I mean by that is everyone’s afraid of what AI is going to do. I think there’s a good historical analogy here, which is Galileo. A couple hundred years ago, Galileo popularized the telescope. He got in a lot of trouble for that. And the reason he got in so much trouble was the telescope allowed us to see some truths and allowed us to realize that the universe was larger than we imagined, and it made us feel really, really small.

And over time, we’ve come to realize that while we may be small, the universe is grand and it’s beautiful. I think over time, we’re going to realize that LLMs are the telescope of the mind, that right now they’re making us feel really, really small. But in a hundred years, we’re going to realize that intelligence is more vast than we could have ever imagined, and we’re going to think that’s beautiful.

Harry Stebbings77:45

Jonathan, dude, I I always end up taking copious notes in our conversations. Thank you so much for doing this with me, man. It’s so lovely to do it in the studio, and you’ve been fantastic. Thank you.

· Sponsor read0 min · 605 words
Harry Stebbings

It was so special to have Jonathan in the studio there. And if you wanna watch the episode, you can see it on YouTube by searching for 20 VEC where you can find all of the episodes live in the studio. We’d love to see you there. But before we dive into the show today, I love seeing the team come together to make this show happen. What I don’t love is trying to keep track of all the information, the data, and the projects that we’re working on across dozens of platforms, products, and tools.

That’s why we use Coda, the all in one collaborative workspace that’s helped 50,000 teams all over the world get on the same page. Offering the flexibility of docs with the structure of spreadsheets, Coda facilitates deeper teamwork and quicker creativity. And their turnkey AI solution, the intelligence of Coda Brain, is a game changer. Powered by Grammarly, Coda is entering a new phase of innovation and expansion aiming to redefine productivity for AI era. Whether you’re a startup looking to organize the chaos while staying nimble or an enterprise organization looking for better alignment, Coda matches your working style.

Its seamless workspace connects to hundreds of your favorite tools, including Salesforce, Jira, Asana, and Figma, helping your teams transform their rituals and do more faster. Head over to coda.iotwentyvc right now and get six months off the team plan for startups for free. That’s coda, coda,.io/20vc, and get six months off the team plan for free, coda.io/20vc. And while coda keeps your team aligned, Radix makes sure your startup’s name is just as sharp. This one’s for all you tech founders out there. You finally come up with the perfect name for your startup, then you check the.com and, damn, it’s taken.

Parked, unused, or priced like rent in Palo Alto. So you settle with extra letters, weird spellings, whatever it takes. But, hey, you don’t have to compromise because now there’s finally a domain for tech founders like you, .techdomains. Get the startup name you actually want on .tech. No compromises. What’s more, when you use .tech, you signal to your customers and investors that you’re building tech with just your domain name. Isn’t that cool? So if you’ve got a name in mind, search for it now with .tech on a trusted platform like GoDaddy or visit get.tech/20vc to grab it.

You’ve got the name locked down with Radix. Now it’s time to get the fund structure just as solid. If you’re listening to twenty VC, you know we have a really freaking high bar. Well, AngelList is the modern platform used by the best in class venture funds where over 40% of top endowments and banks are LPs. Their customers include a top five venture firm, 20 VC, and they now have, check this out, a $171,000,000,000 of assets on the platform. They combine an all in one software platform with a dedicated service team that moves as fast as you do.

One manager said this awesome quote, AngelList feels like an extension of my fund. Another said, AngelList gives me total peace of mind, the attention to detail, lightning fast response time, and just real sense of ownership from the team are exactly what I need to stop worrying about back office ops. So if you’re starting a new fund, don’t be a moron. Just use AngelList. They’re incredible. Head over to angellist.com/20vc to learn more. As always, I so appreciate all your support, and stay tuned for an incredible episode coming on Thursday with the one and only Jason Lemkin and Rory O’Driscoll.

↑ Top