# AI Chip Wars: How Cerebras Plans to Topple NVIDIA's Dominance

Why We Have Not Reached Scaling Laws in AI · What Happens to the Cost of Inference · How We Underestimate China and Shouldn't Sell To Them with Andrew Feldman

20VC · Mar 24, 2025 · 63 min · 11,549 words
Speakers: Andrew Feldman, Harry Stebbings
Source: https://www.996.fm/episodes/20vc--ep-fcbd3b23/

## Cold open

**Andrew Feldman** [0:00]:

Our AI algorithms today are not particularly efficient. In a GPU, most of the time, it's doing inference. It's five or 7% utilized. That means it's 95 or 93% wasted. We won't be as dependent on transformers in three years or five years as we are now. 100%. The fundamental architecture of the GPU with off chip memory is not great for inference. Now they will continue to do well in inference, but it can be beaten, and I think they know it.

**Harry Stebbings** [0:29]:

This is 20 VC

## Intro

**Harry Stebbings** [0:30]:

with me, Harry Stebbings. Now we did a show with Jonathan Ross at Grok, and it blew all numbers out of the water. Millions of plays. Everyone loved it. And everyone said that we had to get Andrew Feldman from Cerebras on the show. So I'm so excited to make this episode happen today. Joining us in the hot seat is Andrew Feldman, cofounder and CEO of Cerebras, the fastest AI inference and training platform in the world. Now in September 2024, the company filed to go public off the back of a rumored $1,000,000,000 deal with G42 in The UAE. They challenge NVIDIA in the inference market. Andrew is the leading expert for all things inference. This show was I have the best job in the world. I sit down with the smartest people and learn from them, and this show is exactly that. But before we dive in today,

## Sponsor read

**Harry Stebbings** [1:18]:

turning your back of a napkin idea into a billion dollar startup requires countless hours of collaboration and teamwork. It can be really difficult to build a team that's aligned on everything from values to workflow, but that's exactly what Coda was made to do. Coda is an all in one collaborative workspace that started as a napkin sketch. Now just five years since launching in beta, CUDA has helped 50,000 teams all over the world get on the same page. Now at twenty VC, we've used CUDA to bring structure to our content planning and episode prep, and it's made a huge difference. Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time. With Coda, you get the flexibility of docs, the structure of spreadsheets, and the power of applications all built for enterprise, and it's got the intelligence of AI, which makes it even more awesome. If you're a startup team looking to increase alignment and agility, Coder can help you move from planning execution in record time. To try it for yourself, go to coda.io/20vc today and get six free months of the team plan for startups. That's coda.io/20vc to get started for free and get six free months of the team plan. Now that your team is aligned and collaborating, let's tackle those messy expense reports. You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on a month end panic when you realize you have to reconcile it all. Well, Pleo offers smart company cards, physical, virtual, and vendor specific. So teams can buy what they need while finance stays in control, automate your expense reports, process invoices seamlessly, and manage reimbursements effortlessly, all in one platform. With integrations to tools like Xero, QuickBooks, and NetSuite, Pleo fits right into your workflow, saving time and giving you full visibility over every entity, payment, subscription. Join over 37,000 companies already using Pleo to streamline their finances. Try Pleo today. It's like magic, but with fewer rabbits. Find out more at pleo.io/20vc. And don't forget to revolutionize how your team works together. Rome, a company of tomorrow runs at hyper speed with quick drop in meetings. A company of tomorrow is globally distributed and fully digitized. Instantly connects human and AI workers. A company of tomorrow is in a Rome virtual office. See a visualization of your whole company, the live presence, the drop in meetings, the AI summaries, the chats. It's an incredible view to see. Rome is a breakthrough workplace experience loved by over 500 companies of tomorrow for a fraction of the cost of Zoom and Slack. Visit Rome, that's or.am, for an instant demo of Rome today. Nobody knows what the future holds, but I do know this. It's going to be built in a Rome virtual office, hopefully by you. That's roam ro.am for an instant demo. You have now arrived at your destination.

## Conversation

**Harry Stebbings** [4:23]:

Andrew, it is such a pleasure to meet, man. I've wanted to do this one for a while. I've heard so many good things from Eric for a long time. So thank you so much for joining me.

**Andrew Feldman** [4:30]:

Harry, thank you for having me. I appreciate it.

**Harry Stebbings** [4:32]:

Not at all. This will be a fantastic conversation. I have my pen ready. I feel like this is gonna be a learning experience for me. I wanna go back to 2015. What did you and the team see in the AI landscape in 2015 that led to the farming of Cerebras? We saw the rise

**Andrew Feldman** [4:47]:

of a new workload. This is every computer architect's dream. We saw a new problem to solve. What that means is maybe you can build a new machine better suited to that problem. And so in 2015, and the credit goes to Gary and Sean and JP and Michael, my co founders, They saw on the horizon the rise of AI. And what that meant was there'd be a new problem for computers, that what the AI software would ask from the underlying chip processor would be different. We came to believe that we could build a better machine for that problem. That's what we saw. You know, obviously we didn't see it exactly right. I underestimated it. You know, this is my fifth startup and the first time I underestimated the size of the market by a lot. But what we did get right was that this was going to be big and it would put a different type of pressure on a processor, and that it would put pressure on the memory bandwidth, that it would put pressure on the communication structure. That's what we saw, we dove in, it's been an extraordinary nine years.

**Harry Stebbings** [5:57]:

How does the movement into an age of AI change the requirements from a chip perspective of what is needed for a provider and how that then resulted in how you built Cerebras? The way to think about

**Andrew Feldman** [6:09]:

a chip is it does two things. It does calculations and it moves data. This is what a chip does. Sometimes along the way it stores data. And so what AI presented was a very unusual combination of challenges. First, the underlying calculation is trivial. It's a matrix multiplication. And an FMAC can be developed by any second year electrical engineering student. So you say to yourself, holy cow, this has a huge number of very, very simple calculations. The hard part with AI work is results and intermediate results have to be moved a lot. Therein is the most complicated part. They have to be moved to memory and from memory, and they have to be broken up and moved among GPUs. And what we saw was that this was going be the hard problem, and that if we could solve for that problem, we would build an AI computer that was

**Harry Stebbings** [7:07]:

faster and used less power. When we think about how we're going to build and what we're building for, to me, kind of a couple of core elements, is like, where are gonna focus? You focusing on, you know, fine tuning? Are you focusing on training? Are you focusing on inference? Three. You chose all three. Yeah. Why? And I'm sorry for my base questions, but I thought like GPUs were specialised towards training and they weren't specialised towards inference. Can you have a mono architecture that does three best?

**Andrew Feldman** [7:40]:

The first step in computer architecture is deciding what you're not going to do. What are we not going to be good at? Is really the first important question. To answer your question, you say, Is the computational work for training from scratch different from fine tuning? And the answer is, it's not different. It's approximately the same. Now, inference and training have some different requirements, and generative inference in particular has some very challenging requirements on exactly the communication dimension that I mentioned. In generative inference, you have to move all the weights from memory to compute to generate a single word, and you have to move them again to generate the next word and again. So if you have a 70,000,000,000 parameter model, not a giant model, and each weight is 16 bits, you're moving, what, 140 gigabytes of data to generate one word. This is an enormous amount of data movement across memory, and that's called what's consumed that needs is memory bandwidth. If you have an architecture like we saw in the GPU, that is your fundamental limitation. It's a fundamental architectural limitation. That was what we went to Wafer Scale to solve. They use memory, a memory called HBM, the type of DRAM. It is phenomenal memory, but it's slow and high capacity. And when they set the architecture for graphics, that's what you wanted. You didn't have to go back and forth to memory very often. SRAM, on the other hand, is unbelievably fast but has low capacity. And so we wanted to use SRAM, but if you build a normal sized chip, you can't hold a model. And so by going to wafer scale, we were able to put down a huge amount of SRAM and get the benefits of speed and enough capacity. If you build a normal sized chip with SRAM and you want to do a 400,000,000,000 parameter model in inference, you might need 4,000 chips. Or if you want to do a DeepSeek six seventy one, you might need six or 8,000 chips. What an administered nightmare. And if you can keep it on as much as you can on one wafer, two wafers, or four, or 10, you get all the benefit of the SRAM, and because you've been able to use the wafer, you get this tremendous capacity as well.

**Harry Stebbings** [9:59]:

Can I ask you first? I totally get you on HBM and kind of the slowness of it. Why is it then that bluntly so much of the market just continues to use it and 40% of NVIDIA's revenue is using DeadShips for inference?

**Andrew Feldman** [10:13]:

Unless you went to wafer scale, there wasn't really a credible other choice. This is the way GPUs had always been made. It's called a graphics processing unit. That's the way they were built. It was part of their advantage against a CPU, was they were built this way. But now they're dedicated chips like ours. What used to be their advantage is now their weakness. That's a fun market to be in when, over a very short period of time, what you're good at becomes your weakness.

**Harry Stebbings** [10:38]:

With a market cap like they do, and with Janssen as good as he is, which I'm sure we both agree with, they must know

**Andrew Feldman** [10:44]:

this. They do know this. A, they don't make memory. So they're a consumer of other people's memory, and that's SK, the highness guys, or Samsung, I mean, Micron. There are only three or four or five companies that make huge amounts of memory. Not many choices. But it's part of a complex architectural trade off. On the flip side, you could say it's worked really well for them. Right? Look at where it's taken them. But in comparison to those of us who do wafer scale, it's a small set. It's a set of one, us. We have a real advantage against them on inference.

**Harry Stebbings** [11:16]:

How do LPs fit into this? We've got HBM, we've got SRAM with you and Botany having many more of them to make it work and scale. Where do LPUs fit into this mix?

**Andrew Feldman** [11:27]:

In our business, there are a lot of ways to skin a cat. Our way is different than NVIDIA's way. It's different than the TPU. It's different from Tranium. They're different. Right now, and every day since August 26 when we launched inference, our way has been the fastest way across a whole set of models tested by artificial analysis and others.

**Harry Stebbings** [11:47]:

Can I ask, when we think about kind of that speed, I am actually you said that kind of your one on one with wafer and kind of the architecture associated? What does that mean in terms of cost? With such efficiency, is it inherently more expensive? And what does that look like from a cost profile?

**Andrew Feldman** [12:02]:

This isn't our first dance. We've been building computers for a long time. When you make a choice like Wafer Scale, you have to weigh the trade offs. We use less power because one of the most power hungry things on a chip are the IOs, are moving data off chip. And so if you are moving data off chip frequently, you're using more power than if you can keep it in the silicon domain on chip. So we knew we would use less power. We knew if you went to wafer scale that you had to solve some problems that people said were impossible to solve, like yield. So we had to invent techniques that allowed us to yield wafers. In fact, we invented techniques that allow us to yield as well or better than others who are building much smaller chips.

**Harry Stebbings** [12:46]:

What is yield and why is it impossible to solve?

**Andrew Feldman** [12:49]:

A wafer begins a 12 inch diameter circle, a slice of silicon, and your chip is punched out of this the way your mother might take a cookie cutter and cut out cookie dough. During the process, at some point, just like your mom might have done, she lifts up the edges and all the little bits are removed and what's left are just the cookies. Those are your chips. Now what happens is there are a set of naturally occurring flaws, and that's like your mother closing rides and throwing up a handful of of M and M's. Now the bigger the cookie, the higher probability you hit an M and M. The bigger the chip, the higher the possibility that you have a flaw. And traditionally, what you did when you had a flaw was you threw away the chip or you sold it as a less valuable part.

**Harry Stebbings** [13:38]:

You

**Andrew Feldman** [13:38]:

shut down part of the chip and sold it as a less valuable part, something called binning. So every wafer is going to have flaws. The bigger your chip, the higher probability you hit a flaw and the more part of silicon is wasted when you throw it away. This is what everybody thought was known truth. And one of the things our team realized was that there are other ways to handle flaws. What if instead you built your computer, you built your processor out of hundreds of thousands of identical tiles and say there was a flaw, say you just shut down that tile and worked around it. Say you had a row or a column of redundant tiles that when you needed them, could just pull in. Now, that had been traditionally the technique used in memory making, and the memory yields are extraordinary. And so it occurred to us that if we could build a computer, build a processor, build hundreds of thousands of identical tiles, we could use redundancy such that when there was a flaw, we could just leave it there, shut it down, work around it, and pull in one of the redundant tiles. And that had never been done in a in a computer before, and that's at the heart of our architecture. That allowed us to yield and deliver whole wafers. Nobody had ever been able to do that in the seventy year history of our industry. Really, really smart people struggled. Mean, Jean Amdahl, one of the fathers of our industry, had a company called Trilogy that crashed and burned trying to do this,

**Harry Stebbings** [15:05]:

and we figured it out. When you speak about being the fastest, and across all benchmarks being the fastest, what matters the most? Is it being the fastest? Is it being the most efficient? Is it being the least costly? How do you think about the stack of prioritisation for your customers? I think it varies. If

**Andrew Feldman** [15:23]:

you go to get a cancer diagnosis on, God forbid, your mother or your wife, I think 93% accuracy is just plain not as good as 94% accuracy. And you pay a lot and wait another week to understand what the accuracy is, right? You pay a lot. Now, on the other hand, if you want Llama 405B to generate data to help you tune Llama 70B, maybe you can wait a few days, three days, a week, more. There's no urgency there. On the other hand, if you want an answer from Perplexity, you don't want to wait forty five seconds for a search answer. You don't want to wait in a chat. You don't want to wait three minutes for R1 on GPUs to give you an answer. What we know is that in interactive mode, milliseconds matter. In interactive mode, what over Google years ago showed was that you can destroy your users' attention with milliseconds of delay. So being the fastest matters everything in that domain. So I think what you have to do is sort of be thoughtful and say in some cases being the fastest doesn't matter. We'll call those batch. Lots of maybe the cheapest matters there. In other domains, there is no search if you've to wait eight minutes to get an answer. That's not a product. When you go fast, a whole set of new opportunities open up. Netflix used to mail DVDs. That's what happened when the Internet was slow. They mailed DVDs. Remember

**Harry Stebbings** [16:50]:

I look young, Andrew. I'm not that young. I I remember Blockbuster. Yeah.

**Andrew Feldman** [16:54]:

Well, if you remember Blockbuster, right. First, I mean, let let's look at the history of that. You're exactly right. First, we used to drive to Blockbuster to get a DVD. Then Netflix was mailing them to us, and then we got broadband. And there suddenly, Amazon's a studio. Right? I it changed everything. And speed in inference does the same thing.

**Harry Stebbings** [17:14]:

When we chatted before, you gave this great equation for inference. What was the equation that you gave for inference? Because it was really helpful for me in understanding.

**Andrew Feldman** [17:21]:

It begins with the following. Training makes AI. That's how we make AI. And inference is how we use or consume AI. And so understanding how big the inference market is is understanding the number of people who are going to use it, how often they're going to use it, times how much compute each use takes. And right now, we are in this rare time where the number of people using AI is growing, the frequency with which they use it is growing, and the amount of compute used in each instance of use is growing. That's why you're getting this extraordinary growth, and that's why it's off the charts right now.

**Harry Stebbings** [17:58]:

When we think about the distribution of resources between training and inference, what will that look like in five years' time? Because we've seen all focus go to training, well, not all, but a lot of focus go to training and not as much go to inference. What does that look like?

**Andrew Feldman** [18:12]:

What we made until the middle of twenty twenty four, what we made in AI was a novelty. It wasn't very useful. Late in 2024, what we made began to be useful. What was the turning point? If you look at the models, they became I mean, ChatGPT was not really a technical innovation. It was a user, user interface invention. But it gave more people access. But we didn't really right away know what to do with it. It was cool. Right? That's what I mean by novelty. Was like, Woah, this is cool. Now, if your marketing team isn't on an LLM each person several times a day, they're not doing their job. That difference between novelty, it's cool, and this is part of everyday workflow. That's what changed starting sometime in Q4 last year and running into this year, is AI became useful, not just to a select group in Silicon Valley, but to my dad, to my brothers, the doctors, to ordinary people who aren't buried in the Silicon Valley discussion. And when you get them, then the market is ripping.

**Harry Stebbings** [19:20]:

Do you not still think we are so incredibly early though? Going back to your point of like how many in five years' time then, where are we? Are we a 100 times bigger? Are we a thousand times bigger in terms of demand? I

**Andrew Feldman** [19:31]:

think we're way over a 100 times bigger. Yeah.

**Harry Stebbings** [19:33]:

What does that mean in terms of what we need to equip ourselves to deliver? These are incredibly energy utilizing. It is incredibly difficult.

**Andrew Feldman** [19:42]:

Our our industry consumes a lot of power.

**Harry Stebbings** [19:44]:

Yeah. And a lot of water, and we're seeing that come down. But are we equipped from an energy and a data center standpoint to deliver the inference requirements for a population that is as AI hungry as we are?

**Andrew Feldman** [19:57]:

I think a couple of things. I think the the first thing is to admit that this is a power intensive problem. We we consume our industry consumes an enormous amount of power. Second thing to say is, therefore, the burden is on us to deliver exceptional value as an industry. You take both, the good and the bad, right? In order to make it worthwhile from a societal perspective to expend all this power, you better deliver the goods. We better use AI to find cures for diseases. We better use AI to solve a bunch of different societal problems. That's the macro view. Do I think that we are equipped? I think we are in a very unusual situation in The US where we have plenty of power, but it's in all the wrong places. We have power in Niagara. What we don't have is power where you want to build data centers, where we have good fiber. What we don't have is a national way to relax the local regulations that make getting power. And so when you go to Silicon Valley, if you want to build a data center, you're dealing with local government and installed interests, and that is not an efficient way to decide if you want to build a power plant or put a new data center in, especially if it's large. I think those places that have ripped out some of that burden in Texas, for example, are getting a huge amount of data centers built.

**Harry Stebbings** [21:14]:

You know, when I spoke to, you know, Jonathan at Grok before, he said there were a huge amount of data centers being built that were not actually really equipped properly, And then we're seeing this massive supply of data centers that are really come down by tourists, so to speak. And that is a massive problem and that the provisioning of these data centers isn't there. Do you agree?

**Andrew Feldman** [21:33]:

A data center is a construction project to begin. It's an access to power, and then it's a construction project, and it's got a design engineering component. I think there's been a huge push for new construction data centers. We will see. We we don't know if they're gonna be good enough. I I think many of them will be fine. The guys who were there early were some of the Bitcoin mining companies, Terra Wolf, the guys at Cursor, and guys in Europe. They were early in building buildings near low cost power in order to run compute that used a lot of power, and they are some of the leaders now in some of the largest projects. Now those are certainly not tourists. Those are extremely sophisticated data center builders. Sure, there's some tourists, but there are a lot of very, very knowledgeable data center builders building huge facilities right now. I mean, gigawatt scale facilities, both domestically and internationally.

**Harry Stebbings** [22:28]:

How do you think about how the cost of inference goes down? With the surge of demand that we mentioned, you know, over 100x, does the price reduce 100x? Does it follow Moore's law continuously? How do we think about the ever reducing price of inference?

**Andrew Feldman** [22:43]:

The cost of inference is built up of several pieces. Right? There's the power and space that is consumed to generate the response. That's a data center cost. That's an OpEx item, number one. Number two, there's the cost of the computer. We can drive down the cost of the computers with each generation by driving up their performance, etc. The other thing we can do is we can develop more efficient algorithms. Our AI algorithms today are not particularly efficient. There's a tremendous amount of room. In a GPU, most of the time it's doing inference. It's five or 7% utilized. That means it's 95 or 93% wasted. Over time, I think as an industry, we get better at things. We can drive the cost of compute down. We can build more efficient data centers with lower PUEs, and our algorithms will get more efficient so that our utilizations on our now cheaper computers are higher, so you get a higher percentage of the maximum number of FLOPs. You get more tokens per unit time for the same power.

**Harry Stebbings** [23:45]:

When you look at the inefficiency of the algorithms, as you mentioned that, and what that means for the utilization of the chips, why are people suggesting that we're at scaling laws already? That seems to suggest that there is so much room for improvement. How do you think about what you just said in conjunction with the idea that scaling laws were hitting this asymptote point? How do you reconcile the two?

**Andrew Feldman** [24:06]:

I don't think there's a lot of debate among senior ML thinkers that we have tremendous room for algorithmic improvement. I don't think there's a lot of debate there. There's even debate about whether the scaling laws are over, whether we ran out of mojo to keep making data or gathering data to fill these ever bigger models. But OpenAI's work on o one shows me that the scaling laws, certainly for inference, are fully functional. Right? The more compute you put on inference, the better answer you get. Many of the leading models are now MOEs. They're not presenting all of the weights to each token, and that's one way to do it. Present the important stuff, not the unimportant stuff. There are other ways to do it that we will invent and learn over time, but we have human models that aren't all to all connected. Many of our models today are all to all connected. That's a lot of unnecessary connections, connections that don't produce anything that we still end up doing math over.

**Harry Stebbings** [25:03]:

I'm sorry. What does all to all connected mean?

**Andrew Feldman** [25:06]:

In many of the layers in a neural network, every element is connected to every other one. That's not the way actually the learning happens. Some are more valuable and some are not valuable at all. Imagine you're gonna read 50 books, you wanna learn something, you can read all 50 books or you could read three books that are really important or you could read summaries of the three books that are the most important. The problem is we don't know which they are at the beginning, and there's a process that you could learn. There's things called dropout and all these other techniques to use sparsity to help solve these problems. We are early in the evolution of AI. Plays right into this point that we'll get better at these algorithms. You know, transformers aren't the end of the world, right? We'll get better. Better will mean faster, more accurate, and more efficient. That's what's exciting about an ever changing industry. That's why I'm not in all these other industries that don't change quickly. Same nine years ago as are today.

**Harry Stebbings** [26:05]:

But this show is kind of strange for me because I speak to a lot of people and they think about the three pillars and they're like, you know, compute algorithms and data. A lot of the common refrain is that actually we're very far along in all of them, and that has been the refrain. And when I hear you, it's like, actually, it's very exciting. Think they're wrong.

**Andrew Feldman** [26:21]:

I think they're wrong. I don't think we're very far along. And it's very difficult to say that we are early in an industry, but we're far along on all its underpinnings. I think we are early in all of them.

**Harry Stebbings** [26:32]:

If we just take them one by one, in five years' time, how much synthetic versus human data will be used to train models if you were to put a percent on it? Almost all synthetic. And the utility value of synthetic is the same as human?

**Andrew Feldman** [26:45]:

When you teach a pilot to fly in a simulator, there is a lot of potential data that isn't very useful in teaching a herd to fly. They've spent a lot of time going straight doing nothing as a pilot. Now, takeoff and landings are where you want to spend your time, and that's why when we put them in simulators, that's what we have them doing. And in simulators, we can create data where engines blow, where there are a whole set of problems, where learning can take place. That's simulated data. And in the same way, as we think about creating data, whether it's for self driving, whether it's for other forms of AI, what we want is the data that's hard to gather, right? Otherwise, we just have a bunch of data of people driving straight on a freeway. Not difficult. We've been able to do that for a decade. What we want is an unprotected left turn in the snow. It's snowing. It's hard to see. You've got an unprotected left turn. That's a difficult thing. And you want that thousands of different ways, millions of different ways. That's where the synthetic data comes along, is to use it to fill in the empty parts where it's really expensive or painful to get that type of data. Think of the pilot. You want them spending a huge amount of time on things that are rare in their training. Same with a surgeon. A huge amount of time on things that are rare. Most of the time, it's carpentry. But their expertise is only when it's rare. Something happens, the unexpected occurs. That's when their metal is shown. And we will get better synthetic data by a great deal.

**Harry Stebbings** [28:13]:

I love it. I get it. From a consumer perspective and from an expectations perspective, if we move the needle on compute algorithms and data, what does that mean for the experience of AI?

**Andrew Feldman** [28:24]:

Faster and cheaper is the first answer. The second is, when things become faster and cheaper, new applications emerge. It's used everywhere, right? When computers became faster and cheaper, suddenly they were in cars, and then they were in your pocket, and then they were in your dishwasher, and in your TV. That's what happens. I mean, we, thirty years ago, you're like, need a computer in my TV? Are you kidding me? I need one in my pocket? Now you've got powerful computers in your pocket, you've got them in your TV, you've got in your kids' toys, you've got in the car. That's what happens. Diffusion of innovation accelerates when you make things faster and cheaper.

**Harry Stebbings** [28:57]:

This is Jevan's paradox and Satya's belief there, mate.

**Andrew Feldman** [29:01]:

Yeah. I know in the VC community, you've cite nineteenth century English economists. I'm English.

**Harry Stebbings** [29:08]:

Come on, if I'm not allowed to cite an English philosopher, what am I here for? Are you just like, oh, he's fucking VC. He's just being like, oh, Javan's paradox.

**Andrew Feldman** [29:17]:

That's right. It's like, make stuff cheaper and faster. There are very few examples in our industry, actually none in compute in fifty years, in which by making things cheaper and faster, the market got smaller. The market always gets bigger, always.

**Harry Stebbings** [29:30]:

Can I ask from an architectural standpoint, you mentioned transformers there, is there a world where we move past transformers?

**Andrew Feldman** [29:37]:

There is a world, 100%. We won't be as dependent on Transformers in three years or five years as we are now, 100%. They're not the end all, be all. Why is it? What will replace it and what does that look like? I don't know. I don't know whether they're going be state based models, I don't know whether they're going to be other types of models, but what I know for sure is that innovation doesn't stop. The transformer has some weaknesses that people are desperate to overcome. There's a quadratic effect in the attention head. There's all sorts of things that could be improved, but it's pretty darn good now. It's the best we have, and that's what you run with. You run with the best you have, and the minute it's not the best you have, you drop it in favor of the best you have. I mean, the number of innovative companies designing models is large. And what DeepSeek showed us is you don't need 5,000 people and billions of dollars of gear. You can do it with 200 smart people. More gear than DeepSeek said they had, but less gear than others had.

**Harry Stebbings** [30:32]:

Were you very impressed with DeepSeek, and what impressed you most?

**Andrew Feldman** [30:36]:

I think it was a result of focused engineering, and that impressed me. It was designed to be better. They weren't confused about being model intellectuals or they weren't confused about whether it was important to break new ground or they were interested in being better. From an invention standpoint, that's a little boring. But from an engineering standpoint, that was sweet effort. They really built a model that was just plain better at many, many things, and that's cool. I like good engineering projects. Now that they chose to announce it right around Trump's inauguration and the politics of it, that's all a separate matter, and we can talk about that later. But Is distillation wrong? No, I don't think distillation's wrong. Is summarization wrong? I'm

**Harry Stebbings** [31:22]:

a VC, you kidding me? That's what we do.

**Andrew Feldman** [31:24]:

If you didn't summarize, you wouldn't know anything. That's exactly right. No, I don't think distillation's wrong. And if distillation is wrong, then certainly using people's copyrighted data is wrong. That's the problem. The problem is you gotta be a little bit consistent.

**Harry Stebbings** [31:39]:

Well, Sam's been a guest many times and we hope he will be again, so I'm not gonna ask you answer these

**Andrew Feldman** [31:44]:

two. No, I think neither are wrong

**Harry Stebbings** [31:45]:

actually, but I think you have to be consistent. Well, thing is with it bluntly, DeepSeek is open. Everything that they did innovate on, OpenAI can learn from and take to.

**Andrew Feldman** [31:55]:

I think there are a few examples of an open source anything having the sort of immediate impact that model had. I mean, that model had a giant impact in a technical community of really smart people, and there are very few examples of other open source software projects that had that type of impact in that amount of time. And, you know, you're you're in the business of betting on these guys. They ramp up and they, oh, look, 10,000, now it's a 100,000 users, it's now a million users, and we better start a company around that, get those grad students. But this had a loud boom in the industry immediately. It was like, woah.

**Harry Stebbings** [32:34]:

The thing I have to think as a venture investor is where is enduring and defensible value simply? And how do I get in early and build that over In hardware, Harry. Well, this was my question. Well, I mean, you have to be a very smart investor like Eric Rischer to do hardware, to be clear. But on the model side, do you think there is value when you look at the sheer number of players all with relatively comparable models?

**Andrew Feldman** [33:00]:

To demonstrate enduring value, you need both immediate value and a trajectory for more. I think the problem is in some industries, you are capable of demonstrating a leadership position for a short period of time. And then someone else, maybe the next generation, they generate the next, and the next generation the next. And I think that ends up in the soft world being you're competing against other people's release cadences. You're four months ahead. They're six months if that's really where you are, there's not a lot of value. But if you can stay at the top over years, right, even if you're not the best, even if you're top decile over years, and the people above you are changing constantly. Very large Silicon Valley companies have been built with not the most compelling technology. It might have started the most compelling technology, and then it got to a point where it was good enough, It was easy enough to use. That's when you're at the mature market. But we're a long way from there right now. Right now we are in the early phases. You characterized my position exactly right. Data, compute, algorithm. I think we have a ton of room for improvement on all of them.

**Harry Stebbings** [34:07]:

When we look, you said that compute and hardware, that's where the value is. How does that value distribution shake out? We've obviously got the 800 pound gorilla that is NVIDIA. How do you think about how the distribution of value shakes out in hardware and in compute over the next five years?

**Andrew Feldman** [34:23]:

Historically, one of the barriers to entry was sort of the capital intensity of a project. And in the world of building chips, there's both scarce resources and expertise and it's very expensive. Historically, it hasn't fit very comfortably in a software company. And the things that modern software companies value are not entirely conducive to chip making. So when I look down the road, who has endured in much of infrastructure tech? People who build systems. Cisco, Juniper. Chip makers have endured. There's a reason that Apple and NVIDIA are among the most valuable companies on earth. There is what they do is hard. That's why it's worth challenging. If it weren't hard, if it wasn't enormous and difficult, why spend time being the underdog and and challenging it?

**Harry Stebbings** [35:15]:

A lot of people place defensibility around NVIDIA's kinda CUDA lock in. To what extent is that real versus hype?

**Andrew Feldman** [35:23]:

In inference, it's not real at all. There's no CUDA lock in in inference. None. You can move from OpenAI on an NVIDIA GPU to Cerebras to Fireworks service on something else to Together to Perplexity with 10 keystrokes. Anybody who actually uses AI knows there's no CUDA locking in inference. I think there was a fundamental effort to disintermediate CUDA, first by Google with tensor flow, and first by some grad students with cafe and some of these early efforts, but later by Google with tensor flow and then Facebook or Meta with PyTorch. I think today, most AI is written in PyTorch and you ought to be able to compile it and run it on your hardware. NVIDIA has many moats. When you are a dominant market share leader, that in itself is a moat. That you're the default solution is a moat. That everybody learns to think about AI in your structures. Those are moats. The software, compilers are hard, but they're tractable.

**Harry Stebbings** [36:28]:

I completely agree with you in terms of kind of being the leader is a moat in itself. It is.

**Andrew Feldman** [36:33]:

It's never talked about that way.

**Harry Stebbings** [36:34]:

Would you put OpenAI in that same, it is the leader, everyone's mother knows ChatGPT.

**Andrew Feldman** [36:40]:

Let's look at Intel, right? Intel has made, until hiring Llabboo prior to that, nearly a decade of catastrophic decisions. And they still own 80% of the x86 market. 75% of the market. AMD is worked up to like 25% or 30%. After a decade of screwing up and you ask yourself, that's a moat. Right? How big is my moat? I can make a bunch of bad decisions for a decade and only lose 20% share. That's extraordinary. The moat was just unbelievable. Now we'll see. I mean, I'm a huge fan of Litbuzz. He's an investor in our company. I I I wish him well and I I think if anybody can change that company, he can. But I think we rarely talk about what being the market share leader means in terms of a mode, in in the right context, because as a challenger, we have to think about it exactly, because it's exactly that that we need we need a bridge for. It's exactly these characteristics of the moat that we need to get over.

**Harry Stebbings** [37:39]:

In five years' time though, is it Uber or is it like AWS and cloud? And what I mean by that is like cloud is an industry market where a a couple of players or several players have relative segments, 25, 30%, and it's shared relatively evenly between them. Not exactly, but relatively. Or is it one like Uber, where Uber has 90%, Lyft has 5%, and then there's alternative providers with the other five?

**Andrew Feldman** [38:03]:

I think it's gonna be between those two. In five years from now, NVIDIA is gonna have 60. Right? I think right now they have approximately all of it. I think they will come down over time. Of

**Harry Stebbings** [38:14]:

NVIDIA's usage, what percent will be training versus inference?

**Andrew Feldman** [38:17]:

I think they will continue to have a meaningful business on both sides. I think they're exceptional at training. They will not roll over and play dead in inference. Think they're a world class company. I mean, they've had one of the great decades of any company in history. Right? I mean, from 2014, they were worth, what, 10,000,000,000 to where they are right now. That's one of the great decades in corporate history. I don't think they're going to roll over and, oh yeah, we're not going to be in the inference market. That's not going to happen. They're going have a meaningful share. But the market's growing, and we'll have a piece. I think others will have a piece. There'll be some very big companies made in this 100x growth.

**Harry Stebbings** [38:55]:

Do you think chip providers will be far larger than model providers in terms of enterprise value? In the five year timeframe? Yes. How does that prediction change in a different timeline?

**Andrew Feldman** [39:06]:

I think in a shorter timeline, you know, when you price an option, variance and uncertainty increases the option's value. If you look at the way Black Scholes works or if you look at any option pricing model, uncertainty is a friend. Variability is a friend of the value of the option. And when people are paying these extraordinarily high prices for our model companies right now, I think part of that is this extraordinary uncertainty, this wild variance. And so in the shorter run, it it might not be the case. But in the longer run, as markets mature, as we begin to understand the value of these models, we understand what their businesses look like, what their long term net profitability looks like. What did Warren Buffett say about markets? In the short term, they're a voting mechanism, and in the long term, they're a weighing mechanism. At some point, the weighing kicks in. Usually, it's in the public markets. And then investors say, which is likely to give me better growth in the future?

**Harry Stebbings** [39:59]:

Mean, listen, you mentioned the word public there. I do wanna just hone in on your business. You're cash flow positive in a world where everyone else literally bleeds cash. Help me understand, what do you do to make you cash flow positive when everyone else is bleeding or hemorrhaging cash?

**Andrew Feldman** [40:16]:

Traditionally, your gross margins were a measure of your technical differentiation, right? If you're running a negative gross margin business, it speaks for itself. You're selling commodity. Your value creation isn't being recognized in the market. And so I think our technology is creating an opportunity for us to maintain margins where some others can't.

**Harry Stebbings** [40:39]:

A lot of your revenue is concentrated to the G42 deal. To what extent is that a strength or a weakness?

**Andrew Feldman** [40:45]:

It's both. The way you catch three large customers is to catch one first. The way you build three large strategic partners is learn to be a strategic partner. That's a learned skill. We didn't arrive knowing how to be a strategic partner at G42. Now that we've worked at it and worked at it, it's a muscle we can replicate. We could be a better partner to any of a dozen different companies in the world.

**Harry Stebbings** [41:05]:

What have you learned in the G42 relationship build process that that may see CUDA a good partner in a way that you weren't? We

**Andrew Feldman** [41:13]:

we've deployed tens of exaflops of compute vastly more than than anybody else that isn't AMD or NVIDIA, right? I mean, at a a huge amount of compute. Our software has been hardened on some of the largest AI clusters in the world. We've gone through the growing pains of increasing manufacturing, 2x and 5x and 2x again, I mean, through unbelievable growth in manufacturing. We've worked with our supply chain partners to be sure that they're ready for this extraordinary growth. When you work with a strategic partner of this size, your organization comes out different on the other side. There are things you've learned and there are mistakes you've made and I hadn't done a big relationship in The Middle East. There was a huge amount to learn. I think you come out a much better company, and much better prepared to do business with a hyperscaler, to do business with another massive partner, to do business with another sovereign. It takes real work, and your team has to learn.

**Harry Stebbings** [42:09]:

You said you'd come out better. Why go public when you did? When this happened, I was like, it seemed preemptive, respectfully. And my question now to companies is why go public at all? There is so much private capital. The Collisons have shown, I think, very clearly that you can stay for a lot longer than you plan to. Data

**Andrew Feldman** [42:27]:

breaks have certainly shown that, right? I mean, those were historically public market valuations. You know, the valuations that Anthropic and OpenAI and some of the others are getting are historically public market only valuations. And like you said,

**Harry Stebbings** [42:42]:

S1's live. Anyone can read it. I wouldn't want people reading mine. We have nothing

**Andrew Feldman** [42:48]:

to hide. Mean, I

**Harry Stebbings** [42:49]:

think No, but your competitors have got asymmetric information.

**Andrew Feldman** [42:53]:

Yeah. We've got asymmetric technology. To be public, you have to be ready organizationally, be ready with your processes, you need to be ready to forecast and predict, to be held accountable in a way that private companies historically haven't been. We think that there's tremendous value. We think that we will be among the first in the category. We think that some of our largest targets would have a stated preference for doing business with public companies. Large enterprises in The US, have done that historically. Those were some of the reasons it led us to.

**Harry Stebbings** [43:28]:

How many G42 relationships, a la G42, will you have in the next twenty four months? How fast can you ramp them?

**Andrew Feldman** [43:37]:

That's a good question.

**Harry Stebbings** [43:39]:

Several. Those are big numbers. Sorry, remind me, how big is the G42? It's 87% of revenue, I know that.

**Andrew Feldman** [43:45]:

It was big. I mean, when we announced it, it was, some estimated, it was north of a billion.

**Harry Stebbings** [43:50]:

Well done. That must be a bit of a high five, doesn't it? Look,

**Andrew Feldman** [43:55]:

think Come on. In the sales team you're like, you know, First, dollars yeah, there's tremendous excitement. And then there's sort of every entrepreneur's reality is, I got to make a lot more gear. Right? I need to make and you make a list of your top 10 vendors and you fly it to the mall and say, Big orders are coming. Be ready. Right? You work with all your partners to get ready because you need to make a great deal more stuff. And that's one of the real differences between hardware and software is when we grow fast, the number of people you need to work with in your supply chain and the amount of collaboration that needs to happen is truly extraordinary.

**Harry Stebbings** [44:37]:

Are anybody going to have a clusterfuck of unhappy customers who bluntly have waited so long for chips? By the time they get them, the chips are outdated and they're going, what?

**Andrew Feldman** [44:46]:

All of that's an opportunity for us and others. That's opportunity. I think being a market share leader isn't easy either. But when you're late, when the bully falls, everybody wants to give them a kick. I mean, that's a lot of that happened at Intel. They'd been the dominant player and when they fell, everybody was happy to jump in and kick them when they were down. I think there is a real opportunity in the potential for NVIDIA customer unhappiness, for sure, for those of us who are competing with them. I mean, if you can't get your gear, may as well test somebody else's. That's a huge opening.

**Harry Stebbings** [45:22]:

Head over to Cerebras and use the promo code Harry20 for your chips today. Do that. There we go. I'm here for you, baby. Influencer mode turned on. No worries. It's fine. I hope if we could do a 20% take on the billion deal, I'm happy. That's

**Andrew Feldman** [45:41]:

fine. I know this venture business hasn't been so good to you, Harry, and you gotta get shoes for your kids and the like. And yeah, we're happy to donate to The the Harry

**Harry Stebbings** [45:51]:

400,000,000 fund and fees when I have no kids as well.

**Andrew Feldman** [45:55]:

No, no, look, 220 is a rough way to make a living, Harry.

**Harry Stebbings** [45:58]:

Dude, you don't get it, okay? You and hardware. You said about the complexity of hardware. Are export controls being implemented properly? Do you think that is a good idea? Everyone was going with DeepSeek, Wow, how did this happen? They must have stolen chips. How could this be? It

**Andrew Feldman** [46:20]:

turns out that they probably did use chips in Singapore. I think the following. Think managing software and managing hardware compliance are extremely different things because their vector of diffusion is different. There's different weights. If you sell a server that weighs 500 or 600 pounds, arrives on a pallet, you can go visit it. You want to deploy it in Kazakhstan, you can put a data center in, you can have somebody from the embassy visit it, take photos of it once a month. It's not going anywhere. You can keep track of who uses it and provide logs and that's much, much harder with software. And open source is whole another. That's the first observation. The second is that I got to know the leadership in commerce in the previous administration. I didn't always agree with their policies, but it is a world of unintended consequences. You sought to limit Chinese access to EDA tools to delay the growth of a Chinese chip market. And so US venture capitalists backed tons of Chinese companies in Shenzhen to build EDA tools. Right. This is an unbelievably slippery, dynamic, challenging problem. I don't know if it's a tractable problem. To delay another nation's progress on a technical trajectory is an enormously challenging thing. I certainly came to appreciate just how difficult it was for well meaning people to predict the impact of policy during the last two years, for sure.

**Harry Stebbings** [48:02]:

Do you think this administration is better for AI than the prior administration? I

**Andrew Feldman** [48:08]:

don't think there's any doubt that's the case. The past administration lined itself up against big tech. That was a mistake. AI is also in a different place, so it's easier to be for it. It's less scary now than it was. We sort of have a better picture of the trajectory, both the risks and the benefits. This administration sort of had the foresight to put in place an AI czar or leader to be a focal point for discussions. Yeah, I think it's probably net a fair bit better.

**Harry Stebbings** [48:39]:

You said it's very challenging to kind of hinder a nation's development, adoption, progression of a technology. Respectfully, you chose to not sell to China.

**Andrew Feldman** [48:49]:

Yeah.

**Harry Stebbings** [48:50]:

Why was that? And does that not go against the difficulty in hindering progression?

**Andrew Feldman** [48:54]:

No,

**Harry Stebbings** [48:54]:

I have

**Andrew Feldman** [48:56]:

a very simple rule and I encourage the team to use it. You don't need a big handbook to help you make good decisions in a company. Just ask yourself, Would my mother be proud? And would she be proud if I did this? Would she be proud if I explained exactly the situation? And would she look at me and say, I'm proud you're doing this, son. And I asked myself that, and I came to believe that the deal on the table wouldn't be used for good, and I wasn't comfortable with that, and I wouldn't have been able to explain it to my mother. And that's a moral compass. Do you mean it wouldn't have been used for good? Do facial recognition, to identify minorities, for persecution, build military equipment, to things that I either couldn't see or what I saw didn't didn't feel right, it's more important than money.

**Harry Stebbings** [49:41]:

Do you think we fundamentally underestimate the Chinese's A

**Andrew Feldman** [49:45]:

100%. And it is one of the most obvious and frequent errors in judgment, is that you underestimate the other side. You have to look carefully at what they're doing and their investment in infrastructure has been extraordinary. The rate at which they generate engineering talent is exceptional. The government's ability to have a policy and implement it, that's not a democracy. They weren't designed to have checks and balances there. The funding that flowed into the development of AI technology, that their venture capitalists were backstopped by their government. They have national champion companies, that they've developed a belt and suspender strategy to sort of make much of the third world dependent on them and their technologies. I think they absolutely should not be underestimated. They have a lot of people, and we see a tiny fraction of it. They have produced industrial policy that has moved their nation forward.

**Harry Stebbings** [50:41]:

What was most significant, do you think?

**Andrew Feldman** [50:44]:

The creation of economic zones like Shenzhen was clearly a visionary move. They knew that their own system was in the way. They created zones that relaxed their own system. Could The US

**Harry Stebbings** [50:57]:

learn from them that way?

**Andrew Feldman** [50:58]:

We did some of the same things in Trump won administration, right? What did we do? We relaxed our own rules in the development of vaccines. We knew that in this time it would be very difficult to go through the steps that we always go through, and we tried to implement some thoughtful workarounds, rather. I think that you know, why are they committed to trains as a mode of transportation and and we can't build a decent train system in The US or in California or why we have three different standards for train rails and the rest of the world can build extraordinary high speed trains linking important cities. What are we doing wrong in the building of our infrastructure that our bridges and our freeways are in disarray? Those are questions we got to ask ourselves when we see other people doing it differently. If you watch a good football team and you say, Woah, that's interesting offense, and you're not thinking to yourself, How could our team learn? What could we do? Why did that work? What was it about the people they had or the talent or the structure or something that made that a successful series of plays? And what can I take away from that? How can that inspire me to do better? I'm always looking inspiration in others and competitors and partners. We have some of our partners at G42. I mean, the work ethic is unbelievable and inspires me. And the scope of the challenge that I'm taking inspires me. I think I'm always looking for that.

**Harry Stebbings** [52:27]:

Andrew, I could talk to you all day. I do wanna do a quick fire with you. So I say a short statement. You ready? Yeah. Sure. What do you believe that most around you disbelieve?

**Andrew Feldman** [52:36]:

I think we're closer to peace in The Middle East than people believe. There is a rise of a moderate of business focused Arab state that wasn't there twenty five or thirty years ago. If you visit The UAE or Qatar or even KSA, what you see is amazing transformation, a desire for to be included in the West in their own way, but also to to to enjoy the benefits of it. We are closer than people think.

**Harry Stebbings** [53:04]:

What's the most underrated threat to NVIDIA's market share dominance? The

**Andrew Feldman** [53:09]:

fundamental architecture of the GPU with off chip memory is not great for inference. Now they will continue to do well in inference, but it can be beaten, and I think they know it.

**Harry Stebbings** [53:20]:

What's the crazy AI prediction you have that most people would call science fiction? Dario at Anthropic said we'll live to 150.

**Andrew Feldman** [53:27]:

I don't think we're gonna live to 150. I don't think 90% of our code will be written by machines in this year. But I do think that within a year or two, most people in The US will engage with an AI every single day in one form or another whether they know it or not. That AI might be in their mapping program that helps them pick a better route to work. It might be any number of different things within a year or two. AI's penetration will be approximately the same as telephones. What have you changed

**Harry Stebbings** [54:01]:

your mind on in the last twelve months?

**Andrew Feldman** [54:04]:

Many decisions I made

**Harry Stebbings** [54:06]:

turned out to be wrong. What was the most wrong decision?

**Andrew Feldman** [54:09]:

There are two ways you can be wrong. You can actively be wrong or you can fight against what was right. In 2016, JP, one of our cofounders and chief system architect, laid out a plan that would have us doing water cooling and for our systems. Nobody else was doing it, and I fought so hard and I was so wrong. JP was right. About a year or two later, Google announced that the TPUs were gonna be water cooled. We were first, and now NVIDIA is only selling water cooled parts. I mean, I was dead wrong and JP was right. Many, many instances when you make a lot of decisions every day where you're wrong. I've been wrong about people. People I thought were pretty good turned out to be extraordinary. People I thought would be extraordinary, were really smart but couldn't finish projects and get stuff done. If you're not prepared to be wrong a fair bit, you ought not to be making a lot of decisions because it comes with the territory.

**Harry Stebbings** [55:01]:

As a venture capitalist, I'm never wrong, so I don't know what you're saying.

**Andrew Feldman** [55:04]:

As a venture capitalist, you're wrong nine times in 10, and everybody forgets as long as you're really right.

**Harry Stebbings** [55:08]:

And I get a picture of you signing the term sheet with me then he's That's in right. Then I go

**Andrew Feldman** [55:13]:

think yours is a perfect industry in which nobody cares about the average. On average, you're wrong all the time. And what they care about is the occasional time you're really right. That's what moves a fund. That's different than being a CEO. I think we've got to be mostly right most of the time. But if you're making a lot of decisions, you're still making a ton of mistakes.

**Harry Stebbings** [55:32]:

This is your fifth startup. I mean, you are a sucker for punishment, aren't you? I mean, like five times. Christ, Andrew, did you not get beaten alive enough? My question to you though is like, I believe in the value of serial entrepreneurship. I've spoken to many who don't. How do you think about the inherent benefits that you have having done it four times before?

**Andrew Feldman** [55:57]:

I think if you are in a business in which running a business is a benefit, then experience matters a great deal. I think if you are in a business in which you look like your customer, there was a reason why social networks were started by people right out of college or in college, because dating is top of their mind, and they look like their customers, and that was more important than knowing anything about running a business. In that environment, it will certainly select for people who are of the demographic that their customers are. They know that backwards and forwards. But if you want to have a business that has manufacturing in it, that has a supply chain, that has you managing hundreds or thousands of engineers to a timeline, to a schedule, I don't think anybody would turn around your statement and with a straight face say, know what I'm looking for is an engineering leader with no experience. Right? No. And I don't want somebody who's led a team of four or 500 who has experienced the challenges of growth. What I'm looking for is somebody with no experience.

**Harry Stebbings** [56:58]:

Naivety is a bonus here.

**Andrew Feldman** [57:01]:

Just right. I think the people who sell that sometimes are consultants, Oh, look, my guys have no experience in your industry. They're not biased. Maybe a little bit of experience in the industry would help, right? Come on.

**Harry Stebbings** [57:13]:

Where are people investing today in AI? Across the stack, you can choose any part where you're like, why is so much cash going to that part? I'm not saying that company, I don't. I think part

**Andrew Feldman** [57:23]:

of the dynamic in your industry is sometimes money needs to find a home. Some guys have raised really, really big funds and they've got to find a home for their money. And some people don't like to be left out, they're willing to make investments for maybe for some status purposes or other reasons that don't seem to make sense. There are some underappreciated places of investment, I'd say, in the chip world, the sub milliwatt, really tiny, tiny little chips that live next to sensors that do inference. These are tiny little things that will only send back useful data. It's an extremely interesting market and they will sell enormous volume. Now, it's not a part of the market I love to play in. I like to build bigger things and sell them to the data center, but I think that part is extremely interesting. I think they'll be fundamental for robotics. That's an area where extremely underappreciated.

**Harry Stebbings** [58:17]:

Final one. If we think about Cerebras in ten years' time, where do you envision the business in ten years' time? If everything goes well, where are we in business having that conversation?

**Andrew Feldman** [58:28]:

So ten years ago NVIDIA was worth $10,000,000,000 That's a long run-in our world right now. I think in three to five years, I would like our technology to have been used to solve two important societal problems. I would like it to be used to have found a therapeutic for an affliction that impacts more than a million people a year. I would like our inference to be powering a collection of apps that don't exist today. And I would like that a meaningful portion of the population in The US and in Europe inadvertently uses our technology. So uses something that we power and that they don't even know. I think those are things that make me really happy.

**Harry Stebbings** [59:10]:

Andrew, I've wanted to make this show happen for a long time. As I said, I heard so many good things from Harry for for many years. There have been so many requests to have you on the show. My team is just like, just get Andrew on the show, Harry. I'm like, okay. Okay. And I tweeted it, obviously, which is how we got this. Thank you for joining me.

**Andrew Feldman** [59:29]:

Tweeted it and like 40 people sent me a note saying, how come you're avoiding Harry? How come he has to go? Tweet it. And I was just like, alright. Just call me. It's good. Send me a note. Happy to come Really thoughtful questions, Harry. Really thoughtful and interesting, a really fun conversation.

**Harry Stebbings** [59:48]:

So I have wanted to do that show for a while, but frankly, I was just blown away by Andrew's humility, his no BS approach. He was incredible to work with in the process, and I just so appreciate his time today. If you wanna watch the full episode, you can find it on YouTube by searching for 20 VC. That's two zero VC on YouTube. But before we leave you today,

## Sponsor read

**Harry Stebbings** [60:07]:

turning your back of a napkin idea into a billion dollar startup requires countless hours of collaboration and teamwork. It can be really difficult to build a team that's aligned on everything from values to workflow, but that's exactly what Coda was made to do. Coda is an all in one collaborative workspace that started as a napkin sketch. Now just five years since launching in beta, Coda has helped 50,000 teams all over the world get on the same page. Now at twenty VC, we've used Coda to bring structure to our content planning and episode prep, and it's made a huge difference. Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time. With Coder, you get the flexibility of docs, the structure of spreadsheets, and the power of applications all built for enterprise, and it's got the intelligence of AI, which makes it even more awesome. If you're a startup team looking to increase alignment and agility, Coda can help you move from planning to execution in record time. To try it for yourself, go to coda.io/20vc today and get six free months of the team plan for start ups. That's coda.io/20vc to get started for free and get six free months of the team plan. Now that your team is aligned and collaborating, let's tackle those messy expense reports. You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on the month end panic when you realize you have to reconcile it all. Well, Pleo offers smart company cards, physical, virtual, and vendor specific. So teams can buy what they need while finance stays in control, automate your expense reports, process invoices seamlessly, and manage reimbursements effortlessly, all in one platform. With integrations to tools like Xero, QuickBooks, and NetSuite, Pleo fits right into your workflow, saving time and giving you full visibility over every entity, payment, subscription. Join over 37,000 companies already using Pleo to streamline their finances. Try Pleo today. It's like magic, but with fewer rabbits. Find out more pleo.io/20vc. And don't forget to revolutionize how your team works together. Rome, a company of tomorrow runs at hyper speed with quick drop in meetings. A company of tomorrow is globally distributed and fully digitized. The Company of Tomorrow instantly connects human and AI workers. A Company of Tomorrow is in a Rome virtual office. See a visualization of your whole company, the live presence, the drop in meetings, the AI summaries, the chats. It's an incredible view to see. Rome is a breakthrough workplace experience loved by over 500 companies of tomorrow for a fraction of the cost of Zoom and Slack. Visit roam, that's or.am, for an instant demo of Rome today. Nobody knows what the future holds, but I do know this. It's going to be built in a Rome virtual office, hopefully by you. That's roamro.am for an instant demo. As always, I so appreciate all your support, and stay tuned for a fantastic episode coming on Wednesday with, I think, one of the most under discussed firms in venture capital, Lead Edge Capital and their founder,
