Cold open
Do you feel like you have enough cash now? I
guess a startup is always fundraising.
Do you think enterprises are ready for open source?
The most technical savvy enterprises are definitely ready for it. In order to widen the adoption, there’s definitely some tooling to be brought to the market.
What are the biggest barriers to Mistral ral today?
We are still bottlenecked by compute for sure, but that’s because we don’t have many of it. We have 1.5 k 800, which is a few percent of our competitors.
Can I ask finally, was it a mistake for you to not scale that quicker?
I mean, you can’t really scale that quicker. You can’t raise, like, 2,000,000,000 on the seed round. I mean, at least you couldn’t in 2023. Which
competitor do you most respect and admire?
We were surprised by recently.
It’s a competitive landscape. We are always mixing it up with new intro styles. Let me know what you think of that new intro style, and what a show we have for you today on twenty VC.
Intro
Mistral is one of the most exciting AI companies today, at the forefront of the foundation model charge. Joining us is Arthur Mensch, co founder and CEO at Mistral, where he’s raised over $520,000,000 in funding from the likes of Andreessen Horowitz, General Catalyst, Lightspeed Venture Partners, and Microsoft. And before founding Mistral, Arthur was a research scientist at DeepMind, one of the leading AI institutions in the world. But before we dive into the show with
· Sponsor read0 min · 431 words
Arthur today, we’re all trying to grow our businesses here. So let’s be real for a second. We all know that your website shouldn’t be this static asset. It should be a dynamic part of your strategy that really drives conversions. That’s marketing one zero one. But here’s a number for you. 54% of leaders say web updates take too long. That’s over half of you listening right now. And that’s where Webflow comes in. Their visual first platform allows you to build, launch, and manage web experiences fast. That means you can set ambitious marketing goals and your site can rise to the challenge.
Plus, Webflow allows your marketing team to scale without relying on engineering, freeing your dev team to focus on more fulfilling work. Learn why teams like Dropbox, IDEO, and Orangetheory trust Webflow to achieve their most ambitious goals today at webflow.com. And speaking of incredible products that allows your team to do more, we need to talk about Secure Frame. Secure Frame provides incredible levels of trust to your customers through automation. Secure Frame empowers businesses to build trust with customers by simplifying information security and compliance through AI and automation.
Thousands of fast growing businesses, including Nasdaq, AngelList, Doodle, and Coda trust Secure Frame to expedite their compliance journey for global security and privacy standards such as SOC two, ISO 2,701, HIPAA, GDPR, and more. Backed by top tier investors and corporations such as Google, Kleiner Perkins, the company is among the Forbes list of top a 100 startup employers for 2023 and Business Insider’s list of the 34 most promising AI startups of 2023. Learn more today at secureframe.com. It really is a must. And finally, a company is nothing without its people, and that’s why you need remote.com.
Remote is the best choice for companies expanding their global footprint where they don’t already have legal entities. So you can effortlessly hire, manage, and pay employees from around the world, all from one easy to use self serve platform. Plus, you can streamline global employee management and cut HR costs with remote’s free HRIS. And, hey, even if you are not looking for full time employees, remote has you covered with contractor management, ensuring compliant contracts, and on time payments for global contractors. There’s a reason companies like GitLab and DoorDash trust remote to handle their employees worldwide.
Go to remote.com now to get started, and use the promo code 20 v c to get 20% off during your first year. Remote opportunity is wherever you are. You have now arrived at your destination.
Conversation
Arthur, I am so excited for this. JC introduced us quite a long time ago now. I’ve known you for a while. I’ve been wanting to make this happen for a while. So thank you so much for joining me today.
Thank you for having me. It’s a pleasure.
The pleasure is mine, my friend. But I wanna start. What would your or how would your parents or teachers have described the young Arthur? I’m just always intrigued by the characteristics and traits of the best founders. How would they have described a nine, 10 year old Arthur?
I guess I was a bit curious and a bit stubborn, not very nice to my to my brothers. I was the healthiest of of them also. I don’t know. You should ask them. I don’t know. I think they they have good memories, hopefully.
Do you know what? Sadly, your mother wasn’t in our reference list, so we missed that one out. But, you know, JC provided some great commentary. So I do wanna start, though, also. You know, growing up, what was your first exposure to AI? You’re a kid in France. How did you first get exposed to AI and machine learning, and what was that first passion point?
That was Andrew Ng flying a helicopter helicopter backwards. It’s a control problem, which is not easy to solve, and I’m not sure if it was really AI related. I think he was saying that he was using a neural network to control all of this, but that’s the well, the first memory of me being shown what you could do with machine learning at the time. That was in twenty twenty thirteen, I think.
Most recently that you spent, yeah, two and a half years, three years at DeepMind. Can I ask what are the biggest takeaways for you from that experience, and how did they impact how you think about building Mistral?
A team of five is faster than a team of 50, except if you organize the team of 50 to be 10 teams of five that are sufficiently uncoupled. One finding that I learned the hard way at DeepMind and the reason why we created the company in a slightly different way in terms of organization of the science team. And also the reason why we knew we we had a chance to do interesting things with a smaller team.
Can I just ask, sufficiently uncoupled, do you not lose efficiency, or is there not a leakage between those silos and it creates actually inefficiency by having such silos?
You have to share some things. So you share the infrastructure, you share the code base, you share findings. We’re doing general purpose models. In general purpose models, you need to evolve them in different directions. So you need to make them speak different languages. You need to make them be able to code, be able to do mathematics, be able to reason. You need to add multimodality to them. All of these things are loosely coupled. It’s useful if you use the same framework for optimization, for data, for training, but you don’t want to have your team spend their entire day in meetings for coordination, and it’s actually pretty hard to figure out.
I think so far, we’ve managed to scale it relatively well. The team is only 25 people, so that’s actually not super challenging. It will become more and more of a challenge. That’s what I remember from DeepMind. It was very well at the beginning. Gemini was a bit too slow, and I think they recover sufficiently well since. We we have optimized the team to be as fast as possible and to ship as fast as possible.
Was it an easy decision to leave to start Mistral? You know, you’re at DeepMind, one of the best institutions in the world for AI with some incredible talent around you. Was it an easy decision? And just take me to that moment when you decided to leave to found or cofound Mistral.
So it’s not a zero to one decision. It’s not a binary decision. You start to think, like, I’m 10% leaning on leaving, and then it grows. And at some point, you cross the threshold, then you say, okay. Well, I guess, know that I’m sufficiently decided there’s no way in which I stay more than a than a few days. Otherwise, I wouldn’t be counted with my colleagues. And so that’s that’s how you you you get started and say there’s no turning back. Need to leave.
What was that point for you?
That point for me was probably around March end of March at the last year where I decided to leave on Friday, and I resigned on Monday. You you can’t stay if you’ve decided to resign. It’s all the way, it’s not very fair.
No. I I totally agree with you. Now I do wanna run this with some chronology. I spoke to so many of your advisers, investors, and I wanna start with actually the first model, you know, Mistral 7B, being one of the most popular, released, you know, a while ago now. Why do you think it was so popular? What do you think you did so right, and what did you learn from that?
I think it it serves two purposes. So the first was to show that there was a lot of slack in compressing models. And so from a scientific perspective, it was a good finding and and a good learning from from from the community. It also filled the gap in the efficiency to performance to the space of models where there was definitely something missing. Seven b is the size that allows to run efficiently a model on your MacBook or on your smartphone. And we made it sufficiently smart so that it was still useful.
So there was already 7B models before, but they weren’t good enough to do interesting applications. And so by by targeting this specific space, we talk to the developers immediately because developers like, the casual developers running on a gaming GPU or on its MacBook. So it it created a lot of curiosity and adoption because it was a missing spot in the in the performance to efficiency space.
When you look at, like, lessons from that and how it impacts future releases, any that really stand out for you?
I guess it taught us that there was a lot of interest for efficiency other than scale, and so that’s why we continued of targeting very efficient models with the Mixtral 8x7B and and more recently, Mixtral 8x22B, ensuring that for a certain cost and for a certain size, we were reaching the top performance of the market. That has been our major motivation for us to target efficiency while simultaneously scaling to other larger and larger model.
I spoke to Sarah before the show, and she said the cool question that I think is, you know, with the focus on efficiency and the efficiency frontier, does scale matter?
Scale matters in the sense that if you spend more training compute, you can make the models more compressed. So you do need to have some compute to compress models. No. Scale isn’t the only recipe the only ingredients to the recipe you need to scale, but you also need to have proper data. Otherwise, you reach some data quality limit. You need to have proper techniques for training. I mean, people call it compute multiplier, I guess. How do you actually make some efficiency gain that are not costing you compute because computing is expensive.
And so one of the thing that we do at Mistral is to try and harvest this compute multiplayer.
Can I ask, in that chasm of efficiency gain without costing more compute, is there much more efficiency we can eke out? Like, is there a lot for us to eke out, or are we working really at marginal improvements already?
I think it’s an open question. I believe there is. I believe we can make models that are much better for a certain size, but it’s it’s as open a question as can you find can you make a much better model on the same kind of data by making it bigger and training it for longer? Things you need to discover them also on the way. You can try and predict the kind of performance you reach out you you will achieve. At the end of the day, you need to try it out.
So that’s the I mean, it’s really much a research field. You need to do the research, and you need to try things.
So I asked Sam Altman this question. What is the end state for the model landscape? Most people say, ah, it’ll become commoditized. And, actually, there will be 12 players, and it’ll be a race to the bottom. What is the end state for models in your mind, and how do you think about the commoditization question?
I think the end state is to have more more features on developer platforms that allow us to do customization, that allow us to make low latency models that serve a certain purpose, that allow us to evaluate them and to improve them over time. And so the model is only, like, a tiny part. Mean, I it’s a central part, but it remains a tiny part of an application. What you want to do across time and when you deploy an application that you expose to users, you want to ensure that it works, ensure that it’s that it’s latency reduces over time, ensure that its quality increases over time.
And so I think that’s the the end state is models are effectively going to be a starting point for any AI application developer. They need to be surrounded by tools, by life cycle management platform, basically, and that’s the one thing that we started to build. Like, general purpose models are a bit undifferentiated, but the differentiation that you need to create for for your application comes from the data you put into it, the user feedback that you gather, and the intelligence that you have to figure out what the application should be doing.
And that is not commoditized at all. There’s no recipe that allows to go from a a general purpose model to model that is super good and better than all of the others at your specific task. This is a missing piece in the puzzle, and that’s one of the aspects where we’re putting our strength on the product side.
Sam and and Brad said the other day that models just aren’t actually that good any, like, yet, and they need to improve a lot in quality. What are the largest constraints or bottlenecks on model quality today, and what needs to change for them to improve?
I think the data quality is is a constraint. How do you leverage the entire world knowledge and ensure that the model follows a certain path toward learning more and more complex things? That’s a very important part, and I think it has been a neglected part. There’s obviously compute, but given the amount of data you have we have at hand, compute is already running into is no longer the bottleneck. The bottleneck is more the data at that point. You should look at text to text models. And so the question is, how do you refine the data and how do you feed very high quality data to the model itself in order to improve it over time?
And I think in in that setting, it becomes a bit one bottleneck that is associated to bringing better model performance is the question of how do you evaluate these performances. You need to have very good evaluation that targets very specific topics. Like, you want the model to be good at helping diagnosis in in in hospital but in French. And oftentimes, you’re a bit out of domain compared to the data you have. And that’s where you you should identify a gap and you should try and fill it out.
The pushing the model capabilities become also a question of mapping where they’re failing and figuring out ways of improving it. For instance, they are failing at mathematics. How do you improve their mathematic thinking? How do you improve the way they demonstrate theorems? The answer to this is very different from the way you answer the question to how do you improve medical diagnosis in French, for instance.
Will we see large scale generalized models that are able to answer huge waves of very complex problems? What do you think we’ll see much more vertically specific, smaller, more specific models that are much more vertically aligned?
Yeah. We we believe that. And, actually, these vertical models are not going to be out there. They’re going to be built by the application makers. Because the only way you can make a low latency model that is super good at a specific task is to get rid of the general purpose aspect. Because a general purpose model is a bit bloated. You can think about everything. But if you you want your model to think thoroughly about a specific topic so that you can call it in your AI application while maintaining a good user experience with low latency.
What role do you play in that world? If it’s actually in the application layer where you have that specific model creation, where that kind of value accrues, where do you play in that?
It’s a very hard job to make a specialized model. So it’s actually very tied to the way you create a pre trained model. And so bringing the tools that allow to do it in a foolproof way, so allowing developers to create customized model that are performing very well at their task, but that doesn’t require expert AI knowledge, which is hard to find, is definitely something where we’re insisting.
So I’m in investors’ day, and I’m I’m pleased that you just said that there will be value accrued at the application layer. Because I look and I worry that, bluntly, everything is going to get steamrolled by some of the players that we mentioned. How do you answer the question of will value accrue at the application layer? And for me as an investor, say, Arthur, you know me, how would you advise me?
There’s two opposing directions. So the first is that the models are getting better and better. So it means that creating a verticalized application, as long as you have the data for it and a good understanding of the use case you’re facing, is going to be easier and easier if you have access to the tools that facilitate it. So that’s the first aspect, which would make me think that the application layer is going to grow thinner and thinner. But then there’s also the fact that the models are are getting cheaper and cheaper because we managed to compress them because we make a lot of improvement on their efficiency.
And so that means that effectively, this plus the competitive pressure there is on the model layer means that the price around the model, the dollar per intelligence unit, let’s say, is definitely going to reduce. So there’s these two aspects of growing ability, compressed price, which on one side says that the application layer is going to grow thin, and on the other side says that the model part is going to grow thin. So for us, the the approach that we are taking is that the model part is still going to be big enough and that we need to build this platform on top of that, because that’s where we are going to enable all of the vertical applications that will be interesting for humanity.
How do you think about that positioning and brand? Because there are other players who are much more direct in saying, hey. We’re we’re gonna dominate a lot of different verticals and kind of be afraid. How do you think about that enabler to vertical applications or not in that positioning?
We are not a verticalized company. We started Mistral to bring value to developers and to bring freedom to developers. So when we started, there was basically one API out there, soon too. And the field of generative AI was starting to look like it would be very centralized around a couple of players. And we took this platform approach where the model that we’re making and the technology that we are making, we are allowing developers to own it, to modify it. And so bringing freedom to developers and AI application makers is, I think, the best way in distributing generative AI as widely as possible, which is our objective as a company.
Making AI ubiquitous, bringing frontier AI into everyone’s head is the reason why we started. We did a good job at it, but this open source part was, I believe, a good enabler for the community and made people realize that they could build very interesting technology by modifying the models themselves instead of depending on the APIs of a couple of providers.
Dude, what do AI developers care about? Everyone kind of gets on Twitter and goes, oh, did you see x’s performance this week is better than y’s performance last week? What do they care about? Efficiency, scale, cost? What drives that usage in decision making?
They care about cost for sure. They care about customization, being able to modify the models at will. And on that aspect, I think we are only scratching the surface of what can be done, Like, the fine tuning aspect that has been, like, the go to solution is probably a little too low level from what we should be doing. They care about being able to deploy anywhere. So they operate in a certain space, in a certain cloud. They might be operating on prem. They might have some edge devices to deploy to, and they want to be able to put their technology there.
And so they also care about portability, which in turn offer data control. Usually, LLMs, AI becomes very useful when you connect it to knowledge bases or to anything that is related to certain business. In that respect, it becomes a very sensitive part of your application because it sees everything. It sees all of the data you have. And so enterprises, for instance, do care about ensuring that the proprietary data they have is accessed in something that they can they can completely secure. And that’s the reason why we deployed our platform on Azure and AWS, for instance.
That is bringing the security layout that they need.
We’re gonna get to enterprise. Can I just ask you, does brand matter in this segment? When we think about building brand, both in terms of developer adoption brand, corporate brand, is brand a large determinant of adoption in this segment?
Brand seems to be critical, and this is something that we have learned on the way. People use certain models because they are known to be good. You can’t afford to evaluate everything out there. And so having some form of community vouching is super important. The approach we took with APHT distributed models has contributed to what I think has become well, at least a known brand, and we believe that it’s definitely going to be important. Brand is important because trust is important in that domain. And open source brings trust in terms that provide some trusted brand.
You mentioned the word open source. I’m I’m gonna get to that. I do just wanna touch on the you mentioned cost also. I wanna touch on cost. How, when, and who will make marginal revenue that exceeds marginal cost in LLM based products?
You should be telling me you you I mean, you are you’re the investor. Right? You have your own company.
Yeah. I’ll I’ll that means I know nothing. Okay? And why
tell you who is doing the most the most margin at the moment so that that it’s it’s probably going to evolve over time.
Who’s doing the most margin at the moment?
NVIDIA is at that point. The cloud providers are pretty much at cost. LLM providers, we are not at cost, hopefully, but the margin that are known to be lower than the typical software margins. AI application makers, some of them, the one that almost used, seems to be doing a pretty good margin. I think it’s going to be quite a moving space. As I’ve said, the the capacity of models makes the the cost of making an application lower and lower. I don’t think there’s any way in which the marginal cost and the margin of the most important part of of that technology, which is really the foundational layer, becomes zero because otherwise, there’s definitely going to be a a fairness problem.
What do you mean by the fairness problem? Talk to me about that.
Usually, the value tends to accrue where most of the difficult part is and most of the defensibility is. It was for a while, it has been on foundational models. I think it’s obviously evolving with time, and there’s no mode that isn’t disappearing or evolving with time. That will remain the path where most of the innovation will will be made and where most of the well, at least a significant part of of the accrued value will well, the value will accrue.
Is there actually much of a barrier to creating a foundational model company today? I know that’s a really broad, question in many respects, but you have so many different players now and new ones popping up every day. Is the barrier just reducing day by day?
I don’t think it is. To be relevant in that space is a very hard topic. You need to be dominating on the cost efficient on the efficiency of performance by auto front, and there isn’t there’s only a few companies that that are currently well positioned. So you can try and do something, but but but if it’s not relevant, if it’s strictly dominated by another model or another technology, then you have a problem. There’s a few barriers that are pretty hard to to face that you need to to raise sufficient capital to have enough compute and be relevant.
You need to have people that knows how to train models, which is still the scarce resource. And then you need to have a good brand because as you’ve said, it’s highly competitive, and this is not something that comes out of thin air. So I think there’s still a lot of defensibility on the market. Although there is a lot of noise, which is different.
How quickly does the cost of compute go down, do you think? Because if you look at those things, actually, said cost of compute, access to talent, and brand. If we drastically bring down cost of compute like many think we will very quickly, you’ve got access to talent and brand. Two of those are more doable.
The cost of compute reduces over time just based on hardware cost. It reduces around 30% every two years if you show a low Nvidia roadmap. The other thing that increases the efficiency of algorithms. So if you look at the way we train models from three years three years ago and the way we train model today, I think we have probably made something around a 100 times algorithmic improvement. That’s probably where most of the gains were actually made in the last three years. Obviously, the cost of of compute does reduce, but it doesn’t produce faster than the mobile load.
So our bet is more on efficiency where I think there’s a lot of of improvement that can still be made.
Given Nvidia’s prominence there and Nvidia being the one where the gains are, as you mentioned, is bluntly one of the single most important things, not simply the quality of your relationship with the core provider being Azure or being Nvidia or being one of these players. Is that not the core determinant of success today?
I guess it’s an important aspect. There is a strategic dependency from the AI layer on the cloud providers and on Nvidia. The competition is heating up as well, but it’s it’s effectively important. It’s effectively useful when you develop a software to also know the hardware provider because they can help you out in optimizing for the hardware. It’s useful when you’re selling your developer platform to enterprises to bring that platform through their usual provider, which happens to be a cloud provider. So there’s definitely some important collaboration to be made there.
When might Amazon invest, like, $2,000,000,000 in Anthropic or whatever it was? Is that not just, like, a trade where, like, Anthropic then spend $1,800,000,000 on Amazon and return it back to them. Do you see what I mean? Is it is it not a bit of a misnomer?
It looks like you’re round tripping. Yes. I don’t know about that deal particularly, but it’s it makes sense from from both perspectives.
Can I ask how does the unlimited availability of open source LLMs impact the answer to the above being marginal cost and marginal revenue? Does it change much?
It moves the value a little higher than the model itself. It moved the value to the platform and customization part, which is really, I guess, something that we’re expecting, and it’s it accelerates that process.
You started off complete, like, very open source, very much open to the community. Now you have small models open and then larger ones closed. Am I right?
We also have large models that are open now. Depends on the threshold for small and large, but the eight times 22 b is actually relatively large by any standard.
What was behind the decision then to close some models? Is it just a business case where you need to make money?
Opportunistically, there was a opportunity to grow the business using that asset as something that we were selling. It’s still the case that we’re growing that our business on top of commercial with it in particular. There’s also a good way of cementing some strategic relationships with cloud providers, and it’s going to be to continue to be the case. Like, we still intend to to be a leader in the open source part and to have some unique assets that we can license and to have some unique platform that developers can use.
It’s hard when you suddenly have, you know, some closed and you start building an enterprise team. For you as a founder now, how do you think about that balance between a research team and a sales team and making sure that the two cultures come together well?
I think some important thing is to create empathy so ensure that the science team also understand the problems that the users are facing. It improves the science because at the end of the day, the general purpose technology we are making is only general purpose if you identify the use cases. So that comes back to the earlier discussion we had. So ensuring that the science team has some relatively direct exposure to the product and to the business team is actually important to make them understand what where the model is failing and how it could be improved significantly.
And on the other side, the go to market team has to understand it. It’s a very technical sales sales motion because you you’re selling not the product, but you’re selling something that is going to pull out the product. So you need to tell the customer how these things should be used to actually make something that brings value to the business, and that only goes through strong enablement of the go to market team. So it’s I think it’s a challenge. They don’t operate on the same scale.
The science team has cycles of several months. The go to market team goes faster, shorter cycles, let’s say. But I think so far, we’ve managed to recruit go to market people that have some technical interest and and technical people that have some business interest, and I think that’s how that’s how you ensure that you don’t have silos at the end of the day.
One of my worries with, bluntly, this space as we move into enterprise is that brand matters so much in terms of enterprise actually, and they already have existing agreements with Microsoft. And I worry that actually product or model quality doesn’t matter as much as distribution. Microsoft just tack on existing clients with new products. How do you think about that as a core challenge? And am I wrong to be worried about it?
So I think it’s true. Distribution is very important. A shortcut to distribution is to create demand through open source models.
Do you think open source is ready for enterprise? Or do you think enterprises are ready for open source, and do they care about it enough?
It depends on the enterprises, but some have been early adopters and are using a lot of these models into in production. So for sure, they’re ready enough in order to bring them to the next level of putting things into large scale production, etcetera. I think they’re still lacking some product around, like, managing correctly, load balancing, customizing the models. Because you can do it with DIY solutions, but if you want to make it robust enough and scalable enough, it’s actually not easy. And if you want to actually increase the quality of the models, the custom models, the recipe are are a bit hard to to set.
So the most technical savvy enterprises are definitely ready for it, and there’s a few there’s actually many use cases that are in production using using open source models. Now in order to widen the adoption, there’s definitely some tooling to be brought to the market.
Obviously, every enterprise today is sitting in a boardroom going, what’s our AI strategy? What do you advise them and what questions should they be asking?
Start thinking about how they are going to change all of their products using AI as a premise, using the existence of, like, very clever agents because you can build very clever agents today. Assuming that presence and working backward to understand the consequence in term of organization. Not thinking about generative AI as a way to as a way of increasing productivity in a word processing, but rather as a way to change completely the way you operate your core business, which usually involve taking models and customizing them pretty heavily to create the differentiation that you will need in, like, five years’ time when everybody will have adopted the technology in its core business.
So my question to you is you’re in France. I’m in London. We both know that European enterprises do not move very fast. Most do not even have Slack today. My concern is that we drastically overestimate adoption in the near future and maybe underestimate it in the ten year, twenty year future. Do you think that’s the case here? And do you worry about the lethargy of a lot of enterprises, especially in Europe, in adoption?
I mean, it’s a general phenomenon in in tech that you always overestimate the speed but underestimate the impact. I think it’s probably occurring today. It’s slightly different in the sense that there’s some executive support for pushing generative AI solutions even in Europe. So there’s some delay compared to The US market for sure. I wouldn’t say it’s it’s very significant. It’s one year maximum in term of delay. The challenge here is that it’s a technology that is that can take many forms. And so trying to focus on some specific thing that you can bring to the market that have AI in it is a prioritization challenge, and so you need to be very strategic around that.
I don’t think this is super easy for enterprises generally. It will become easier once they try out off the shelf solutions a bit more. Once they realize that there are some developer platforms that allows to do it without hiring very expensive and hard to find AI scientists in house. And so we expect that this is going to accelerate in the in the coming years.
Do you think that we’re still just playing in the experimental budget game, or do you think that we’re moving into core budgets as well?
It depends. It’s moving into core budget for customer support, for instance, where like, areas where the application of AI is pretty obvious. It’s definitely moving into core budgets. It’s also at the experimental stage in many other functions. And for core applications in the industry, the telecom industry and in health care, this is still in the playground, but I I think it’s going to evolve in the next year.
Can I ask, as you build out enterprise, it’s another expensive thing to build out on top of compute and talent? It costs real money. And I spoke to Paul at Lightspeed before, and he was mentioning to me bluntly how much less capital you’ve raised compared to a lot of your competitors, most obviously OpenAI and Anthropic. He said in a world where capital equals compute equals quality of model, how does Mistral keep up and stay relevant?
So the good thing is that capital is correlated to compute, but then compute is correlated with quality. It’s not completely dependent on it. And as I’ve said, there’s some strong opportunity for providing models that are the best of their class. They might be sufficient to actually solve certain use cases. That’s where we’re playing. In addition to be playing on the scaling part, because obviously, you do need to keep stay relevant. You need to to keep your technical team motivated. And to keep your technical team motivated, you need to give them the experimental bed they need to make new discoveries and to to progress science.
And that that is where you need compute in addition to growing the model across time. I mean, we’re growing our compute like every companies, we are convinced that we don’t need to grow at the same rate because there’s a lot of barriers that are not compute related that are appearing on the way and that we’re already seeing. So we think we we can we can scale, and we are also convinced that we on the efficiency front, we are already very well positioned, and we’re strengthening that position.
What are the biggest barriers to Mistral today?
We’ve had a few delays with our compute providers for sure. That has been a barrier. So the last answer to your question is to be taken with a grain of salt. We are still bottlenecked by compute for sure, but that’s because we don’t have many of it. We have 1.5 k 800, which is a few percent, I think, of the capacity of our competitors. And so that’s definitely a bottleneck that is going to improve significantly in the coming months. Can I
ask finally, was it a mistake for you to not scale that quicker with the benefit of hindsight now, which you wish you’d scale that quicker?
You can’t really scale that much quicker because you you can’t raise, like, $22,000,000,000, like, on on the seed round. I mean, at least you couldn’t in 2023, but you can today. But you can only hire that fast. You can only scale your infrastructure to manage more GPUs that fast, and you can only raise capital that fast. So there’s some acceleration constraints that are pretty hard to fight, and that are pretty much the the first principles of of starting a business.
You mentioned about the scaling constraints in in in cash. Does it matter where your cash comes from? Does it matter if you have European funded, Saudi funded, US funded? Did you think that matters?
I guess governance matters. The way it is important for a young company like us is to be in under con under the control of of the funders because there’s a lot of things to be invented, and and vision can only be carried by them. We have very good governance down, a very simple and clean governance that makes us a for profit company, growing a business to actually push the science frontier. This is something that we’re very attached to, being able to control the company, leverage our funding partners appropriately to grow in different parts of the world.
In The US, in The EU, it has been critical as well. So it does matter in the sense that we want to have partners that are supportive and long term because we are in the field that is fast moving where we don’t know yet exactly where the value will accrue. And so being flexible and being smart is definitely a requirement when you raise money.
Would you take money from Saudi or China?
Good question. It’s it depends depends on the term. China is a bit hard. For us, it’s even hard to operate in China. I mean, don’t operate in China because you can’t really operate in US and China without being, like, a very, very large corporation, And so you need to make some choices.
What chance do you think that Europe has in AI? I know it sounds deterministic and defeatist, and so you might be going, oh, fuck, Harry. Shut up. But it’s like, what chance do you think Europe has in AI, and what does it take for us to stand up as a serious AI industry with Europe?
I guess the chance it has is that it’s a revolution. It’s changing the way we do software. And so as every revolution, it opens a lot of opportunity for for new actors, And there’s no reason why there shouldn’t be an actor that is that was created in Europe that could grow pretty fast. And that’s the mission that we we gave ourselves. We have the talent. Capital can cross oceans without without too much problem. We have the market. The market is more fragmented than in The US for sure.
The ecosystem, the digital native ecosystem is definitely smaller, but it exists and it’s growing. There’s local opportunity for business development. On the talent side, we can hire 23, 24 years old people that we can onboard in four months, months, and they operate as well as any software engineer in the valley. So people are quite talented here. And so we if we manage to keep them and to to convince them not to go to The US, we have a lot of opportunities.
When we look at computers, mobile, cloud, the kind of core technology shifts, the way that it’s worked is Europe has kind of ceded control to The US and then just taxed US companies for access to our citizens if one’s being defeatist. Is it different now?
I mean, Europe is paying the the the price of not setting up a VC system in the sixties, but setting it up, like, forty years later or even fifty years later. And so
the And way, the dirty secret is that the VC ecosystem in Europe is US funded.
Yeah. It was. I think it is it still On
honestly, in large part, yes. There’s government institutions which are backfilling it, but largely backfilling it with bad players who aren’t very good. But the best providers in Europe, largely US funded by top US institutions.
Okay. I think, yeah, as I’ve said, it takes time for an ecosystem to build. So you have layers of entrepreneurs and investors that stack on top of each other. The US has sixty or seventy years of venture capital investments. I think Europe has only twenty years. I mean, it takes time. It takes an incompressible time to build an ecosystem. It takes also some willpower. And I think now we are seeing that willpower. We are seeing entrepreneurs creating companies. We are saying, this is like you, not going to The US.
Everything is positive. It just takes time. I’m adamant that we’ll manage to do something interesting.
On the engineering side, do you feel like you have the depths of talent pool to hire from as you scale on that? You do.
On the engineering side, on the AI side, do. We have a team in The US, though, which is working on, like, specific topics. Like, for senior AI scientists, you find them more in The Valley than than in France. For junior AI scientists, there’s a wealth of talent in France, in Poland, in The UK. I think one of the strength of the area.
When you were raising money, was it very different speaking to European investors versus US investors?
I guess in the seed round, no. It wasn’t that different because it was a seed round. For the series a, which was a bigger round, it was European funds were structured to do the kind of deal that we were proposing. We didn’t even have a lot of conversation because they they just couldn’t get their head around the investment that needed to be made, whereas we were approved in your company. Yeah. I think what is lacking and it’s related to to the ecosystem parts in Europe are are growth funds that are able to take huge bets with lots of conviction.
And that in turn should improve over time, especially if if we manage to to use a European wealth and channel it more into their growth funds than it is today.
I think you have more hope than me on that one. That is not gonna happen. We are not gonna see many more European growth funds be built in the next few years for sure. Not in the next three to five.
Yeah. It it hinges on a few political decisions.
And I think it hinges on supply of capital and belief in a future European ecosystem that can contend with other large ecosystems.
It’s a chicken and egg problem. This could be nudged into the right direction if if politics wants to do it. If a couple of companies show that you can actually have companies that grow fast in Europe, and that’s what we are trying to do. I’m not too pessimistic. I mean, I I find you too pessimistic. You should come to France. I think that you will get more optimistic.
Do you know what? If a Parisian is telling me that I’m too pessimistic, then shit. I really need to be more optimistic. My my question to you is, like, you know, you just mentioned that the speed of scaling. Hardest thing, dude, is scaling with your company at the same speed. What was the hardest thing about yourself scaling as CEO with such speed of scaling of the company?
I mean, we are learning on the job. It’s effectively, you have organizational challenges. How do you how do you ensure that 45 people communicate well together? How do you manage your time in terms of representation time, in terms of business development time? Because we’re still at the stage of the company where we get involved a lot in the in the deal making aspect. And how do you, yeah, ensure that you set proper directions and maintain the team in the state of tranquility despite the amount of noise that there is on the on the competitive side.
The fact that the direction is obviously going to be changing over time because there’s a lot of uncertainty in in in that field. So this is, I think, the hard part. I don’t I don’t think I’m doing it properly, but we are actively trying to find sources of information to learn new things, let’s say.
If you could cool yourself up to, like, the night before you became CEO and founded Mistral and give yourself some advice with the now knowledge that you have, what would you say to yourself, Arthur?
Maybe stage a bit more of the product development and go to market development. We did start the go to market motion at the time where we had absolutely nothing to sell. It did work out. It did create some brand awareness despite the absence of anything. I think it might have been slightly simpler to save things maybe a little more, developing the product a little before developing the go to market. But since it’s such a fast moving field that we did start everything a bit together with some organization that was a bit lacking, and now we are we are solidifying it on the fly.
It has worked out. It hasn’t been optimal, for sure. And so in hindsight, you can always give you I I could give me, like, a few tactical advices on who to hire when. Generally, I think the strategy we had one year ago hasn’t changed much. We did realize that we would need more capital and that we could not operate only from Europe and that we needed to go to The US very quickly. Those were findings that we did on on the way. I don’t think that they would have helped it would have helped that much to know it a year ago.
One, do you feel like you have enough cash now?
I guess the startup is always fundraising. It’s it’s a field where for the years to come, the investment are going to exceed the revenue by design because you do need to scale and you do need to stay relevant as a company. So effectively, there need to be some investment. The revenue is ramping up, so there will be some some revenue to reinvest. But today and for the years to come, the speed for developing research should be faster than the speed at which you can develop your your go to market.
Before we do a quick fire, when you look at the landscape today, which competitors did you do you most respect and admire?
I mean, they all delivered. We were surprised by Cohere recently, came up with new good models, and I think that was a surprise for us. And, obviously, OpenAI and Anthropic and my friends at Google are also doing a good job. So, it’s a it’s a competitive landscape, and we respect all of them. We also all work in the same direction, and, eventually, we’ve, the same, higher goals. So, it’s great to have respect for one another.
Is it too late to start one now? We see, like, holistic starting now. Is it too late?
Holistic. I know them well. Is it too late? I I wouldn’t recommend going into the foundational layout business. I know Sam didn’t recommend to do that one year ago, and we did, and it seemed to have so far changed a few things. So I I think it would be arrogant for me to say that there’s no chance for a new competitor to arise and beat us for sure.
Listen, my friend. I wanna move into a final thing, which is just a quick fire. So I say a short statement. You give me your immediate thoughts. Does that sound okay?
Yeah. Let’s do that.
So what worries you most in the world today?
Global warming. There’s the race of, the planet heating up and us finding solutions for it. I think AI is part of this of the solution. It brings more control. It brings potentially more efficiency in some of the processes, but there’s effectively a race for survival. So I think this is something that we should be a bit more aware of.
What have you changed your mind on most in the last twelve months?
I think I’ve changed my mind on a lot of management premises that I had and that I had never tested in real.
What was the biggest one?
Transparent feedback is actually super useful for our company, and so operating in a almost fully transparent manner has helped out growing at without breaking.
What element has been the most unexpectedly challenging in the scaling of Mistral?
The amount of demand that we had to manage, which is too high for what we can handle. The brand success, the fact that people knows us was a bit unexpected. We knew that it would be noticed. We had no idea that people would start using us that fast.
What do you do to calm down? You have a lot going on now, Arthur, and you have a lot of expectation and cash on your shoulders. What do you do to just
I run. I cycle. I think my partner will yell at me, but I I try to take care of my daughter.
Okay. You’ve recently become a father. What do you know now that you wish you’d known when you first had your daughter, you know, very recently?
I had no idea that you needed so much energy to to care for for small children for a small child, let’s say.
Where do you think AI will take the world in the next ten years? Like, what does the future of society look like in a world where AI is embedded into everything?
Well, it’s changing the way people work significantly in the sense that it requires to be more creative and to bring more value beyond what can be automated. So it’s a very structural change on the job market, which means that there should be some adaptation that are taken pretty quickly in training, in education, that people can get a sense of what is going to be expected from them in their daily job, assuming that there’s some AI out there.
Do you think the fears of job replacement are grossly overexaggerated?
I think they are. I mean, it depends on who you’re speaking to. I think job are going to be displaced for sure. Some will be replaced. Some will open up. We’re just trying to move humanity to a higher level of of abstraction. So we can now talk to machines, and machines can understand and answer in a human like fashion. This is not so much of a badging change compared to what we were doing with computers. I think what’s happening right now is that probably the speed in our elevation toward higher abstraction level is probably occurring at an unmatched rate in history.
Though that means that the society adaptation is going to be more challenging and needs to be anticipated.
Final one for you. We do a show in twenty thirty four, ten years’ time. If everything goes right, where’s Mistral then?
Mistral has some very relevant models, commercial and open source, and it has a very strong developer platform that allows to do everything that you need to create your AI application. That would be a good achievement.
Arthur, listen. I’ve so enjoyed doing this. Thank you for putting up with me going in many different fast moving directions. You’ve been incredibly patient and a brilliant guest, so thank you so much, my friend.
Thank you for hosting me.
What a fantastic guest to have on the show. I wanna say a huge thank you to Arthur for being so patient with me there and for being so open with some of those answers. If you’d like to see more, you can of course find it on YouTube by searching for 20 v c. But before we leave you today,
· Sponsor read0 min · 481 words
we’re all trying to grow our businesses here. So let’s be real for a second. We all know that your website shouldn’t be this static asset. It should be a dynamic part of your strategy that really drives conversions. That’s marketing one zero one. But here’s a number for you. 54% of leaders say web updates take too long. That’s over half of you listening right now. And that’s where Webflow comes in. Their visual first platform allows you to build, launch, and manage web experiences fast. That means you can set ambitious marketing goals and your site can rise to the challenge.
Plus, Web allows your marketing team to scale without relying on engineering, freeing your dev team to focus on more fulfilling work. Learn why teams like Dropbox, IDEO, and Orangetheory trust Web Flow to achieve their most ambitious goals today at webflow.com. And speaking of incredible products that allows your team to do more, we need to talk about Secure Frame. Secure Frame provides incredible levels of trust to your customers through automation. Secure Frame empowers businesses to build trust with customers by simplifying information security and compliance through AI and automation.
Thousands of fast growing businesses including Nasdaq, AngelList, Doodle, and Coda Trust Secure Frame, to expedite their compliance journey for global security and privacy standards such as SOC two, ISO 2,701, HIPAA, GDPR, and more. Backed by top tier investors and corporations such as Google, Kleiner Perkins, The company is among the Forbes list of top a 100 startup employers for 2023 and Business Insider’s list of the 34 most promising AI startups of 2023. Learn more today at secureframe.com. It really is a must. And finally, a company is nothing without its people, and that’s why you need remote.com.
Remote is the best choice for companies expanding their global footprint where they don’t already have legal entities. So you can effortlessly hire, manage, and pay employees from around the world, all from one easy to use self serve platform. Plus, you can streamline global employee management and cut HR costs with Remote’s free HRIS. And hey, even if you are not looking for full time employees, Remote has you covered with contractor management, ensuring compliant contracts, and on time payments for global contractors. There’s a reason companies like GitLab and DoorDash trust remote to handle their employees worldwide.
Go to remote.com now to get started, and use the promo code 20 v c to get 20% off during your first year. Remote opportunity is wherever you are. I so hope you enjoyed that show. As always, it means the world to me that you listen. You can check it out again on YouTube by searching for 20 v c, and stay tuned for an incredible episode this coming Wednesday with an OG of the venture space, the one and only Marc Souster at Upfront Ventures.