Go back

How Mistral Is Building Frontier AI for the Enterprise | NVIDIA AI Podcast Ep. 301

21m 33s

How Mistral Is Building Frontier AI for the Enterprise | NVIDIA AI Podcast Ep. 301

In this podcast, Tim Laquois, co-founder and CTO of Mistral AI, discusses the company’s philosophy on open models, its collaboration with Nvidia, and its enterprise-focused platform. Mistral, founded 2.5 years ago and now with over 700 employees, releases open-source models to reduce duplicated pretraining efforts and foster global innovation. The company tailors models for specific domains—such as languages, private codebases, and manufacturing—to make them faster and cheaper for enterprise use. Mistral Forge, their platform, provides training frameworks, data pipelines, and customization tools, enabling customers to build on reusable use cases. Laquois highlights the Neematron Coalition with Nvidia, which aims to create a new open-source frontier model, leveraging Nvidia’s large-scale infrastructure. Mistral benefits from Nvidia’s GB200 and GB300 hardware, achieving 2.5x speed improvements for mixture-of-experts models, and uses NVFP4 precision for inference efficiency. Key challenges include managing inference at scale and developing a robust permission system for AI agents. Looking ahead, Mistral focuses on agentic AI, platform usability, and partner enablement through 2026. Laquois emphasizes that open models allow enterprises to control, customize, and trust their AI, even if slightly behind frontier capabilities, balancing innovation with security and governance needs.

Transcription

3349 Words, 18481 Characters

English
The benefits for everyone involved really is that we will have a new open source frontier model that everyone can build off on. Welcome to the Nvidia AI podcast. I'm Noah Kravitz. My guest today is Tim Laquois, Tim is co-founder and CTO of Mistral AI. We're here to talk about Mistral's philosophy on open models. Their collaboration with Nvidia and the Neematron Coalition and their new framework for Tim. Thank you so much for taking the time to join us. Thank you for having me. So maybe we can start with each of the audience a little bit about Mistral and about your role there from the beginning as co-founder and of course the CTO. Sure. So we started Mistral about two and a half years ago with a human author. At the start we were all three of us fresh of our researchers role in Big Tech and what we knew how to build were models. And so that was what we started with and we showed the world that we knew how to train models that we could be efficient with the infrastructure and that we could deliver high quality models that we decided to release open source. That was our claim to fame but our goal as an enterprise company was to provide the value of these models to enterprise. And so quickly we realized that just chucking weights over the wall wouldn't achieve that. We went ahead and built a service part of the company that would go with our customers and help them realize value with those models. We also started building a platform to inference those models and enable our customers to really actually use them. That platform has grown a lot over the years with the industry really with the rise of connections and the need for more context with added mcp connections with added a lot of nice days to the platform to an authentication and things like this with the rise of agent AI. We're also adding a lot of hosting capabilities so our customers will require to be easily to easily deployed things like sandboxes for their mcp microservices they might also need some sort of hosting and auto scaling there. And so we're really building up that those platform capabilities in a way that stays something that we can deploy on prem for the customer and where they have full control. And so through this we've also seen the need to extend at the lower layer into infrastructure. And so June of last year we also announced missile compute which is our initiative where we're building our own infrastructure and standing up around data centers. And we've started to train on them and it's an infrastructure that we can also ship for customers using how long has the company been around now. Two years and a half. A lot packed into those two years and it's been quite a journey. We started with three of us and today we're north of 700 employees. So alongside all of this we've also grown and managed to come from right and interesting. What is it about open models that really accelerates global innovations so much. So especially around pre training. I think there is a lot of wasted resources because everyone is taking the same raw data which is the data that's available on the web and doing their best to compress it and to a fixed amount of weights. And so there is certainly a lot of know how into how to select that data, how to curate it, how to do the right optimization procedure, how to train models at scale. But essentially everyone is doing the same thing and ending up with pretty much the same artifacts. And so one of the main frustration with you and when we're working together at meta was that the entire world of research couldn't really benefit from this and couldn't really work on top of all of this effort. Because models weren't open source and not no other academic lab had the resources to create something like this. And so by creating models and releasing them as open weights. We're still free to work on the licensing and free to provide a software platform and services around it and we can still build a business. But we also enabled the entire community to really create around those models and it's one of the most exciting things I've seen is how quickly the open source around the open source community around open weight model is grown and the crazy things that has built both in terms of the creativity of it, but also the quality of the infrastructure that it now provides. I want to ask you about the Neematron coalition. It's something that Mistral has joined and I'm wondering what your perspective is on it. What you think it's going to bring to bear both for Mistral, but you know, for the coalition in the industry as well. Yeah. So for us working with Nvidia isn't really a new thing and we've already trained a model in the past, which was called Missile Nemo 12 B and that was our first experience training models together. And so our teams know how to collaborate and what this will provide for us is Nvidia's expertise in larger scale infrastructure because Nvidia has more resources and more experience and running those large data centers at scale. We have a lot of expertise in the various ways to pre-trained models. We have expertise in multi-modality training as well. And so as we speak our teams have been running experiments and exchanging on what model it is that we want to train together and release to the community. The benefit for everyone involved really is that we will have a new open source frontier model that everyone can build up on. When it comes to customizing or as Mistral says, tailoring AI models, can you speak a little bit about, I mean the importance of it, I think as we're talking about large language models and world models, but now small language models and models that are kind of tailored to really do a specific purpose or work in a specific sector industry. Talk about Mistral's philosophy and why tailoring models has always been such a core part of what you do. Sure, the reasoning behind it is quite simple. It is that when you think about agentic system or all meaning workflows, not all of the intelligence in all of the step has to be this big, very powerful thing. However, a lot of the steps are going to be repeated. When you go at scale, they're going to have to run fast. They have to run cheaply. They have to be efficient. And so once you reduce the domain of decision that you're mobilized to make or you reduce its input space and output space, then you can really reduce the size and energy that it requires. And so there are a lot of customers with whom we work to specialize models to make them faster and cheaper to run. But customization is also about extending the capabilities of models are already there. So typically if you go to different areas of the world where English might not be the main language, then it's beneficial to also address this in the training mix and maybe continue pre-trained to some model to re-release. And then you can add some model to really add to be some Southeast Asian languages to the mix to get a model that is now fluent in that language. You mentioned this kind of the bias towards English and a lot of these models and the importance of, you know, in your example, you know, you know, the Southeast Asian language of, you know, pre-training model that way, do you do a lot of work? Do you think a lot about kind of these less serviced languages and cultures and this idea of needing to preserve intelligence through models. So the kind of thing that comes through in work. This is something that we work on with clients from these areas of the world where for addressing their own customers or for their own needs, they will need to improve the capabilities of the models in those areas. And so they're usually more adept than we would be at finding a good source data in those languages. And so we help them with how to address the right data, make sure to really make the models better with through that data. So I want to ask you about forage and about how forage works together with Nvidia technologies and the Neematron family. But maybe first you can start just in case all the listeners aren't familiar. Just describe what Mr. Al forges a little bit before you get into it. So Mr. Al forges is a platform that we're releasing, which is really the distillation of all of our training capabilities in house. And so this is always a set of things, a set of capabilities. So you have the training framework. So really the part that takes inputs provides creating steps to the model and updates to provide a better model. And so, one of the. this mechanic, all of the hosting, all of the runtime check pointing and all of this is something that we provide from our own training capabilities. But it's also the tooling around it. So the data pipeline infrastructure, the evaluation infrastructure, having something that we can use with our customers that's close to what we use for our internal research is also a lot easier on us because we can validate the results. We know what to expect if there is an evaluation that is interesting to us and where we're failing for some reason, then we can transfer it. And so it's really valuable for us and it's been something that we've been using with a few companies in different industries. So typically in manufacturing, you would have companies that have a gigantic amount of specifications for what they're building and what they're engineering is doing, having a model that's fluent in that company's domain is really helpful to provide assistance. As I mentioned, we've done a lot of model customization around languages and one particularly requested use cases also around customizing to code basis that are private and have never been seen on the web. So if a company develops a very complex domain-specific language, most modern language models will struggle to provide assistance with this because they just haven't seen a lot of it. Whereas if we work with the company and deploy our solution and specifically train our coding models on that company's code base, which can be massive, then we get improved performances with the same style that the company expects with the same guidelines and influence and maybe languages that don't exist anywhere else. Right. Is there a tension that you find with enterprise customers between open source and whether it's a real tension or maybe it's just in the expectations but around things like performance and using open models and staying true to that but also staying up to date with what's better than I do, just a furiously moving industry. How do you balance that? So it depends. For some customers who are running to become an air gap environment, they don't really have much of a choice. They'll have to run the models themselves. And for this, my personal conviction is that we as a company and in particular with collaboration with NVIDIA will be able to push the frontier of open source model to be able to match what other companies are doing. And if we weren't in this for this, I wouldn't be doing this job really. And so I truly believe that we can provide models that are frontier in their capabilities and maybe we'll be six month late, but a lot of the customers that are running with us are fine with a six month delay. If that means that they completely control the models, if they can customize it, if they control its runtime, there are also many benefits to it. And so one thing that is important for us and for a lot of our clients is to know what the gap is and whether we're addressing it and whether we're, you know, we indeed are only like six month late, which is completely acceptable. And so to address this, we can potentially provide evaluations through other third party models. It's also helpful to build solutions with the latest technology when it's available. I mean, we can't stay a state of the art forever in all of the domains. And so it's completely fair and fine for our customers to go through other models and other providers to build their stack. What I want them to be confident in is the fact that there will be missile models or other open source models that will be able to provide those capabilities quite soon. And it's typically one of the benefits of the collision with Envilia. What are you seeing Tim from the customers themselves? What are your enterprise customers thinking about or what are they looking for, you know, maybe this calendar year from their AI investments? I mean, everyone's looking for value in solving use cases. When we engage with an enterprise, we often try to target an iconic use case, something that's really hard and that really provides value. This lets us dive really deep into what the enterprise and question does. And it also lets us set up a lot of infrastructure to be able to solve that use case. And this is really important for us because in doing so, not only do we get a deep understanding of this company, but we also enable a lot of further progress, once that use case is in front of completion. So typically interfacing with all of the company's connectors, setting up a system for sandbox, setting up all of the company's contacts that will need to be accessed. All of this is reusable, right? So whenever we set up with a customer, we try to develop all of our use cases in a way that compound for that customer and so that it's value that accrues for them so that the next use case is going to be easier and the use case after that even more so. And there is a lot of work that goes into this. And sometimes it's plumbing work, sometimes it's work around defining the right roles and the right access control lists. All of this ensures that after this, when people are adopting more and more AI bottoms up, they will do so safely and easily. Did Mralsey benefit in training your model using Blackwa? Yeah, I mean definitely the GB 200s, which we've been using since June 2025, I believe. We quickly saw a 2.5 X improvement, at least like out of the box, when training large-sparse mixture of experts models, especially. And yeah, no, it's definitely been a great acceleration for us on that class of models. And yeah, I think we're seeing further improvements with the GB 300 as well. Yeah, great. How is adopting NVFP for precision, impacted model efficiency throughput and also cost in your inference pipelines? Yeah, I mean, it's always great to see the, that there seems to be no limit to how much we can compress the models really. And it's a pleasure to see that Nvidia is addressing this with hardware that can natively run those operations at high speed. And so typically on infrastructure that supports it, now we run a lot of our inference in NVFP4. The challenges that we've seen with it is really around the manipulation of the attention and the longer context, where that's where things start to break down for us. But it's also part of the game to address this and make models, or quantization better at addressing those issues. Kind of along those lines, is there a big hurdle in front of your team right now that you're really focused on getting around? I mean, there are many big hurdles. I tend to focus on all of the things that we haven't been doing, and that's all of the things that we should be doing better. I think managing inference at scale is something that everyone is doing with infrastructure that's shifting, infrastructure that's new. The open source world is also going at full throttle. And so keeping up with this, maintaining our inference both stable and correct has been quite a fun challenge. Sure. The main thing that's keeping me awake and thinking is really how do we make the permission system of AI agents something that's not a headache to configure, something that's simple and natural to configure, but also robust and something that people can trust. Because in a way, it's easy enough for someone to configure an open-claw or an e-mail claw in a way that's going to work for them, and they'll likely be safe, but also as the CTO, I worry around, how do I set that up in the most efficient way in my company, in a way that's respectful of all of the data. And typically, one of the challenges is that we often think about what an agent is going to be able to read. We more rarely address where it's going to write the results. And so thinking about audiences and what restrictions we should put depending on all of the content that went into the third process and into making up the result is something that I think is not addressed while in the industry right now. And is that just a matter of just things move fast and it takes time to sort of in retrospect, kind of put these guardrails in protection zone? Yeah, I think it's something that's also-- it's one of the benefits of the open-source community, really. They create a lot of amazing things. And typically, with this claw technique, I think what we saw was how important it was to have something that has a consistent context that's always on always working for you and always available. Now, it's our role as well to understand the needs of our clients, their requirements in terms of security and governance, and understand their fears. And let them really address those fears in a way that's understood. by everyone and that lets all of the people then build on top of it in a way that's where they feel no restriction because they know everything is safe. And that's where you really see explosion in adoption in use cases which is we reviewed to see. So what's next for Mistral? specific to the Nvidia collaboration we talked about and also just on the company's roadmap for the year. Yeah so on the collaboration I mean I guess our teams already at work collaborating on figuring out what architecture we're going to train, what techniques will adopt and working on the data mixture together figuring out the scale and the model and so this is going to go on through 2026. For Mistral as a company it's going to really be about execution in the agent equal to really providing a platform that's easy to use, easy to deploy and maintain for customers and enabling a lot more people to build on top of it. So it's the platform has been getting more complete being like a real tool that we've built on and so now we want to also propagate this to the world through partners and enabling others to build on top of our technology. Excellent. Tim Lequah thank you so much for taking the time to join the podcast and of course best to look to you and everybody at Mistral. Yeah thanks so throughout me.

Podcast Summary

Key Points:

  1. Mistral AI was founded 2.5 years ago by three former Big Tech researchers, growing to over 700 employees, and focuses on open-source frontier models and enterprise AI solutions.
  2. The company emphasizes open models to reduce wasted resources from duplicated pretraining efforts, enabling global research and community innovation.
  3. Mistral is part of the Neematron Coalition with Nvidia, collaborating to train a new open-source frontier model leveraging Nvidia’s large-scale infrastructure expertise.
  4. Tailoring models for specific domains (e.g., languages, codebases, manufacturing) is central to Mistral’s philosophy, making them faster, cheaper, and more efficient for enterprise use.
  5. Mistral Forge is a platform offering training frameworks, data pipelines, evaluation tools, and customization capabilities for enterprises, with applications in manufacturing, language adaptation, and private codebases.
  6. Enterprise customers prioritize value from AI investments through deep, reusable use cases that compound over time, focusing on plumbing, security, and governance.
  7. Mistral benefits from Nvidia’s GB200 and GB300 hardware, achieving 2.5x improvement for sparse mixture-of-experts models and using NVFP4 precision for inference efficiency.
  8. Key challenges include managing inference at scale, maintaining stability with new infrastructure, and developing a simple yet robust permission system for AI agents.
  9. Mistral’s roadmap focuses on agentic AI, platform ease-of-use, deployment, and enabling partners to build on their technology through 2026.

Summary:

In this podcast, Tim Laquois, co-founder and CTO of Mistral AI, discusses the company’s philosophy on open models, its collaboration with Nvidia, and its enterprise-focused platform. 5 years ago and now with over 700 employees, releases open-source models to reduce duplicated pretraining efforts and foster global innovation. The company tailors models for specific domains—such as languages, private codebases, and manufacturing—to make them faster and cheaper for enterprise use.

Mistral Forge, their platform, provides training frameworks, data pipelines, and customization tools, enabling customers to build on reusable use cases. Laquois highlights the Neematron Coalition with Nvidia, which aims to create a new open-source frontier model, leveraging Nvidia’s large-scale infrastructure. 5x speed improvements for mixture-of-experts models, and uses NVFP4 precision for inference efficiency.

Key challenges include managing inference at scale and developing a robust permission system for AI agents. Looking ahead, Mistral focuses on agentic AI, platform usability, and partner enablement through 2026. Laquois emphasizes that open models allow enterprises to control, customize, and trust their AI, even if slightly behind frontier capabilities, balancing innovation with security and governance needs.

FAQs

Mistral releases open-weight models to allow the community to build on them, avoiding wasted resources from everyone retraining on the same data and enabling global research and innovation.

Mistral Forge is a platform that includes a training framework, data pipeline infrastructure, and evaluation tools, distilled from Mistral's internal capabilities to help customers customize models for domains like manufacturing or proprietary codebases.

Mistral believes open-source models can match frontier capabilities, possibly with a six-month delay, which is acceptable for customers who value control, customization, and runtime ownership.

It is a collaboration between Mistral and Nvidia to combine expertise in large-scale infrastructure and model training, aiming to release a new open-source frontier model for the community.

Tailoring reduces model size and cost for specific tasks, making them faster and cheaper to run, and extends capabilities like supporting non-English languages or private codebases.

Mistral saw a 2.5x improvement in training efficiency for large sparse mixture of experts models, with further improvements expected from the GB300.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.