Go back

Demis Hassabis on Gemini 3, Google Antigravity, Medical-grade AI, and more

19m 51s

Demis Hassabis on Gemini 3, Google Antigravity, Medical-grade AI, and more

The conversation revolves around the significance of Gemini 3 as Google's premier model, emphasizing improvements in performance and reliability from Gemini 2.5. There is a focus on enhancing tool calling, usability, and overall user experience, including improvements in style and persona. The dialogue also touches upon the integration of personalization, memory, and context into Gemini 3, aiming to enhance user interactions and connections with other Google platforms. Additionally, there is a glimpse into the potential future applications of Gemini 3 in fields like health, education, and research, highlighting its multimodal capabilities and the goal to create a universal assistant for everyday activities. The discussion also hints at the launch of anti-gravity, a genetic development platform, and the vision to revolutionize the IDE experience from an agent-first perspective.

Transcription

3636 Words, 20382 Characters

- Demis. - Hey, Ron. Good to see you. - Good to see you too. Thanks for your time. Today we're talking about Gemini 3, the most intelligent flagship model from Google. If you had to explain it in one sentence, why is this launch important? - It's important I think because, you know, it just continues the progression I think we've been on with Gemini over the last couple of years and we're really happy with the overall performance of this model. I think people are gonna be very pleasantly surprised by it. I think it just continues the overall performance increase across the board and you can see that from all the benchmarks, from reasoning to tool calling, reliability and creativity. I think it's better across all those measures. - If we rewind back to when Gemini 2.5 was launched to now with Gemini 3, what breakthroughs have happened since where Gemini has gone to this level with the benchmarks? - Yeah, well we've really focused quite hard on, I mean 2.5 was a great model too and we're really pleased with that and you saw how well it did in the market with developers and also in the Gemini app. But we wanted to improve things like tool calling and tool use and just sort of the reliability of that. Of course that's important for coding, which is one of the big use cases of these models, but it's also important for just general reasoning and generally how you use it. The other thing we did, I think we've done is improved the style and the persona a lot. You know, I think it's more succinct, more to the point, more helpful. And at least I, you know, and the internal testing shows people enjoy using this model even more. - So definitely a step up in coding reasoning. But for like an everyday normal tech worker that's not a developer who already uses Gemini today, what are the noticeable concrete things though so they'll be able to do tomorrow that they couldn't before? - Yeah, it depends what your use case is, but I think almost everything we've tried, like if your brainstorming ideas or your vibe coding or you're just doing some creative writing or summarizing things, you should find it's meaningfully better at all of those things, more reliable, much smarter, and I think stylistically better too. I think the tool calling and things like this, you'll sort of feel it under the hood that it's using search better and it's basically more accurate on things because the tool calling is better and more reliable. So I think across the board you should just feel this is a much more, if you're using it as a general Gemini app user for example, you should feel it's just across the board much more capable and more pleasant to work with and use. - One thing that I didn't notice from any of the announcement was memory. I'd just love to hear take on this. I think Google has a real advantage across all the tools, given how much data you guys have across Gmail, YouTube, Maps, everything else for the user. And for me like candidly, the stickiest part of Chatchapiti has been the small memory component they've added. How do you think about integrating that long term across Gemini? - Yeah, we're really sort of going deep into personalization and memory and context and I think this is part of the 3.0 era. You know the Gemini 3 era is to kind of double down on those things. So you're gonna start seeing a lot, us discussing that a lot more as we move into the Gemini 3 era. Obviously there's more models to come. There's the family to fill out, family of models to fill out and we'll be doing that and more features and capabilities that are already built into the model but we'll start exposing more and more in our products and our developer surfaces. So including kind of deep personalization and connection into the rest of the Google ecosystem. Gmail and Calendar and so on. And you're seeing bits of that already happening but it's just sort of scratching the surface of what we have planned. And Gemini 3 is a really capable model that's able to do that. And again, things like tool calling and tool usage are gonna be really important for reliably connecting in to these other surfaces. - Yeah, it does seem very capable with all the benchmarks. I'm just like, I just wish it came sooner. Like I'm using chat to me so often and Gemini, the beating everything in benchmarks and it has all this access. I know it's hard to give a timeline on these things but do you have any like rough estimates on when that real memory is gonna start rolling out on 3.0? - Yeah, we're in testing in sort of dog fooding all the time internally. Lots of different ideas around this and when those things are polished enough and we feel are reliable enough, we'll put them out as soon as we can. We know users want it. We're also building more efficient versions of the model, flash versions, these kinds of things which will allow us to serve it at scale. So, we're very excited about the sort of prototyping we're doing and you'll see the fruits of that very soon. The other thing I should mention is I think that I'm super impressed by with these new model is the kind of multimodal capabilities. Gemini, as you know, has always been really strong. Think best in class, Soto on multimodal reasoning, multimodal understanding and generation, things like Nanna Banana. And we're gonna like up level all of that with this new model. So, I think there's gonna be a lot, I think the general public, the general user will see a lot of benefits of that and we're also gonna start plugging that into other surfaces with YouTube, AI Studio and so on where those will come through, shine through I think those new kind of multimodal capabilities. - I'm excited to be fully testing out and see what the world does with the models as well. Alongside the new model 3.0, you're launching anti-gravity which is a new, a genetic development platform. And it sounds like the platform enables every developer to have almost like a AI co-worker that can operate across the editor, terminal and browser now. But in your mind, what's the differentiator between anti-gravity and the other major agentic coding apps that are out there right now? - You know, I think it's gonna iterate over time but I feel like we're really trying to re-imagine the IDE from an agent first perspective. You know, I think we sort of have the roadmap of where that's going, where we wanna take Gemini, right? And under the hood of that, of course, you can use other models too with anti-gravity. And I think we're trying to sort of re-imagine that and the windsurf guys that we're working on, they windsurf people, you know, they're obviously experts in this. So this is, we're very excited about this area. We're using it internally, you know, which is the first step, right? And people are really enjoying using it. And the productivity is, you know, gains are impressive there. But I think we're still at the beginning of that, right? As the systems become more capable, which we're obviously expecting them to be, what does it really mean to kind of re-imagine that whole experience? And obviously I'm talking beyond vibe coding here, which is more for the amateur coder, right? Let's call it. What is the professional coder one from their dev setup? And I think anti-gravity is our first attempt at trying to sort of answer that and build the roadmap towards that. And then of course you've got things like AI Studio, that's more maybe for the casual developer or single developer, right? Or prosumer, let's call it. So I think we're gonna have different surfaces depending on the level of professionalism and whether you're working in a team, this kind of thing. And I think anti-gravity is, you know, people are really gonna enjoy that. - So anti-gravity is more for that professional coder rather than the vibe coder? - I think that's what we're currently aiming for. Though of course, you know, any developer will, you know, hopefully many, many types, you know, all types of developers will use it. - And speaking of using tools internally, it's just a really curious question I have. I heard Googles using AI to generate a lot of new code now, but are there like internal tools or models that you guys have access to that you're not releasing to the public just so you guys can really get the early benefits of these products? Or how do you guys think about that in terms of like testing tools internally before releasing them and/or keeping it to yourself to get it a leg up over competition? - Yeah, look, we have lots of experimental models and tools all the time. So, and we also have tools that are, you know, at this time too expensive to serve at scale. You know, you could think of like Genie as being an example of a model like that, right? We would love to give access to that, but it's expensive to serve currently. Obviously, we're working on that with future versions of the model. Some of our deep think models, you know, are only available in Ultra because the Ultra tier because they're also very expensive to serve. So we're continually trying to optimize for those things. And then generally when we're able to, it's more of a physical constraint with the compute. When we're able to, we generally put those models for everyone to use as soon as we're actually able to do it from a serving point of view efficiently enough. So it's more that is the dominating factor. We also, of course, do have lots of research ideas and research models going on all the time. You know, that's part of the ordinary course of being a kind of frontier lab with a very deep and broad research bench, I would say, probably broader and deeper than anyone else's. And so we're always trying to pioneer, you know, the next hour for go, the next transformers, what's coming down the line, obviously world models is one of those things. So we're always experimenting. And some of those things, you know, when they're ready, that will put them out into the general public. Then there's also other things too, like hardware and software developments going on, things like, you know, glasses assistant and stuff like that, that we're also, you know, iterating on and starts off experimental before we're ready to show the world in general about it. - Are you guys slowly getting quicker on these releases though? 'Cause I noticed with 3.0, for example, you guys are launching in search off the bat. - Yeah. - So are you guys just slowly getting faster? How are you thinking about that? - Yes, that's great spot actually. So we worked really hard. I think 2.5 was the first real version of that where we had a, you know, world-class model, SOTA model and deep integrations into the main Google surfaces very, very quickly, right? And I think you saw that at I/O, which is what I think a lot of people were impressed at the I/O. I think with 3.0, you know, Gemini 3, we're taking it to the next level and SIM shipping, as you say, with search and AI mode and so on. And I think that's the direction, you know, we worked really hard over the last few months, I think, to, you can think of Google DeepMind as being the engine room of Google, right? So we've tried to make sure we're plugged into all the PAs and powering up every big product. And there's so many amazing products at Google from maps to YouTube to search and of course workspace. And we want all of those, the goodness of everything that we're doing with Gemini and the underlying models to really power amazing new capabilities and features in these products that billions of people use every day and laugh. And I think we're seeing that flywheel really starting now. I think we're still only midway in that evolution. There's a lot more exciting stuff to come and I think we can go even faster. I mean, I think search is a poster child for how we want it to be. And then we now need to do that across the board. - Speaking of useful apps within Google's whole ecosystem, Gemini, the Gemini app has just hit 650 million monthly active users. - Yep. - Thanks. - Yeah, thanks. We're very proud of that. - You guys are catching up to JGPG really quickly. But I'm really curious, at the scale you're at now, is there any specific use cases you're seeing across the Gemini app other than coding that are really useful to your users? - Yeah, we're seeing, actually, I think the Gemini app is really good for multimodal. So I think with NanoBanana, that was a big driver of usage for us. From very fun things, like planning your surprise birthday party invites, whatever, to, like in certain part territories, like making little figurines. So there's so many fun things. One can do comics. So I think using the multimodal capability is something pretty unique that Gemini app is good at. And I think that's driven a lot of interest and I think that will continue. And I think we're also doubling down and thinking through on things like education and health and other stuff that we know users like to use chatbots for. And we want to be absolute best in class in that. And I think Gemini 3 is gonna be the foundation stone for that. But I think, yeah, multimodal and at least for me, I love brainstorming with these things, whether it's like naming a project to sort of sense checking an idea. And I think that the app is really good for that too. - You said something there that was really interesting that Gemini might be like the cornerstone for the health questions. Is there any like more detail you could get into that? 'Cause obviously your background with health. - Yeah, yeah, we've got all these sort of other projects if you like, like co-scientists and we've done a lot of work on this. We've got a system called Amy, a medical diagnostic kind of system that are more in the science team. And what we'd like to do is bring all of those capabilities into the main Gemini. So that's where we're looking at that. I would love it to be what all scientists use to kind of riff ideas on or to do some research. And I think Gemini 3 is a good foundation stone for that. And you'll start seeing rolling out those capabilities, the various forms of Gemini 3, including things like deep research and deep think that are built on top of it. But now with the extra reliability that Gemini 3 has due to the reasoning and tool calling and so on, that should come through in citations and understanding literature. And again, Gemini should be amazing for that because it's so good multimodally and a lot of health and education questions and what users want to do with it are multimodal, right? Here's a diagnostic image, what does it mean? Here's a paper, here's the figure and the tables. What does that mean with the text? Or vice versa, you know, in education, I've got to make a poster about this subject, you know, help me lay that out, right? And generate the visuals for that. I think that is what I'm hoping and we're expecting people to use the Gemini 3 systems for and including of course, the Gemini app primarily. - I'm very excited for that, especially on the AI and healthcare and education, both are very interesting to me. - Yeah. - Looking ahead and again, you might not have an answer to this, this is kind of further down the road, but are you open to and/or looking at using AI for proactive preventative healthcare at all? - We are looking at that in the science team and health teams, like some kind of, you know, medical grade thing, but that would need obviously additional approvals and checks and balances with the regulators and so on. So you have to be careful, obviously, the Gemini app is not a medical grade tool, right? It's for personal use and you still need to consult a doctor and all of those things, but it could be really useful in places where, you know, poor parts of the world, where there isn't very good primary healthcare or education, right? And we're very excited about that. And I know you're very interested in this round as well. And so we think, and because of Google's reach and distribution and Android and things like that, which already, you know, are kind of lifelines in some of those places, I think that we can, you know, it could be very good to at least get a basic level of care and knowledge to those kinds of people, and which could make a difference to them, right? And I think we can continue to kind of improve that with our, and then look to these more medical grade applications and see when are we ready for something to be like a doctor's, you know, assistant or companion or research assistant or something like that. I think we still need some more levels of reliability and I think we're on the right direction with Gemini 3, but there's still a lot more, I would say, that's needed. And we're researching that heavily, as you know, it's a core passion of mine, right? Science and medicine are using our systems for that. And of course, Gemini, we'd love that to be the main foundation stone, and that is the plan for those additional capabilities to be built on. So, you know, we're very excited about that. I'm very passionate about that, as you know, and also with our work with isomorphic and so on. And, you know, I'm pleased with the progress we've made with Gemini 3, but it's just the beginning of that we need a lot more if we want to be really reliable for those types of use cases. - Got it, that's gonna help billions of people. I'm very excited for it. So, switching gears here a little bit to just the real world work, what we currently have with Gemini. Another thing that stood out with the launch was Gemini agents within the app, which is new, which allows you to connect to things like Gmail, which was already there before, but now it gives you tailored steps and actually allows you to execute tasks like sending emails directly within Gemini. As we kind of get towards this AI assistant that is almost your life assistant directly in Gmail, what's your like dream vision for this like digital coworker? Do you want Gemini to be like this standalone assistant type platform like that people use every single day and work like Slack? Or is it just like a separate tool? - Yeah, I would love it to be, we have this idea of a universal assistant, which is a future version of Gemini, that's useful in your everyday life, every moment of your life, where it's a great assistant for anything you might be doing productivity wise, but also in your leisure time, recommending you cool things, giving you ideas about things, riffing with you on things, and maybe also comes across multiple devices, right? So it's on your computer and your browser, but it's also, it's a work, it's a home, but it's also comes with you on your phone and maybe some other devices like a smart glass. And I think that is the future. I feel very strongly about that, that's the future. And I think you need a really capable base multimodal model like Gemini to be able to do that, because you've got to understand the physical world around you and the context that you're in, and of course, to call and use all these other applications, starting with all the amazing Google ones, like maps and workspace and email and so on, but then eventually becoming fully general so it can call any tools. And then I think we're into a new era where you have, just like if you have a really good personal assistant in real life, those of us lucky enough to have really good personal assistants, bringing that helpfulness to everybody's lives, right? And I think that's gonna be a real boon for what you want to do with your time. So eventually, I hope eventually we'll get time back through this, right? And our attention space back as well, even more importantly, so we can actually spend it on the things we love doing and we want to be doing, rather than the things we just have to be doing. - I'm very excited for it. I think that's all the time we have here, but thank you so much for your time. - Great, thanks.

Podcast Summary

Key Points:

  1. Discussion about the importance of Gemini 3 as the flagship model from Google.
  2. Improvements in tool calling, reliability, and creativity from Gemini 2.5 to Gemini
  3. Plans to integrate personalization, memory, and context in Gemini 3.

Summary:

5. There is a focus on enhancing tool calling, usability, and overall user experience, including improvements in style and persona. The dialogue also touches upon the integration of personalization, memory, and context into Gemini 3, aiming to enhance user interactions and connections with other Google platforms.

Additionally, there is a glimpse into the potential future applications of Gemini 3 in fields like health, education, and research, highlighting its multimodal capabilities and the goal to create a universal assistant for everyday activities. The discussion also hints at the launch of anti-gravity, a genetic development platform, and the vision to revolutionize the IDE experience from an agent-first perspective.

FAQs

It continues the progression of Gemini models, offering improved performance across benchmarks and capabilities.

Focus on improving tool calling, reliability, and persona, resulting in better performance and usability.

Enhanced reliability, smarter tool calling, and improved performance in brainstorming, coding, writing, and summarizing tasks.

Deep focus on personalization, memory, and context to enhance user experience and connectivity within the Google ecosystem.

Anti-gravity re-imagines IDE from an agent-first perspective, emphasizing productivity gains and collaborations with expert developers.

Google has experimental models and tools, some awaiting optimization for efficient serving, and ongoing research for future releases.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.