If I look back at the cloud migration, the one thing that the companies that defined the cloud at the time, so mostly Amazon, did extremely well, was that they handled the security aspects extremely well. That was a big fear of everyone. I am going to share my data or my compute with just about everybody else. Won't they be able to see what I have? And they've been extremely strong about it from day one. I think it would be good for the same to happen with the everything that relates to AI models today. So we can take those main fears after table and let everybody else be able to confidence. You're listening to the AI native dev, brought to you by Taster. Hello, everyone. Welcome back to the AI native dev. Today we have a very special guest that I'm really excited to have on. We have Olivia Pomell, who is the co-founder and CEO of Datadone, if you happen to have not heard of it, which wouldn't be very odd at the audience listening to this podcast. He's a pretty amazing company. A whole growth story. We're not going to-- we have so many things to talk about. We're not going to cover that too much. But it was founded in 2010. Now today it's valued as the market moves, but I think well known for about $35 billion. It has over $3 billion in revenue. Where we can make that in 2025, $29,000 customers, all sorts of massive numbers. And I think importantly, it is the leader, the pretty obvious leader in the world of observability, really formed and shaped the category on it, with the long series of products that help explore new domains and new areas in a very efficient fashion. So Olivia, thanks a lot for coming onto the show. Tell me if I did a Mr. present, or I missed anything important in that weekend, Sean? It's all good. That's definitely me, definitely, Datadog. And thank you for having me. Super happy to be here. Cool, cool. Olivia, I think there's a million things we'll talk about. But really, we're going to probably sink into AI and observability. I think a few people are more qualified than you do share both a perspective of where we are and where we're headed. Maybe we just start off a little bit with just some taxonomy here and some definition. So when you think about observability, how do you define this world of observability and what it contains? Yeah, so observability, by the way, is the worst category name to pronounce. And I do that in any role in any hundreds of times. So first of all, I'm not in love with the word, because it seems so much reductive and passive, whereas a lot of what happens is very much not passive. But the observability is the modern version of what used to be called a monitoring before. And I think what it represents is, and obviously, the idea of understanding everything you need to understand to see who an app is performing, which is actually happening inside of it, whether it's doing what it's supposed to be doing. It's an evolution of what used to be different categories. In the like, five, 10 years ago, we used to have infrastructure monitoring, used to have application performance monitoring, used to have log management, used to have network monitoring. These were all like little microcosms of little categories. And I think in part, thanks to the work we've done, those have converged into a bigger category, which is called observability. That's very much end to end, goes from what runs on the servers when it goes over the network, but also what are the end users of these applications doing and what is it doing for the business? Yeah, definitely have been a part of that journey with a variety of solutions during that time. It was interesting to see it coalesced today. That scope is expected, it's table stakes, when you think about building an observability solution. Yeah, and I like the point around it not being as passive as it implies. We're going to dig into that little bit when we talk about AI and agents. So that's the definition of observability. And I guess at a high level, what do you see as the key opportunities for AI within this world? We're going to start a bit meta and that we'll drill in. Yes, so it turns out there are big opportunities that every single layer of the stack with AI. The first one, I would say I would call it maybe the most dreadful word and boring one, is that there's just more applications being built with AI. And those applications look a little bit different in terms of the stack they use, so more GPUs. For example, a lot more data, things like that. And we see that today already, we see companies that are building AI, that are consuming a lot more infrastructure building with a lot more data. So that's the, I would call it the level 0 of the opportunity. And these are applications built that are powered by AI. There is AI in the functionality of the application. And even more, these applications that are implementing AI, largely, say if you're going to consume a lot of GPUs today, you're probably a model builder or a company that is powering the agents that others are going to be using. The second level on top of that is not there's a lot more applications that are built on top of models. And this is a different way of building the application. The application is not as deterministic. The model build a lot more. And the way you manage, you build, you manage, and you understand what's happening is fairly different. So we'd say that's another new version of observatories. That's the second level of that opportunity. And then the third level is, OK, great. Now we have all of these AI technology, all these new building blocks, all these new smarts. What can we automate? What can we build so that the engineers don't have to go on solve issues themselves, but the machine can do a lot of that for them. It can detect a lot more and can solve more for them. And so we are heavily investing in all three, basically. The lack serving of communities that are building the AI today, serving of communities that are building on AI, and then ourselves implementing with the AI to do a lot of automation. Got it. And this is the actual audience for these things is in reverse order, right? There are probably 1,000 companies that could benefit from AI-powered observability, if you will, for every application that is built using AI on top of the models, and probably 1,000 of those for every model builder of this. Is that how you think about it as well? Yes, yes. At the end of the day, we think just everybody will want automation, not even in the world will want automation. That's the dream, right? The dream is you never have to wake up in the middle of the night ever again to fix an issue. It's been for you, and maybe you cheat it when you're up the following day. Then I would say most companies will build on top of AI models, and they will consume these models, and they will have to understand who they work. And a smaller fraction of companies will build the models themselves and operate vast fleets of GPUs and AI infrastructure. And there's an inverted parameter there in terms of the number of companies that these apply to. It just so happens that today, the biggest needs are with the companies that are building the models, because they're upstream of everybody else in terms of the consumption of AI. Yeah, and they're using. They also became big operational consumers very quickly, and every time also they need the uptime, they need the inference, stability, and then eventually they need the training and large quantities as well. Yes, OK. So I think still focusing a sec on the first on that third category on improving the products using AI. So improving, I guess, the large swath of data dogs products of it. You've launched a bunch of things already in this domain. So again, just starting from setting the stage here, what have been the biggest need of movers or the biggest opportunities you tap into leveraging AI to better serve the broader set of customers, even though it was building a classic applications if you will. Yeah, so we have just a few things we built, a few products with names on them that we've pushed out. One, historically, we've had a product for anomaly detection, that's called Watchdog. And the idea there is basically it watches absolutely everything that's going on. And it tells you when there's been a spike in errors for specific service or a step change in the request rate for another service, and things like that, which really help. Obviously, you can't be watching millions or billions of time series all the time, and that's the point of it. So that's what Watchdog does. A lot of it is, I would call it machine learning 1.0, I mean, that it's not using financial models largely. It's using statistical models that run very cheaply, and that can run all the data all the time. That's been there for quite a few years. Yeah, if you ask me the way the chat deputies that's going to tell you, it's good old-fashioned AI, which I cannot entertain. Thank you, chat GPT. By the way, Chat GPT also says that I'm a snowboarding champion, which I don't know, I've never been in a snowboarding. OK, interesting. Some hallucinations that to play with over there. Yes. There's another product we've really is more recently, which is called BTAI. And that's the more hygienetic version of what we do, which is bits, runs investigations. It facilitates response to outages and things like that. And you can interact with it, and it can run on its own. And so this one is much newer. Obviously, it's built on the newer, junior generation of LLMs and all of the other even newer models that have come out more recently. And this one is heavily developing as being the future of automation for customers. With applications, initially, for two yards on call situations or after situations or fixing errors that are detecting on the production or helping optimize costs of performance or helping identify root cause of incidents. So these are the kind of use cases that are implemented with bits. And there's a couple more products. One is LLM observability. This one is built to help operators of applications built on LLMs understand how they behave. And we can dig more into that. The last one is we've announced also a foundational model for time series, called Toto. And then left to the Wizard of Oz. The dog in the Wizard of Oz over there. Just interesting character to name it after. Exactly, exactly. And there's a clever background. And the idea behind Toto is built basically the best possible time series prediction model. And we can also talk a bit about it. I think it's fairly interesting. So our very first model, and obviously, there's a lot more building on top of it. So these are, I would say, four of the main directions worked, or the products we've crystallized our directions into. Hello, it's Simon Maple here. One of the hosts of the AI native Dev. I wanted to say thank you for listening to the show and to remind you that we post great content here every week. So be sure to subscribe to keep on top of our episodes. Also, you can send your questions to
[email protected] or ask questions directly in the AI native developer discord community to continue the conversation, connect with other developers, and explore more about AI native development. All links are in the show notes and the episode description. Now, back to the episode. Yeah, cool. OK, so thanks for that. Set the scene on it. Maybe let's start poking holes a bit at the opportunities and the challenges involved in them. I guess the first thing I want to talk about a little bit is trusting the insights from these tools. You mentioned bits. And in general, there's a lot of intrinsic understanding that is easy to rock around the fact that observability deals with fast amounts of data. And the AI is amazing at data. It can look at all of these things. And now it can read the lines and make sense of them and give you whatever it is, the root cause analysis, or even choose one to wake you up at night. But as we've just discussed with some snowboarding examples, it also doesn't always get it right. I guess, maybe what's your sort of sense of the current accuracy level of detecting these things and giving you the correct answer? And what's the appetite for experimentation in this domain, its users? Are they accepting, shall we say, of mistakes? Do they expect perfection? Yeah, I think so. I'll start with one of the biggest learnings we've had when we started their dog was that there's one like a submerged will tell you, which is if your system has a smart to detect issues, I prefer you to give me the false positives and then I'll decide whether they're right or not. And that's a lie. The reality of it is that you send two false positives to people in a row and then they turn you off forever. There's very little appetite to chase down the wrong path because of a machine that told you. We'll do it because of a human corker, but they won't do it because of a machine. And so the bar for precision needs to be very hard on anything that relates to ability. And the good news is that I would say three, four years ago being completely right about root cause most of the time was science fiction. And I think now it's definitely within reach. We see it in the way we've been developing. We see it in the progression of the evaluations for that. And as you've been building on AI too, evaluation. The smart. It's a join a nightmare to be building with a predictability, but you definitely see the progress. We see actually very good progress there. And we can clearly see on the horizon the moment where the technology is good enough to be put into the hands of customers in a large amount of situations. So that's where we are. But yes, the bar needs to be super high. One thing we found interestingly is that the bar is a little bit lower in some relative areas, such as security. I think the users and customers are more willing to take a bet on automated decisions or automated notifications for security than they are for operations. And I think the reason for that is that the risk reward is already different. I think it's okay to crash a workload if you avoid a security incident. It's less okay to crash a workload to avoid crashing a workload. Yeah, yeah, there's a little bit of running. It's really interesting, though, because first of all, like the security, my experience at Sneak is actually very similar in terms of the false positive, false negative element, right? It starts off by saying find all the issues. If you're not sure, tell me about it. I want to know the reality is that's actually the primary thing that they would feel. Don't tell you. But it's interesting, though, you're a comment about identifying flaws or maybe having more tolerance around sort of alerts because to an extent, identifying whether something is or isn't an attempted breach is also something that's less definitive. So maybe just people have lower expectations. They just don't expect anyone including human assessors to be able to spot what is correct or is not correct. So it could be the impact or it could just be the sort of the expectation of the art of the possible. Yeah, yeah. Also, just to level set what we scarce, I think when you compare what's happening in operations and security to what is happening in other areas of AI, these are hard problems to solve. When you think about, we think about self-driving facts, for example, which are just about working now. What these do are what pretty much any human over the age of 16 can do without even thinking about it. When we think of the incidents at least the ones we buy the ones we've seen our customers have, that it takes large teams of very qualified people and PhDs and a lot of time to try and what's going on. And sometimes weeks after the incident, you're still thinking about them and seeing working through them and understanding exactly what happened. So these are really hard problems to solve. Yeah, no, I agree. And I guess on top of that, there is a notion of understanding the application itself, right? Yeah. So I guess, so maybe two questions coming out of that. One is trust is imperative, right? People want they have low tolerance for mistakes on it. But the systems are, as you've also mentioned, not predictable. I guess one question is, can they find it? But even if they could find it, it doesn't mean they can find it again, if given the exact same set of information. So I guess how do you think about solving it? Is this solution sort of the co-pilot world and leaning into that until it's good enough? Is it about narrow slices? Is there a different lens? Like how do you still make progress while you don't have perfection? You have expectation of perfection. And the systems don't get allowed. They're getting close, but they're a line of sight for perfection. I don't think it's present, right? Yes, I think there's two issues there. One is, what's the right form factor for the customers to interact with the right way? And that's where you think about, is it better to have a co-pilot or is it better to have an agent that does things on its own? Do you want to interact with it vocally? Do you want a button? Do you want, like, how do you put that in the back? And that's a big question there. Do you why it really sets up the expectations from the end user? Well, there's a big thing there. And I think we're still, as an industry, we're still figuring out what works. What's pretty clear is that the chat interface, it was a great start, it opens everyone's imagination, but it's not the build end all. I think most of the AI-based functionality is not going to work from a chat interface, the long run. The second question is, how do you build, or how do you enter the market with a solution that shows confidence and/or builds confidence with the end user? And I think for that, you can get higher precision by having lower recall. And you want to start in, maybe with corner cases, maybe with some small sub-selection of all the things you can, all the incidents you could possibly solve, but focus on the ones and have the really high precision first. Now, the challenge with the naive approaches is that the LLMs in particular are not good at understanding when they know what they don't know. They'll always gladly answer. And might be wrong a lot of the time. So I think a lot of the technology that we're building there is around understanding when we're right, when we're not. When we have a super high chance of being right, so we can actually volunteer guests and maybe solve the issue directly for customers, as opposed to not seeing anything. One advantage we have as a company that's already used for observability is that we can choose, pick and choose, basically, the use cases and the issues for which we're going to automatically take action. And so that's what we're focusing on today. - Got it, make sense. So you're already monitoring the whole thing. And so you can make a lot of attempt, see where is it, what are the cases in which you can reach a high-confidence conversion of it and then reach in, which definitely is kind of an incumbent advantage of already being in place in order to see the data. - Exactly. And let's start the choice that's so obvious. You don't have that choice in some other parts of the industry. For example, for self-driving cars, you can say I'm only going to take two to make left terms. I think that's the luxury we have in the industry. - Right, in that sort of breadth. So that makes a lot of sense. So I guess maybe leading from this into the sort of the world of agents of it, you're in the spot, you look at all the data, you see the data for the specific cases where you have conclusions. How do you think about actions? What is the, these are the insights, maybe that you have higher confidence. Do you see some that you reach and like such high confidence that you can even act on them? How do you think about the agentic responses to this data, the non-passive version of observability in its stance today? - I think that's definitely the goal and we see areas where we can do it. I think the lowest hanging fruits there are things such as errors, for example. So you have, you ship a new application and you have some 500s on the servers or that's pretty clearly an error, pretty clearly needs to be remediated. It's usually not that hard to generate the code that will fix the errors, but then there's quite a bit more work involved into getting that code across and validated and running in production. So that's one area where you can take action, I would say, pretty automatically and pretty quickly. And that's definitely something we're working on right now. There's a few other areas. I think it's a, luckily a few years ago, we started building the ability to take action for customers and so we built a workflow automation product. We built functionality. So customers could, could be an application that automate the way they run their own operations internally. And that's a big departure from what we had been doing for the first 10 years of the company. The first 10 years we were very careful about only receiving data and never taking action, never reaching back. And I think all of that is really coming to fruition now that we're trying to close that loop and automate response for users. I'm a firm believer that for AI to show its value, you need to automate action. I think if you have to keep pestering the humans to do things, you're not going to be successful. I'm a firm believer in what you say. Like you have to work towards that sort of autonomous piece. But really the key question is that sort of build up of the trust. And I think much of the industry has taken the kind of co-pilot approach of it. I think you can have another unique opportunity of it, which is to take the cherry pick on ProJump. I'm just going to look at all that data. And I'll pick what works. I guess what has been the most surprising or the cases that you gave me an example of something pretty obvious like the system is basically just dead. It's just returning all the sort of 500s on it. Have there been like a positive surprises about things the LLM starts showing you that are insights you didn't think they'd be capable of? In general, I think every thing that's impressive, you've probably seen on Twitter already. Or X, I should say. That's the blessing on the curse of LLM is that it's so easy to make a mind-blowing demo. And so hard to make that work reliably all the time for everyone. And you've seen examples of you. You have a crazier and immediately understand what happens and write beautiful good to fix it. The problem is it doesn't always work that way. And then understanding how to get that good to actually shape it actually is usually pretty difficult. I think we're still very much in the face of making sure all the plumbing is there so that this works reliably all the time for everyone. - Got it. So I guess let me take this a little bit like in that practicality of taking these lenses and maybe we'll talk about all aspects of it. And one that comes to mind is security. And so you're dealing with analyzing logs, analyzing systems clearly a lot of untrust the data. And we've had several conversations here probably because of my kind of passion for the topic of it about the security and how those for an opportunity for prompt injection that might even come from those those sort of untrusted fields that might try to elude or fool the agent, be it source of some other product to take action when it shouldn't. I guess my kind of attitude of two questions. One is how limiting is this right now? Is this a theoretical concern you're seeing in practice? And then how do you think about overcoming that? It's another aspect of trust, right? That is not about capability but rather about adversarial activity. - Yeah, it's always been a concern, right? So it's already the case that and we've seen that you can try and inject stuff through logs, through traces, through pretty much the thing that submits data. It used to be cross-site scripting and you inject with things into the logs and we've definitely seen that. So that means is that everything needs to be heavily sandboxed. Everything needs to be, like you need to code, defensively against any data that comes in. Obviously with the LLAMs, like their surface of attack is a little bit fuzzier. So the problem is I would say harder to fully contain as long as you end up passing some data to the LLAM in the end. That's definitely something we are aware of and something we're building against. We see that problem too in that, when what I just mentioned, to solve, to resolve errors and fix the code for that, you do end up having to generate an execute code. Anytime you execute code, you add great risk. Everything there also needs to be heavily sandboxed, heavily isolated so that you can do that and not fear the consequences of having all of that trusted code that's going to execute everywhere. But in general, it's good to have it to imagine that anything that processes external data might end up executing and trusted code. I think the interesting, I had a Mateo on the show on it who is a Lakerra, which is a security company. We talked about two things. One is how indeed the control plane and the data plane insecurity or sort of in the world of LLAMs, the control plane and the data plane are merged. And so you're basically executing everything, like all the data gets run, right? It gets executed by the LLAM, so it's hard to separate it. And two is that from an agent perspective, if you think about maybe one definition of agent security versus a LLAM security, is that LLAM security tries to fool the human, right? I might try to get the LLAM to give an alert that shouldn't try to reach inside the student, but then agents actually take action on those. And so we're taking it up to the next level, right? You're actually trying to guide the system. And so it's interesting that you say that you run code to generate it, I guess that's a more powerful, but more dangerous application to it. But in general, do you think this is an eventual limit to the autonomy of these systems? Or is it, I guess, do you anticipate that our sort of trust level in the agent's decisions would reach par with our trust level, with the sort of the humans that make the action? Forget how you execute it well, but you've chosen to restart a server, you've chosen to redeploy code. Those are decisions. Yeah, I think what this means is we probably just won't have completely end-to-end models, so we can build the right limits around the models, so we can build enough trust there. I think that might be different for the desirable end studies, maybe in self-driving. Instead of driving, you hear that it now is building around end-to-end models. You get cameras in and Accelerator.app, and that's it. I think that's going to be a little bit different for the kind of systems we build precisely because we want to have all sorts of checks and balances and validations along the way. So you can't have adversarial systems or injections and things like that yield to long actions there. And on our end, it's also an opportunity because we also build the IT tools, both for correctness, but also safety and performance. And so these questions we have when we build those agent-to-systems around, how are we certain that they're not being abused, and they're doing the right thing, and we can trust them. These are all our customers have, and they want to answer for the one applications, and so we build a tooling for that as well. Yeah, you're able to monitor your own actions, to an extent, and have the time at the right competency when they both attack and the team itself, which is pretty cool. I think that's really interesting. And I wonder, I spent a bit of time thinking about the scopes of autonomy, and oftentimes, like self-driving cars, a good example of it, you don't hold them accountable, or you don't know who to hold accountable if they make a mistake. And so agent systems might have the same problem of the limit might not be a technological one, but rather worries it that you're whatever. What is a reasonable liability for a data dog as a company to take for taking an action, otherwise would have been a user's decision for it? Does that come up like this sort of risk management? Almost like compliance risk, legal risk threshold on it versus technology, as you think about it a bit, though. It's not uncharted territory. You have systems not automate updates, for example. And as we've seen in several very public cases, like an update can have really a horrible adversarial impact on the customers. So that's not completely new. Of course, there's going to be maybe extensions of that. Maybe there's going to be more refinements of the legal framework around it, but I trust that that's going to get figured out. But to your point, every company, not every company, but every other company, has had an issue where the intern dropped their production database some day trying to fix an issue that might happen with all the versions of these agents, too. That does create an interesting question of who ends up being responsible for that in the end. But again, I think it's not very different from all the values ways in which automations are living used. Yeah. And say how much you've seen firsthand or been part of the immigration to the cloud. And I think it's true that in cloud, you had to be a little bit embracing of the failures to be able to be in the lead, right? Those who embrace the cloud at the beginning suffered from some gaps from whatever. Container sandbox vulnerabilities, the reliability issues of some of the clouds at the beginning, to things that were not irrational fears, but actual kind of flaws. And you had to maybe embrace the failures, right? And lean in and accept the slightly finicky behavior of this system at the beginning to be cloud native and benefit from those advantages before the rest of it. Do you feel it's the same here? Is that all of these concerns that I'm raising and all of these fears are ones where, hey, if you want to be one of the leading companies of the future, then as a user, I'm not even talking about data dog with a company, although that's also true, but as a company building right now and using tools, embrace it, embrace yourself for having some false reports, for having some sort of system that they needed, because if you don't, you're just going to be left behind. Does that resonate with you? Yes, I think the technology is definitely rough around the edges for anybody who's building applications on AI today. I think that's definitely the case. And that's also why, in a way, there's a relatively small number of AI native companies that are developing fast and shipping a lot, but the broader set of companies that are building on AI are still mostly testing the water. Like they're still having to have applications in validation, in early tests with their own customers, in proof of concept, pilots like you name it, but it's harder to get all of those into production, just because the failure modes are not well understood yet, and the technology is changing fast. There's a bit of discomfort with all of the new exposure you might get with it, so we'll definitely see that. I would say, if I look back at the cloud migration, the one thing that the companies that defined the cloud at the time, so mostly Amazon, did extremely well, was that they handled the security aspects extremely well. I am going to share my data on my compute with just about everybody else, won't they be able to see what I have? And they've been extremely strong about it from the one. I think it would be good for the same to happen with everything that relates to AI models today, so we can take those main fears off the table and let everybody else be able to be confidence. Yeah, I think that's a really good observation, and it's interesting, because while I 100% agree with what you said, I also think cloud security misconfigurations are probably today the biggest sort of sort of breaches and security flaws within systems. Those are not direct flaws of the AWS cloud providers, however they are created by the ease of creating infrastructure and then configuring it. So you do have to wonder whether prompt injection will basically replace that world and it's be the next vulnerability. There's a risk of that happening. Yes, we replace the open S3 bucket by head. There's a prompt in there, there's a chat in there, so I know exactly what I get out of it. Basically, with the statistical attempt, the source says whatever, send me all your data from any sort of chat interface that you have that you come along. Hey, let's get the technical a second year and talk about how to get observability right. So you've made the decision to build a foundation model in total. The more I think about it, the more I love the name is for a data dog. It's the penny took a while to drop here. And what prompted that decision? Why not build on top of the existing LMS? And because do you think of it as a replacement as an addition? It was got a very different focus. The goal is to work on time series and to be predictive for time series. For that, we, in the broad models, don't work very well with time series. They get you some things numerically. They're not great in general. And for time series in particular, they're really not fantastic. We also needed something that could be very small, so that it could be run in process on very large numbers of time series. So the first version of Toto is purely numerical. So we fed it a number of time series. We have large numbers of time series of data dogs, you can imagine. And this would be like metrics from operating all sorts of aspects of systems. Exactly. So some of the data is traditional time series model training data, which is not for us. Some of it is like weather data and things like that. And a lot of it, or actually, as most of it, is our own observability related data. On which we have some pretty strong signals in terms of which time series are interesting or not, because we see which ones are being used for alerting. We see which ones are being looked at, which frequency things like that, which gives us a good sense of quality of the signal. And what was somewhat shocking to us was that that very first model we built, which was built, I would say, in a little bit of a naive way. It's the first time we build one of those broad, no deep learning models. So the first version of it is already, or was at the time of release, instead of the on time series prediction, including for known observability use cases, which shows, of course, that deep learning is working and scaling. I think we've learned that over the past couple of years. But also the value of having great data to train those models and who really makes the whole difference. So this first version, we haven't released it to the broader world. I think we really wanted to use it as a starting point, as a baseline, for more research we wanted to do on a model. So we're busy building new versions of it today. That are multi-model in that, I would say, it's a lower version of multi-model compared to what you can think of for the-- - The finish of the-- - Yeah, yes. Doesn't have vision, but it uses, in addition to numerical data for time series, it also looks at the text data that surrounds those time series, whether that's the descriptions you have for the time series, the names, the time series, the tags, but also the event data that we're going to see, such as what happens in logs, or other time series, not just one time series at a time, but multiple time series that can be correlated, and we're pointing all of that into the model. And the point there is to make it really good at being productive, so that we can detect normalize better, like the animated detection is all about comparing prediction with this reality. But also in some situations, we can actually get ahead of incidents. We can say, "Hey, this monitor you have is going to appear in one hour." And we know that because your monitor is on the, I don't know, yeah. It's on the CPU of your database, and we see that the queue of queries that are lining up, the database is going up, we don't see it, we don't encode that as a rule, but the model has learned that it works that way. And so we're pretty excited about it again, because we've seen really impressive results in the very first research version of it, and we are really at work to build an expression. Yeah, it sounds super exciting, and I relate to the gap that you see in the model, by the way, a bit of a trend in the world to find specific niches in which you can build slightly smaller, slightly faster, but also better models for specific slices, but using the transformer technology. So love seeing that, and yet clearly have the data. I'm curious like how you think about it from a product perspective. Is this meant to be an engine that only manifests in data dogs products? Is do you think of building an open waste model and having the world build on top of it? How do you think about it from an offering lens? So everything's in the table right now. So it definitely will be part of data that gets on point, and that's why we're building it. So we'd love to have an open waste version of it. I think we're being very careful there in terms of what data we have and how we use it, and contractually what we can and can do, and what's not right to do with the data, but we would love to have an open waste version of it. So we're working on that. Yeah, and do you envision eventually creating for large enough customers models on their specific data? Have you experimented with whether that achieves better results or not? That's definitely also one thing we're experimenting with. Today, we see that the broader models work better across customers, because customers also have a diversity of data they're using. But at the same time, they're using a lot of shared components. So customers are using mySQL and Postgres, and it's the same database across customers, and their sales metrics have sales in them in the name. So it's also somewhat semantically similar across customers. So we're experimenting with that. I think we'll see where we end up. Yeah, that makes sense. I love the analogy is sometimes experienced individuals, right? And so if you think about having a very experienced SRE come into a brand new system, they actually will have a lot of knowledge to come along. Part of it is family already with the Postgres that is deployed here that they've seen before. Part of it is just analytical to say, I know that when memory consumption and CPU consumption behave like this, it's recipe for trouble. Yeah, I like the idea of bringing that knowledge from customers to customers, and I guess it's interesting whether there's a fine tuning layer on top of that on a specific customer's data. Yeah, and one thing that we've learned over time, too, is it's critically important to have lower friction for products to be adopted and to show value. And so it's very important for whatever you have to work well on their one, and to work well with who just gets started. They send your data, and they want to see things happen right away. If you can get better over time, it's better, but never get there if you don't show value very quickly. Yeah, love the product responsibility. So I guess one more question on that domain. Again, when you think about the person, and their knowledge of it, infrastructure knowledge is useful, but also familiarity with a specific app in question. Have you looked into and do you anticipate, including code and application knowledge of it, things that might not have traditionally been in data dogs' remit? Yeah, so application, we get already because we do tracing, right? And we do, we understand what calls what, and what's being run, and things like that. Code, it's been a big effort from us over the past few years to get the code into data dog as well. The main driver for customers to start tracking their code in data dog has been using our profiling product. So we have a continuous profiler that runs all the time in production, and really benefits from having the full source code, fully readable source code, a certain type, which is actually executing to the code you have, and basically we're pushing towards that. We've also been building extensions to the IDs, so that customers can navigate directly from wherever they are in the code to what's happening on production, as it has been a big effort from us. And I guess if you casturize further, I'm getting a little bit ahead of myself because I want to finish with some features here, but oftentimes think about AI kind of impact on software development, or Gen AI impact on software development, as really a continuation of the same trend that we've been seeing with DevOps and convenience deployment and all of those, in which it accelerates software development, it allows more independence, more software gets created, and updated faster and faster. This is another multiplier. And I've heard you say similar things in past podcasts on it. In that case, how do you think about the loop between observability and bringing that knowledge back into code and modifying the code back in? Do you see that accelerating? And does that impact the kind of the remit of a platform like data dogs? Should we see the expected scope of product or scope of platform change? So I would say, so just to level set, I think there's been a continuous, as you said, improvement in productivity for development. And that dates back 40 years, right? If you compare where we are today, to where we are 40 years ago, or 50 years ago, with punch cards, and possibly a great genuine, already a thousand times more productive if not more, in terms of what we can be able. We've had the more advanced languages, computer screens, we have open source libraries, we have the cloud, we have all of those things. So you can do, it can be incredibly sophisticated things, with very little code and very little mental exercise. I think we're looking at another 100 times or a thousand times increase in productivity with AI. And as we go through this productivity increase, developers have less and less of an understanding of how the application's working and what it's doing. And so we're switching more and more of the work there from, hey, let me type the code, to, okay, no, I need to understand, is it working? Is it working well? Is it costing me more money than it should? Is it doing what it's supposed to be doing for the business? Or did it break? Did something happen when other things, other systems change, or the customers change, or something else happen? So these are the big questions we answer. And these are, I would say, the questions that are becoming more and more valuable as the productivity gains stack on top of each other. >> Yeah, so I agree with that line of thinking. And one of the things that I wonder is, if you're not writing the code and you have less visibility, you have a need to understand those systems on it, like today, when you observe, you're still observing in large parts of the world, you're observing code. Or eventually, when you find a problem, when you see a flaw, it's either in the infrastructure or in the code, those are the places you change to evolve. >> How do you think about, if you're less familiar with the code, how do you observe a system? Does that require some fundamental change to the core tenants of how observability is done? >> I think the observability will be end to end. I think what you observe in the end is you observe your business, my users doing the right thing, and they're finding value, and they're being served properly, and that's what you care about, right? And you might have more specific questions, which are, oh, I shift that future, is it being used? And it's near difference. And even more specific questions, such as, oh, I changed this line of code, does it have the right impact? Because sometimes you will still have to do and make these specific changes. I think the future of observability is that covering this end-to-end and helping answer those questions, and plugging back into whatever the developers are using to conceive and build those applications to start with. Today, it's largely a number of different code editors and IDEs and they run the code and they ship it. You're building the future of that, in many ways. Others are too, and we'll see what the end-state is with all of those different new ways of building and conceiving applications. Yeah, I love the user-focused aspect of kind of this or business-focused, really, of so saying. Eventually, without your respect to the code, what you're observing should be your customer's experience or your sort of outcomes and see if those are right. And those, you shouldn't care about the code. It doesn't matter if you've written a spec or written code and doing it, the problem identification should focus on the outcome. From here, you talk about how do you identify why the problem occurred, so you should adapt to whatever it is, whether it is a spec in Tesla land or whether it is code or whether it is something entirely different in some other solution path. Very cool. So we talked about some sort of trust. We talked about just how do you get observability right and maybe some level up into the future of it. I'd like to take a slide detour and talk specifically about your sort of second category from the beginning, which is applications built on top of LLM. I'm going to put aside for this conversation of those building the actual LLM. I think that's specialized knowledge and the drill into that too much. I guess, first of all, clearly, the world of observability as you pointed out, has expanded to include logs, to include APM alongside infrastructure monitoring. When you think about LLM observability and observing those types of applications, do you identify kind of new categories that you think will get added to an observability platform as a default? Yes, look, we call it just LLM observability for now. It might have a completely different name and shape to yours from now for one thing. I don't even know if we'll talk about models as LLMs. Maybe we'll have a, maybe the shape of what's provided by the model vendors is going to be a bit different. Maybe the naming is going to be a bit different. We'll see. Maybe the focus on language will be reduced, I don't know. But-- That's the way you're actually embracing AI as a term because it is a bit more generic, just-- Genie AI. Genie AI, it's a problem is the more you-- the more often you pronounce the words AI or the letters AI, the more bullshit it sounds. How do we care if we buzz word in the world, right? Organic, anything that would happen eventually with this meaning? Yes. But we do think there's a growing part of the applications that's going to be non-deterministic. So you use models. And those models are trained. And maybe they're continuously trained, maybe not. And what you have to measure is outcomes. Like what actually happens, what do they produce? So you can't be in a world where everything happens when you define the application. You write the spec, and that's it. And now you know, from the moment you've done that, you know the application is doing what it's supposed to be doing. You actually need to understand what happens on the other side of it. Now, you've shipped the application. It's being used. Is what is happening there similar to what you thought was going to be happening? Are there new things happening? Is it happening in a safe way? Is it being abused? Is it helping to reach the right business outcomes? When folks use this non-deterministic part of the application, do they buy more? Do they stay longer? Do they do whatever it is you want them to do in the first place? All of that are super hard questions to answer. And you need a new form of tooling for you. What's interesting there is that the companies that are building around that, they're typically not the companies that are building models. Like the other companies are building on top of models. We see their needs evolve as the world of models evolve as well. So I think it's probably going to be a couple of years, at least, before we know what this particular part of the stack ends up looking like. It's super interesting. So if I go with this bag, you're saying you say at a low observability and people might think, hey, you're capturing a chat trace and things like that. And those might be correct in the meantime. But really, the real problem with that low observability is that you really need business monitoring or you need outcome monitoring to be able to define it because the product is going to be unpredictable. And so, yeah, you need the technicalities. You need the traces as well to be able to troubleshoot problems. But if you want to know if there's an outage, you don't know if there's a problem, you're not going to get that from the traces. You're going to get that from the outcome measurement. Did they get that right? That's right. I think the trace is very low level. I think it's rare that you're going to go and read step by step what happened either in a chat or in the agent or see what actually went on there. I think it's going to be a debugging and validation of some hypothesis you have but what happened. Whereas what matters at the end of the end is you answer all the requests. Did you generate transactions? Did you generate some of something else? And what you will see is yes, maybe you did, there used to be 1% of those conversations that were not satisfactory to the end user and now there's 5% do we know why? And then you will dumb, you figure out, okay, actually it turns out those 5% are all in the same region of the embedding space and that's because maybe there's a data source that changed and now we have bad data for that or maybe the models behind or maybe something happened there that changed and that poses an issue and you can validate that by looking at the specific conversations and all the specific interactions between the agents and the various agents and you understand what went wrong there. So this is a whole new world of understanding what's happening and debugging. And again, what I was saying earlier is the main challenge at least for us is that the space is changing very fast and the use cases are changing fast. When we first launched that product a year ago, what we were seeing were mostly chatbots. That's what people were implementing. Now we see a lot more agents than chatbots. So instead of end users asking questions from the models, you have agents asking questions themselves to the model and looping. And so we'll see where we end up, I would say a couple of years from now. Yeah, I love the thinking here and I can easily see that indeed being the future of the reservability. I guess that sort of brings you much more lensed lines with product analytics companies and comes back to basically expect those spaces to merge like the infrastructure application and then also business and product analytics. Yes, we have a product that's fairly new in product analytics. We built that on top of a real user monitoring product, which is basically not track the interactions users have with an application. We started from the performance aspect. We started with user monitoring from the question of do you have errors and what's the random speed and noise fast enough, there is a degrading, things like that. And now we're moving more into the understanding of click stream, what are people doing, what are they using, what are they not using, really get the outcomes you want from that or not yet comes you want. And obviously, as I mentioned earlier, I think this is critically important for anything that is related because for typical application, you can have a, after you develop it and after you ship it or add the money, ship it to production, you have a pretty good idea of whether or not it does what it's supposed to do. For any application, you don't and you need to measure in production. Yeah. And I think that's probably like the core of the problem, which is, I guess, observability or identifying a problem requires first understanding what isn't a problem. And that is just very hard in the world of Genai or LLAMs or how you call it because it's just very loosely defined. The systems are unpredictable to begin with, the types of tasks and missions we give them is very broad. And so how do you define it in the first place? Do you anticipate anomaly detection to even be possible or can it only be anomaly in that sort of slightly more narrow slice of business outcomes that you can identify the anomalies but not in the activity? There's a number of anomalies you can detect. So you can run all sorts of checks on the statistical distribution of what comes out of the models. You can understand when there are changes, when you drift, there's a number of metrics you can track. There's even ways of fairly reliable ways of measuring or estimating whether or not you're getting hallucinations without even having access to the ground truth just by looking into distribution of the data. So there's good things you can do there that are good heuristics, I would say, without having the users define, hey, this is my business objective and this is what I want to monitor. Yeah. Yeah, very cool. Yeah, very much, I guess, at the end of the day, extends the scope of what you are able to do with the data that you expand. So maybe one, maybe bring it back to security once again, apologies for that. Another lens of attack is not just prompt injection, but also as the data that a user interacts with and such blends with the execution and the functionality. You also get into topics like views or fraud or such, or the interaction is one that is very textual. If you have thoughts of those, is that a red line you're not going to look at the data or is that also within the scope just trying to get a sense of where you see the lines forming if anywhere, really? Yeah. It's hard to know exactly what the lines are forming yet in great part because we don't know yet what's the final feature set of the models themselves. So we will, we'll, a lot of the commercial models and then extension all the free models after that, would they include some level of built-in protections, built-in filtering, so we don't know that. Today they don't, but the space is young enough that it might change over time. One thing though that is pretty clear to us is that understanding the applications will involve looking at the data that comes in and out, that goes in and out of the models. I don't think there's any way of building observatory in the future without that. And for us, it changes the profile of the data a little bit. When we were looking at traces and logs and things like that, this was mostly not primary data. Like we wouldn't see the doctor's prescriptions or the social security numbers, like those typically don't end up in logs or traces or anything like that. When you look at the data that comes in and out of the models, or you will get all that primary data. And so that has different implications on what you can do with the data or how you can store it, what access you can give to it. So there's a whole new set of functionality and protections to build on that. Great. Yeah. I guess the good news is from a business perspective, you need to make that investment to get the opportunity to open it up. But then once you have that data, it also opens a set of opportunities for providing value to your customers. Yes, yes. Again, yet as close as possible, it's always been the dream. Like the dream has always been, look, you observe the application and the world is transforming digitally. So the application is running the business, therefore by observing the application, you observe the business. And I think this is becoming even more true in the world of AI, just because observing the application now means actually looking at all of the primary data that is making the business. And that's why it's so exciting to us. Yeah. Very interesting. Let me take it to the world of VUX. And so we've cast your eyes. We've been moving further and further into the future here. As we're having the conversation, five plus years from now, you have good agentic system built up. You've figured out some thresholds from a security perspective that you would engage. You've already mentioned chat is like a limited type path. What do you envision as the interface or how does the interaction with the observability system change during that time, do you? And I guess notably, do you envision it changing to be more like how you interact with the operator today that uses data dog and data dog becomes that interface? Or do you think that's just a fantasy? I think, look, we already see today interfaces that represent the systems as humans, as being appealing to users. So some of it might be chat, though chat in a way that's not in the US in questions, but the machine working with you, because it's hard to know which questions to ask otherwise. We see that also in voice, we think there's many situations where especially when there's a group of people trying to work in real time on a problem where voice will be a good way of interacting. And so we think in the end, the software products and operating as one of them are going to look more like humans you deal with, and that might be the way we interact. I think there will still be UIs on top of it, because UIs are a great way to consume information to understand what it is you can't do, when you start your car, it's good to have buttons and to know what's going on there and not start a conversation with it, something with your oven. I think it helps, and we still will have that, but I think it will be complemented with more open-ended human-like interfaces. Yeah, very interesting. So what we spoke about also very exciting bull case, if you will, opportunities that I think all are very valid and real around expanding it, also just interesting to think about how observability might change. What would you say is the bare case for how observability is done today, where they are? Is there a scary scenario from sitting in the lens of the incumbent in the data dog when it comes to AI? I think the things that are scary to me are hype cycles and things being oversold. So, on earnings calls, I'm not overselling the edge of agents and things like that, because I think as of today, the technology is still not quite there yet, it's exciting. We sit on the horizon and we see all the potentiality, but it's not something that is happening just today at scale on the field. And so my worry is that the technology gets oversold and that trust relationship with the customers, the users, is broken. And it's happened many times before in the monitoring observability. As a company, we've been very reluctant to have AI on our website for the first ten years of the company, just because I associated it so much with bullshit and things that don't work and are oversold that I wanted to avoid it. There's a whole category of product called AIOPS that existed ten years ago, and there's usually nothing AI about it, but it's the AI part of it is a regexp, so that's what it was built. So, I want to be careful about underselling and overdelivering and I think it's the way we like to do things. Yeah, very good. Yeah, fully related to that. I remember explicitly taking AI off the sort of the sneak AI-powered static analysis product, the slides of it because decreased trust amidst customers, even though it was actually doing good old-fashioned AI on it, and then needing to actually battle to get it back on to the slides, as it got the mojo once more. And when you talk to the folks that have been researching and building the technology for AI for the past 30 years, like they've been through cycles of winters, where things were super hot, and then, but very cold and nothing happened for ten years, there was a bit of a fear that after pre-training was picking up, like we would have a little bit of a slowdown, turns out RL decided picking up the slack pretty quickly, and there's another vector of scaling now that we're all excited about. We still have to connect those two ends, we still have to get to the point where most agents are good enough for most use cases, and that's not here yet. Yeah, and reliably, which is probably like a key question. Cool, I've got like a whole horde of other questions asked of you, but I think we're running out of time. Let me ask you two quick-wish questions. One is, as you build the AI-powered products from a engineer perspective, the people that you hire into your development teams, or our ND teams, building it, there are difference in profile, whatever you've seen, to be the people that are most successful, more adept at building those products. Yeah, there's not a big change in who you hire. I think, in general, we've always indexed most on the ability for people to grow. And I've seen that through my career before the drug and that drug, where when you hire people, usually they're super smart, but only a smaller number of people are super smart and willing to grow. And the difference here is, yes, I'm super smart, but if on my second day, somebody do something and somebody tells me, actually, you should try a different way, how about this? The people we want are the ones who are going to reply, you know what, let me try. As opposed to, no, I think I got it right. And very often, people who are super smart are their own worst enemies in that way, in that they know the smart, and they don't label, they don't grow. So I think that part is super similar. There's one new category of people we're hiring, though, which is that, last year, we started building an AI research team. We didn't have one before, like everything was applied, so we were building and everything into our products, but we didn't have a proper research team. And the change there has been that, if you compare to the way research was done 20 years ago, actually, I started my career in research. I brought me to the US with a job, IBM research in New York. It was cool at the time, IBM research. Maybe it is still today, but at the time, I remember it being a very high appeal job. Yes. It was before Google Research and before all of that stuff, but the other time, the old products were built from research that was more than 10 years old. And it was very difficult for anything to cross directly from research into products. I remember that's why I left. And today, if you compare to that time, you find into product research that is less than six months old, and that research for a large part is not ever published, which means there is tremendous value in fusing research and product development, and that's what we're doing today. Great. Yeah. Makes sense. And that is a brand new profile on it, and you've had to staff up. Where in the organization does it sit? It's under my co-founder and CTO, Alexy. So it breaks the regular sort of product division? Yes. But look, it's still research with the. There's one thing that's different about research is that you don't have a scrum every morning when you tell the researchers what to do. Research has to have a bit more independence. Great. And then, I guess, the last question is just, what surprised you the most as you've been building on AI, with AI over the last couple of years? But for good, for bad, what surprised you the most? The right of change is just what's the most surprising. You're in the same situation. Like, you're not new to this business. You've built companies before. Yeah. And I think we've never seen such a right of innovation and a right of change. My favorite joke on AI, and I get asked all the time how much time it's saving me, and my answer is always that I haven't reached break even yet. I'm still spending more time learning about it than to get me. And I expect that to last. I think it's changing so fast that it takes a lot of time to learn about it. And to level up. Yeah. Yeah. I fully relate. And I think it also speaks back to the competency that you want in the team, which is a team that is able to deal with a rate of change that was always good, but now it's probably on steroids. Well, this was excellent. A lot of just brilliant paths and very kind of promising paths of expansion also for data. So that's really great. Thanks a lot for coming on to the show. Thank you. It was fun. And thanks, everybody, for tuning in. And I hope you join us for the next one. Thanks for tuning in. Join us next time on the AI native death, brought to you by Tesson.