Today we’re breaking down Databricks, a $130B private company that helps companies collect, store, and process very large amounts of data, and then use that data to run analytics and train machine learning models.
Databricks sits in the middle of modern data systems, connecting raw data pipelines to the tools teams use to analyze information and build AI. If you’ve worked on large-scale data or AI projects, there’s a good chance Databricks was part of the stack, often operating behind the scenes.
My guest is Alan Tu, portfolio manager and analyst at WCM In...
Transcription
12192 Words, 68809 Characters
This episode is brought to you by Portrait. It's the AI research system that I used to prepare for today's episode and for all business breakdowns episodes. Portrait was built by former buy side investors and they understand great investing isn't just about having more information from low quality sources. It's about having the right information organized the right way. And if you listen to the show, you appreciate diligence consists of many things diving into the history of a business, framing the nuanced competitive dynamics, tracking key signposts around your thesis. And historically, that would take up material time that you do not have. But Portrait is basically like adding an army of analysts to your team. It's powered by an AI system specifically designed for investment research workflows. So you get nuanced idea generation. Portrait assesses the same types of qualitative attributes that we discuss on this show. And that can help identify businesses which fit your frameworks. Portrait also customizes research report generation. And I used Portrait to generate a primer and lay out bull bear cases ahead of today's episode to help frame the conversation. And third, there's intelligent thesis monitoring. And that's where Portrait assesses thousands of data points across value chains. Each day, extracting the insights, driving the business. Again, all this work would typically take hours and hours and hours. It's at your fingertips now. Visit PortraitResearch.com to start your free trial today. This is Business Breakdowns. Business breakdowns is a series of conversations with investors and operators diving deep into a single business. For each business, we explore its history, its business model, its competitive advantages, and what makes it tick. We believe every business has lessons and secrets that investors and operators can learn from. And we are here to bring them to you. To find more episodes of Breakdowns, check out JoinColosses.com. All opinions expressed by hosts and podcast guests are soleater-owned opinions. Hosts, podcast guests, their employers, or affiliates, may maintain positions in the securities discussed in this podcast. This podcast is for informational purposes only, and should not be relied upon as a basis for investment decisions. This is Matt Russell, and today we are breaking down Databricks. My guest is Alan II, Portfolio Manager and Analyst at WCM Investment Management. And you may actually remember, we broke down Databricks with a different guest about three years ago, but given how much has changed in this business and the subsequent capital raises, I was personally interested in revisiting this story. And I had a conversation with Alan about nine months ago, WCM had invested in Databricks in December of 2024, and I was curious just to get a better understanding of what they saw in the business. And we got into that in that private conversation, and then really focused on it in this conversation. So we start with what exactly Databricks does for its customers, and I think this is the large private company that might be least understood by the general public. And perhaps that dates back to the unique founding team and origin story, which differentiate Databricks and probably play a role in terms of how it's evolved from a successful initial product into the commercial platform that it is today. Now this conversation was recorded on December 10th of 2025, so all numbers are reflective of what was publicly available on that date. And please enjoy my conversation with Alan II on Databricks. All right, Alan, I am pumped to have you here to talk Databricks. It's rare that we go into the private sphere, but there are certain companies that are 100 percent worth analyzing in this space. Databricks being one of them. And what I would say differentiates Databricks versus a stripe or a SpaceX or an open AI is that most people understand what those businesses do. I think Databricks is a little bit more of a mystery to most people who are not close to the business, have not invested in the business. So if you could just start off with the simplest explanation you could give in terms of what Databricks actually does. Totally. I think part of the challenge with Databricks is it's actually they address so many different use cases and you may hear folks talk about, oh, we use Databricks for recommending movies or pricing strategy or fraud detection. And it's like, well, these are not necessarily super related use cases, but it sounds like Databricks is very critical for all of these use cases. And then you go a later deeper and you ask, well, what exactly do they do? And you'll hear folks say, well, they process the data. Which then, for me, leads to the question of what does processing the data mean. For someone like myself who I do not have a technical background, even that it can be hard to grok. And so the example that resonates for me is that I think we've all had the experience of getting a data file, a spreadsheet in Excel and you're interested in running an analysis based on that data. And it might be a very simple analysis. It could be something as simple as I want to understand the average price of a bunch of different items that were sold. But I think we've all had the experience of, well, you get the data in the spreadsheet, but it's not perfectly set up. It's not every column is exactly where the price should be. And perhaps in certain cells, the prices in a certain currency in another cells, maybe it's even the price written out with text. So you can't just select all that data and just say, what's the average? And so you actually end up spending the majority of your time, sometimes 80 to 90 percent of your time, just going through the process of unifying all the data into the same format so that you can run that very, very simple calculation of what is the average. So to me, that pain point of actually getting data into a format that allows you to ask even a simple question, is this idea of data processing? And now in the case of Databricks, just think of that at a completely different scale. You're talking about tons of different types of data sources. There's this concept of unstructured versus structured data, anything that's rows and columns that fits in a spreadsheet is more structured data. But the reality is most of the data out there is unstructured. It could be log files that are big streams of text. It could be image or video for folks that are analyzing websites. It could be click stream data. You take that problem of a lot of different types of data formats and how do you get them into a state where you can actually run analysis against it? That's how I would think about the core of what Databricks does. Now they've since expanded and they do all kinds of different things. But that data processing concept is what underpins the primary pain point. And then you can tie that back to all these different use cases. If you think about an e-commerce company that is trying to think about how much inventory should they stock of a particular skew or a t-shirt, whatever it might be, you could imagine that there's a lot of different inputs that might help you make that decision. It could range from how is your digital advertising performing with that particular t-shirt. It could be how has competitive type of t-shirts been selling. It could be credit card data. It could be all kinds of different data. And if you could get access to that data, you could potentially put that into a process of creating a model to answer that question of how many t-shirts should we keep in stock. A lot of sense. I think the painting of the picture of the Excel model certainly will resonate with pretty much anyone. And you can talk to people about how simply taking whether it's unstructured or just not in the proper format of data ends up being 95% of the work for what you're doing. And oftentimes I think that can create this environment where you have the idea of doing more analysis or running a more complicated model. But just that workload up front of doing that stops you and limits ultimately what you're doing. So I think that actually does bring it to life in a really thoughtful way. I want to go back to the beginning here because a lot of what you mentioned ties back into some of the unique origins of Databricks. But I wanted to just start with the founding team, the academic. And there's all these clichés about academics and them not being commercial. And here you have this really impressive story of evolution. So can you bring us back to the beginning stage of Databricks, who it was, what it looked like in those beginnings and tell a bit of the story which I think is really, really interesting for this business. I think it's a very unique part of the Databricks story that to this day the culture and the DNA of the organization, you can trace a lot of the decisions back to this founding story. And so it's seven founders that came out of Berkeley. They were all working in what's called the Amplab at Berkeley around the 2009 time period. And if you roll back the clock to that time period, that was actually the beginning stages of the cloud. And what Ali, the CEO and one of the co-founders would say is that the seven of them were working in this building together, doing research. And actually in the floor below was another team that was putting out some of the very early research around the data center being the next computer, basically the early concept around cloud computing. And so you had, in one part of the building, a lot of innovation around just cloud computing at the hardware layer. And then you had Ali and his colleagues that were thinking about what are some of the software opportunities in cloud. They were so close to the research around the early ages, and unlike with AI today, where there's been a very clear recognition that AI is going to be a big deal. Back then, the idea of cloud was still somewhat controversial. >> To imagine, but yes, totally. This group of folks that were rooted in research gained conviction around the idea of cloud. So then they thought about, well, what are the problems that we should work on within cloud? They actually thought of a few different ideas, but what they ended up was around this idea that data is going to be a really big problem. Data at scale. And that if you think about all the different use cases going back to the beginning of conversation, there's infinite number of applications around data. The other interesting thing that Ali will say is that during that time period, it was also when Twitter was becoming big, and Airbnb was becoming big, Facebook was still becoming big, it was actually a very positive time in technology. And so there was a lot of optimism around entrepreneurism and starting startups. And so that also fed into the energy of the group. The thought was, well, one of the co-founders was actually the creator of Apache Spark. They believed that data was going to be an important thing. And so how do we create a business around that? And the other piece that they thought of was using open source as another key bet. And that really was very aligned with this idea of coming from academia and research. When you roll back the clock, there was really three major bets that they had a view on. It was, cloud was going to be big, data was going to be big, and open source was going to be a good way to build a business. On hindsight, it turned out all three of those bets were very good bets. But the idea that at the time was less clear. But because of that environment where there was a lot of optimism around technology, I think all of that coalesced together to be the beginnings of Databricks. On the point of connecting cloud to data, is it fair to assume that by transitioning to cloud, there would actually be more capacity for data to be stored or to be used? Is that connected in the sense that prior to cloud, the data capabilities might have been constrained? Yeah. One of the paradigm shifts of cloud was generally this concept of scale out architecture, which basically allowed the ability to use more commodity hardware to be able to address larger amounts of both compute and storage and data processing in this case. That was an important underlying trend that enabled this proliferation and this idea of just data explosion that I think you're touching on. When you look at what Apache Spark was, it was leveraging this concept of distributed compute and applying that to data processing that was very important. As I asked, it's always interesting to think about second order impacts or a certain market enabling another market on the back of it, particularly with AI today, everyone looking for second order impacts. Totally. It's interesting to hear how they're connected and to your point, those three ideas, those three bets they made, certainly together, compounded in many ways. On that point of the open source, there are all different ways to approach open source and we have walled gardens versus open source and there's the very famous Apple versus window example or Apple versus Microsoft. Can you talk about the commercialization and how that played into things because I know that was a major stepping stone for the business in terms of evolving from this tool that gained usership oftentimes, offering things for free is a great way to get people in. It has to be great, so it certainly was. But that feels like such a meaningful point in time for this business and I'm curious how you describe what happened and how they really turned this into a great product and evolved it into a great business. There are a lot of examples of successful open source projects and that's one of the magical things of open source is that it does enable a level of adoption of in the consumer world. It's a capturing magic in a bottle and enterprise rolling with technology, open source is a really good way of getting just massive amounts of awareness, mind share and adoption. There are a lot of examples of successful open source technologies but there are actually not very many examples of successful businesses that have been built on top of open source and a lot has been made about Red Hat was one of the first companies that built a business on top of Linux, but the reality is that it's actually one of the hardest things to do. You have to hit two home runs. This is actually one of the ways that Ali explains it is you need to hit the first home run which is to develop an open source technology that gets mainstream adoption but what a lot of folks don't think about is the second home run which is how do you actually build a business on top of it and part of the problem is that the open source technology ends up becoming one of the businesses main competitors because anyone can get the product for free, anyone including your customers but also including competitors, they can also leverage that technology and actually those competitors that have more distribution, more customer relationships can actually do a better job of monetizing that open source technology. This is where again that coming from academia I think really helped inform the strategy for Databricks which is the benefit of not being immersed in the commercial market is that you don't have the pre-existing notions of what you should or should not do when it comes to building a business on top of open source because again the biggest example before that was Red Hat which basically provided services and support for Linux but that was the main thing but because the Databricks team wasn't really aware of the precedence they thought about things in a very first principles way and when you think about this challenge of how do you monetize something where there's a free alternative it's actually a simple answer and the answer is you need to create a better product that is worth paying for. Simple answer maybe not simple execution. Part of the reason why it's not simple execution is that if you've built your brand and got a lot of positive feedback for successfully creating the open source technology it can be weird to then say I'm actually going to create a competing product and I'm not going to put all of the bows and whistles into that open source technology. And for a lot of folks that creates a lot of tension you've got a lot of folks in the community that are like how could you do that there's almost a feeling of betrayal. When you're successful with open source it can almost be a curse because you become very popular you're viewed by technologist as someone who's brought this great thing into the world and then all of a sudden you need to be willing to be a villain. I think a lot of people have trouble making that jump but I think again in the case of Databricks they just very much realize that there's no way we can compete unless we have differentiation. So creating that differentiation was just a very important concept and so one of the things they did with Databricks the company and we can also talk about how they chose the name Databricks was that they created a new implementation of Spark that was completely proprietary that had a lot more performance and things that enterprises want reliability, et cetera and they were just very unabashed about the fact that if you want to use this new implementation you would have to pay for it. Is that comparable to the free tier of an LLM versus the pro tier of an LLM today just in terms of you can feel the difference and not everybody will have used the different tiers but I'm trying to grasp the differentiation to your point needs to be there. It's very easy to describe it when you're signing up for a subscription to a website and it's like you get three free articles but then after that you have access to our proprietary database for the subscription cost. What comparison would you make if there is an analogy for the difference? I actually think that analogy is pretty accurate but there are nuances to that. As a consumer we are used to this construct of there's a premium version and then we pay for the premium version and I think that decision about what goes into the premium version is where the nuances. Again going back to enterprise the typical traditional wisdom is oh well the average developer that's working out their garage we like that they can use this technology for free. We really want to monetize the enterprises and so why don't we monetize the enterprise features. Things like single sign on and again governance and security that actually makes sense but the reality is that while it is true the enterprises will pay for some of those additional features how much will they pay for those additional features and at the end of the day those additional features are not the core product. So going back to your example of the LLM there are different ways of deciding how to pay while a product. For certain premium products you can use it ten times but then the 11th time you have to pay or in other situations there are extra features that you have to pay for and then you have to go to the premium tier. The closest analogy in the case of Databricks would be actually the better model that is smarter and will give you better answers you do have to pay for. And so I actually think that comparison is not a bad one when you think about it through the lens of that core performance of the model not just these ancillary things that you're paying for. That differentiation of it's not added features. The core thing is actually higher quality is a meaningful differentiator and this is something to your point that every business has to think through and I talked to many people who want to roll out a premium tier of something and extra features are just not that valuable and nobody is going to pay for it. I did want to touch on Databricks the meaning. What's the origin story to the name? To draw contrast there are a lot of examples of companies that have been formed on the back of a successful open source technology that have basically named their companies the technology. So whether it's Docker was a startup that was built to commercialize Docker the technology you can think of MongoDB is another example. In the case of Databricks the analogous thing to do would have been to name the company Spark. The reality is that there are a lot of benefits to that because again Spark was a very well appreciated name the brand and awareness that Databricks the company would have gotten from coming out with the name Spark would have been very beneficial actually. But the reason why they went with Databricks is because from day one they always felt like it was going to be more than just Spark. Databricks the way they thought about it was that they were going to be many many bricks that you could all be used to apply around this broader problem around data. A simple idea but underneath the decision to name Databricks as opposed to Spark is this reflection of this long-termism of thinking about what is actually going to set us up to become much more over time. I often think about that when companies evolve there's an enterprise value of a brand that has multiple different products and you could stick with that core product or you can evolve above it and it's a little artsy in terms of the way I describe it but I do think there's something representative of that and I think it alludes to one of the points that you made early on which I think is a good thing to bring in now which is they've evolved in terms of what they offer and they have many different things under the hood now and I think this coincides with you and WCM's involvement in the business so can you talk a little bit about the evolution of what Databricks offers to customers and how it's evolved past that original state. Databricks today has truly reached that point of platform. There's sometimes an easy framework of future product platform with an enterprise software and I think Databricks has made strides over the years along that journey immediately following the success of commercializing Spark what the Databricks team did a really good job of after that was recognizing hey who are we serving in the enterprise it's the data engineer and data scientists these are folks that are after they process the data they're actually building machine learning models to run some of these predictions and forecasts and recommendation engines all those use cases we talked about and there's actually a very complex tool chain to enable that process that Databricks very naturally extended to and one of the products that they came out with which ended up being open source again was called ML flow and so this was another product that just extended the value proposition along the same use case for these data engineers and data scientists they then came out with another product called Delta which was a first step towards data warehouses which will get into the convergence between Databricks and Snowflake but Delta was another step in the direction of saying okay a lot of these use cases for machine learning are for advanced scale out use cases but they don't necessarily need to have the same level performance that allows for what's called transactional use cases that was actually an important next product as well the transactional use cases to set up to do with speed speed is part of it there's a concept in the database world called acid is an acronym this ACID the actual letter stand for atemicity consistency isolation and durability which is a long way of saying that there are certain workloads that require a certain level of guarantees around the quality and integrity of the data so when for a lot of traditional analytical workloads for example just taking it back to a data scientist that's running an analysis about how many shirts we should stock in inventory it actually isn't that important that all the data underneath it is perfectly in sync such that if one of the data sources is tweaked by a little bit that totally throw off the analysis but there are other workloads for which you need to have that guarantee that there are no inconsistencies in the data that there isn't one person is changing a piece of data here which has follow on effects there so that is another segment of the market when you think about the types of workloads that databases can address that was an important extension and evolution of Databricks and so they then came out with a product called Delta which was a first step towards addressing these acid requirements and one other really interesting thing about Databricks is that I think one of their core competencies is marketing there's a funny story where when they came out with Delta in order to explain to folks what Delta was they gave out free t-shirts that said Delta is spark on acid that's amazing so that was another again extension of Databricks introducing another product logical extension of their existing products but also commercializing it well and understanding how to market the product to the broader community and then from there to your question there in my view was a very pivotal moment a couple years ago when we got involved at WCM and investing Databricks which was when following Delta Databricks started to show the ability to address traditional data warehouse workloads which is the end part of the acid journey that we discussed and that was very important because up until that point you could largely say that the products that Databricks had built were all addressing again that core persona of the data engineer and the data scientist but the types of folks that actually engage with the data warehouse are more traditional data analysts so these are folks that typically use SQL to run queries against more structured data versus data scientists running Python and building machine learning models against unstructured data but because Databricks had built the foundation and logically laddered their way to the data warehouse they were then able to come out with a SQL product that was more directly competitive with one of their peers public company snowflake. That was to me an incredible proof point because they already deserved a lot of credit for expanding from a single product to multi product but then to expand further to multi persona for lack of a better term was just a tremendous tam expansion in its own right but also a demonstration that there's so much that goes on underneath that to enable the success of a product like that and so that was around a couple years ago they introduced their data warehouse product earlier this year they announced it's on pace to be a billion dollars in revenue which is just an incredible amount of scale for a new product and that to me really started to demonstrate this idea of Databricks this successfully becoming a true platform. Yeah, it's very interesting to hear and we had discussions before this conversation about your involvement in the business and I'm very interested in late stage investors and private businesses and what insights they glean and when you painted that picture to me of that evolution into a true platform it checked out to me just in terms of okay this was a unique moment they've evolved on that point about competition at the highest level. So I'm an organization using Databricks am I potentially also using snowflake am I using multiples is it not necessarily winner takes all market but how does that work just in terms of how much dominance there is with customers when they choose one versus the other. The reality of the market is that there tends to be multiple vendors that enterprises will use and actually I think Databricks is contributed towards a trend of enabling more types of tools for more types of workloads now Databricks we believe will then be able to come out with more products to address those different workloads but to answer your question it is very much the case that you see customers using both Databricks and snowflake for example. If you roll back the clock to that core use case that Databricks addressed early on around data processing it's actually a very classic situation where an enterprise might use Databricks first to process the data and then store that data in a snowflake data warehouse snowflake has started to try to move upstream to do more data processing and Databricks has moved downstream to do more of the data warehouse but that can give you a sense of the way that these tools can live together within the same company. I will acknowledge this is very easy for me to say from the cheap seats but it would seem as though moving from the processing the unstructured world into the warehousing more structured world might be a smoother evolution than vice versa because in my mind that unstructured world is a very complicated solution that was being offered and not the data warehousing is not but you think that's a fair representation and definitely feel free to push back on that assumption. Well I would say that empirically if you look at the numbers that has played out to be true with the data warehouse products scaling to a billion dollars that dwarfed the analogous revenues the snowflake has had around moving to data engineering but what I would say is that not being able to run an AB test in different versions of the world I would say that there was a lot of Databricks execution that led to them having more success moving towards structured. We touched on it briefly earlier but I thought this was another example of the company being not just great technologists but actually really good savvy marketers and folks with a commercial gut instinct when they recognize that they wanted to move into structured and there was this concept that unstructured that what people would call kind of that world would be data lakes and then structured would be data warehouses and so Databricks actually came up with this terminology called the lake house combining the data lake with the data warehouse and at the time you can go back to some of the news coverage there was quite a bit of ridicule about this idea of lake house it was almost to clever this idea oh you're going to combine the two and you've come up with this name fast forward to today and the lake house is a very real defined category that industry observers have all coalesced around so the credit the Databricks deserves not only for executing on the product and the technology to achieve again that we've talked about with data warehouses but to then do all the work to educate the market around why the lake house architecture is the best of all worlds and why that is the future is an incredible piece of the story that I think Databricks probably doesn't get enough credit for yeah I think those things are very hard to measure but you certainly can appreciate them sometimes more after the fact and I certainly just give companies bonus points when they're having fun while doing this execution there's something that just seems to matter to me it shows a willingness to enjoy the aspect of business and competition and whatnot there's a certain amount of fun and I don't know if they would use these words but I feel like irreverence in terms of I think this ties back to the founding heritage in DNA where it's look let's have an opinion about where the world is going I think as an investor if you go back to the early bets they would tell you these are the three bets that we're making we're making a bet that cloud will be big data will be big open source will be a good way to build a business or at least build an option and here it was we think that the lake house will be big and we think that this is where the world should go we think that this will help customers and we are going to bet behind that it just makes it very clear that if you're betting on data bricks you're betting on this future state of the world and I find that sometimes with companies that try to have it all they say well we'll be good here we'll be good there but the reality is that that detracts from your ability to execute in a way that data ricks did with something like the lake house I think things like first principles can get thrown around a lot but in preparation for this watching reading a lot of things that Ali has done in terms of interviews it checks out in terms of the approach and there's some certain clarity to the academic world and being born out of that in terms of understanding exactly what you're doing that focus having that clarity in terms of why you're going after things which I think can sometimes get drown out by some of the other baggage that comes with academia but putting that aside well I would also say that I guess a different way of describing this is that data bricks is helping to lead the industry to where they think the industry should go they identify where the pain points are and they come up with a solution that they think makes sense as opposed to looking at an existing market and just saying okay well we can do a me to product just for the sake of expanding our tamp what I've found is one of the underlying things that day breaks does a really good job of is recognizing true value creation as opposed to just monetizing and revenue growth if that makes sense it's based on customer challenges but also having a predictive view in terms of what's going to happen in the future there's a little Steve Jobs thing about designing for what the customer doesn't necessarily know what they need yet in there not to draw too hard on analogies on the point on market expansion tams all of that when it comes to both data bricks and snowflake both are relatively young businesses were they replacing industry incumbents was it all new market creation how would you describe the tam that exists relative to what it was whether it's prior to cloud or even within the cloud and using some of the incumbents going back to the starting points of data bricks and snowflakes snowflake was really the next gen cloud version of the data warehouse which was a market that did exist in the case of data breaks it would have been that data lake market but that had never been as well established for basically lack of good enough technology there had been a lot of attempts at creating data lake companies there was a technology that predated spark called Hadoop and there were companies that were built on top of Hadoop like Claudera that went public at one point but the problem there was just that simplistically the technology wasn't good enough so to your question of tam I think the market technically existed but there was this period people forget in the early 2010s when big data was a very sexy topic and there was a period where companies had the recognition or at least the inkling that data was valuable there was a whole period of a number of years where companies were storing a lot of data volumes and volumes of data with this underlying view that we should be doing that because big data why not but the reality was that there was in gardeners term a very hard trough of disillusionment where it was okay now we store all this data what are we going to do with it and turns out it's very difficult to get anything out of it and so that's what again going back to data ricks is core value proposition it was really solving that problem and that was massively tam expansionary yeah that makes a lot of sense and data being the new oil I think there's certainly some truth to it but there's also this massive challenge of understanding that this probably has value but how do we unlock that value and do something with it and that being a problem they solved is quite interesting I want to get a little granular just in terms of a use case so I truly understand what is going on to the extent that you can answer this one of the examples that I saw presented is I make a credit card transaction all of this data is flowing through the pipes of my credit card company maybe my bank data bricks is involved in that and if they see I make a transaction with a non traditional vendor it's for a very large amount and it's in a country that I never make transactions with before I can get a fraud alert and for my understanding is that flows through the data bricks pipes to some extent just in terms of managing all the variables of play that would cause a fraud detection in that example how does it work with data bricks actually having the ability to make the decision on the behalf of the credit card company to send me that fraud alert versus them presenting this alert back to the company and coming back to me it paints an interesting picture of how ingrained they are in terms of with their customers and I know it's going to differ by use case but can you talk a little bit about that sure for any given credit card company the implementation might be different than another but to bring your use case to life it is very accurate to think about that core value proposition that data bricks provides which is that you can imagine the amount of data that goes into making a decision of whether or not a transaction is worthy of a fraud alert there's a tremendous amount of input that can go into that and there's probably never enough you could always add more data to that analysis and that hits on the core thing that database provides which is the pipelines to bring in all that data process that data because all the different types of data are going to be in different formats and feed that into a machine learning model that gets fine tuned by the data scientists and tweaked and will constantly be updated based off of the facts on the ground and the empirical data also evaluating these models are they actually accurate after the fact and then tweaking those models again all of that is core data bricks value proposition now once data bricks helps a company come to that decision who is actually sending the fraud alert again there can be different architectures here but typically companies will build another application that actually takes the action and the model output from data bricks will inform that action is maybe a classic way to think about the architecture that makes sense the application that sits atop it's ultimately informed by the data bricks models that are analyzing all those things and I can imagine it being rules based if a transaction meets these criteria and it gets sent up then that triggers the fraud alert warning where I was going with that is I'm trying to grasp again with you'll have companies that are using multiple different vendors obviously there's so many different use cases the one I just brought up is a cost savings we've already talked about revenue growth and how this can be used but in terms of ways to measure stickiness and ramping with customers what does that look like it just assumes the more you would use the model the more ingrained it would be in your business therefore less likely to churn totally well they disclose that their net dollar expansion rates are greater than 140% so there is embedded in that high level of stickiness and also embedded growth so quantitatively that's the numbers that they've disclosed but I think qualitatively the right way to think about it is that many of these use cases that we've discussed are very core to the fundamental product that businesses are selling sometimes in the world of data analytics you can just envision a data scientist or data analyst in the back office running an analysis for the strategy team that may or may not feel sticky but when you think about these use cases where this is a content streamer that is suggesting the next movie you should watch after you finish a movie that is core to the product and that can oftentimes be revenue generating and very mission critical so there is that level of stickiness in terms of once you get embedded into use cases in production and then there is also the added layer of stickiness around just the fact that there's a concept of data gravity and once you put in the work to store and catalog data within a data platform that becomes very sticky so there's a lot of different dimensions with which data bricks becomes very embedded within a company I think the last thing I would mention is if you've done the work to process data once you could potentially use that for multiple use cases you can then again imagine how that becomes very hard for even if a certain product gets sunset but there's another product that's still leveraging the same data that's very sticky as well you've made the illusion and it's the elephant in every room now of AI I guess I'll just start with the very highest of level of what are the impacts of AI on a business like data bricks to the extent that they're beneficiaries to the extent that there's risks associated with it I'll let you wax poetic there. Well maybe just to start with some quantitative framing so data bricks they're now over four billion of AR they have disclosed that about a quarter of it one billion is AI related revenue and so AI has already become a very large part of the business but I think beneath that there are different ways to perhaps slice and dice the impact of AI on data bricks and I think for me one of the things that I've really liked about data bricks as an investment is that there are multiple ways to win starting with the core data processing piece there's almost a consensus understanding within enterprises now that you don't have an AI strategy without a data strategy. Everyone recognizes of course the model providers are doing what they're doing and every generation of models is getting smarter and smarter but at the end of the day if you don't have good, clean, well catalogged data the models can only do so much. One of the ways that AI has really benefited data bricks is business is that it's just created a tremendous amount of prioritization and awareness of the importance of the core product the data bricks is always provided. In my mind that is a durable tailwind that companies will need to do have always needed to do anyway that is actually not dependent on whether we achieve AGI or not or what does the next opening AI model do the reality is that as long as there is a general belief and understanding that AI is important there will be a driver towards more data engineering and data processing. That's a general tailwind for data bricks that I think when you think about data bricks the business as an investor it actually paints a picture of a more durable growth trajectory that is perhaps not as spiky on the upside but also not as volatile on the downside either in the case that sentiment around AI changes. That makes a lot of sense just in terms of a heuristic at a high level for thinking about AI within the business. Yeah and there's probably a couple other ways to think about AI's impact on data bricks. Another way is actually the fact that there is this huge cohort of AI native companies including the largest AI labs that can and do use data bricks as well for themselves internally. This is something that companies in the public market also talk about is that there's how is AI actually impacting product and use cases and there's also the fact that are you part of the stack that AI native companies are utilizing and data bricks is very much so there. Then I think the final part of AI's implication on data ricks is business is actually product. One of the big picture bets the data ricks is making going back to this idea that they have a DNA of having an opinion about where the world is going and where can data ricks add the most value. It's really around the idea that in the future AI and LOMs have already proven even if the models don't get any better than where they are today, the ability to automate more work. When you think about how big of a TAM that is, it's probably just as infinite as the TAM that we talked about initially around data. What data ricks is doing is building products that in the same way that they came out with ML flow that helped that whole process of a data scientist building a machine learning model. They're creating an entire stack with products called agent bricks and lake base that altogether will help enterprises build their own agentic applications to automate specific use cases and actually automate labor and work, which is just tremendous amounts of ROI. Yeah. I'll reframe it maybe and you can tell me how accurate this is versus inaccurate. First of what we talk about with agentic is oftentimes just shrinking the context, giving it very deep context in specific tasks and therefore the quality of the response is going to be much stronger. It's not going to be pulling random places on the webs. Data bricks obviously has the most rich data to use in terms of informing those models and therefore can partner with businesses who might want to develop those agents and set. Exactly. To your point, part of what the industry is realizing is that there are techniques that are important to be leveraged to build effective agentic applications. For example, RAG which you walk in generation, there's a whole set of processes around enabling that with vector databases. For example, embedding the data bricks as offerings there. There's also a very important part of building agentic applications is around model evaluation. Because of the very unpredictable nature of the large language models, it's not quite as easy to always know exactly is our application or agent acting the way that we think it should. There's a whole set of technologies around model evaluation, around being able to actually quantify how are these models behaving or they doing what we think that they should be doing. Data bricks again is building products around that in the same way that it's parallel going back to the machine learning era where you would go through a similar process around evaluating models there. You can get a sense of the fact that while there's a lot of focus on the core large language models, if you actually want to build applications and production, there's so much around and beyond just the model that is really where I think data bricks has a strong right to win. Yes. The infamous question of where the value might occur within the layers of the AI ecosystem is quite interesting. On the opposite side of the equation, there's been this interesting, solid good relationship with the cloud infrastructure providers, AWS's error. Is that change at all with AI because to your earlier point, the data clean up and all of that becomes even more important. I have personally been able to clean up unstructured data into structured data much more cleanly with AI, different story when it comes to me doing this at cloud scale. But what's your view on where that relationship, where the cloud providers haven't necessarily moved into this category? Does that change at all? Is that a risk? Is that something that you think about at all? It's funny. You mentioned that cloud providers haven't really moved into a market because they actually do have offerings there. I think it's actually more a reflection of the fact that data ricks has done such a good job of executing on product, but also just market positioning that you have now. But I would say that for as long as I've followed data bricks, and I first met Ali over 10 years ago, they had just signed their first strategic partnership with Microsoft that was actually Databricks branded under Azure as Azure Databricks. That was an incredibly important partnership to jumpstart Databricks's monetization. I bring that up because Databricks from day one, Ali as a business leader, has always been extremely pragmatic and strategic about how they operate vis-a-vis the hyperscalers. What I would say is that what has always been the case is that there is co-opitation with the hyperscalers. The reality is that for customers that are using Databricks, they will also be consuming infrastructure, computing storage of the hyperscaler. There is a benefit for the hyperscaler clouds when customers are using Databricks on top of their infrastructure, and that co-opitation dynamic I totally expect to continue in the world of AI. I do think it's a very important question because this is a big enough market that the hyperscalers do care about it, and that it's been another, I believe, underappreciated strength of Databricks, which is their ability to align themselves with the hyperscalers. There are a lot of examples of companies that came out with a great product, had a tremendous amount of momentum, and then Microsoft decided that this was too strategic for them to lose, and so they were going to put all of their weight behind killing that product. I think we're going to all think of different examples here, and that has often times been a real challenge for growth stage software companies. I think Databricks has just done a very good job of never positioning themselves in such a way that the hyperscalers are 100% incented to kill them. There's enough alignment, there's enough opportunity for partnership and mutual growth that that relationship with the hyperscalers has generally been relatively synergistic, despite the fact that they do represent very real competition. I bring it up all the time, but Amazon with something like FedEx and UPS, where they were relying on them, and then FedEx and UPS couldn't deliver during the holidays, and it was a big problem to the extent that Amazon built out a network, and then eventually started to compete with it, but there is something to the co-optition of high enough quality where it's not creating a problem. There are other dynamics that can involve, but that's very useful framing. I did want to get a little bit more into just some of the financial dynamics you mentioned, the 4 million ARR at this point. How does it work from a customer perspective? Is it a simple usage-based revenue model? It is, yeah. The way to think about it is that Databricks charges based on the actual compute that's being utilized for any of the workloads that are on Databricks. You can, again, go back to a tangible example, your credit card fraud example. Every single time the customer wants to run an analysis, and everything that happens underneath that in terms of the pipelines that pull in the data, that incurs compute cost, and that's how Databricks aligns itself from a modernization perspective. I would say that we've talked about this open-source piece, but Databricks has been really smart. Beyond just the core usage-based pricing about recognizing when certain features or products are strategic, but when they actually have a right to charge for those products. One of the examples is that Ali is given is that when you think about your smartphone, you have one of the features of the smartphone is your address book. The reality is the address book is an incredibly important feature. Not just for phone calling, but a lot expands from having an address book in your smartphone. The reality is that no handset maker is going to be able to charge for the address book. Databricks has a lot of different products that are analogous to the address book, where they are effectively, again, in many cases, using open-source to give it away for free in order to get adoption, but that are actually still very, very strategic. One example of that is one of their big value propositions is providing a governance layer on top of all the data, so that enterprises have a single pane of glass for which they can see all the metadata that they're processing. I guess all of this is a long way to say that while Databricks does use a usage-based pricing based on compute, the reality is that, at least from my perspective, they're actually monetizing more than just compute. It's a way of monetizing a lot more layers of value that Databricks is providing. And there's an intertwined nature to all the different things interacting with one another, which represents value, even if it's not directly correlated to the compute cost. On your point of when they realize something has so much value that it's something that should be charged for, does that come from it's burning a hole in their pocket from the compute cost of that's how much value it has, or is it more of a qualitative assessment? I can't speak necessarily for all the conversations that are had internally, but I think my view is that there are a lot of different inputs into every decision around how to monetize what to make open source, and sometimes it's more of a defensive stance of saying, "Look, we need to make sure that we get adoption of, for example, the governance product," or in other cases, it might be offensive, and they've actually also done a very good job of historically recognizing when they can be disruptive. This might get a little bit into some of the technical details, but one of the big things that Databricks did very effectively to compete against Snowflake was embracing open formats where they purposefully decided not to charge for storage, and that was something that Snowflake had historically done. That's another example of where strategically they're looking at a way to potentially be disruptive, not just from a pure lower cost perspective, but actually architecturally to enable customers to keep their storage, keep their data wherever they're storing it. Don't force them to put it into Databricks. You can actually just run Databricks on top of where the data already sits, as opposed to historically Snowflake was, you actually have to move all that data into Snowflake. That's another tangential point around this idea of making decisions of not only just went to charge and went not to charge, but also having a view strategically around, is this effective from a defense and offense perspective? Yeah, competitive forces that come into play. I guess on that point, when it comes to general pricing trends, aside from what they charge for, I can understand that it all gets blended together, but does the pricing trend tend to correlate to the cost of compute, or would you say that they're able to raise pricing or they're pricing more? I'm just curious, there's the quality of the product, which is going to make customers make that decision. Some customers might depend on or lean in on the price. What drives pricing changes? I would say that in this market, it is more about total cost of ownership relative to performance. What a lot of customers care about is, are we able to run our workloads in a performant way? Again, because it's not apples to apples in a lot of situations where there is the infrastructure layer that the hyperscalers monetize, if a vendor like Databricks can actually help customers run their infrastructure more effectively, that may not show up in Databricks' pricing, but from a customer perspective, that factors into total cost of ownership. A lot of that is a technology solution in terms of better understanding how certain workloads are behaving, how do you optimize the underlying infrastructure to serve those workloads? In my experience, it's getting workloads into production and really seeing what is that total cost of ownership to achieve the goal of any particular use case. The reality is that for a lot of these use cases, going back to the fact that Databricks oftentimes is embedded in the core products or perhaps the nature of the decision is extremely strategic. If you can effectively provide the end value, the ROI is typically very clear for customers. It's an interesting business in the sense that the use case on credit card fraud, you can actually draw a very clear ROI in terms of fraud is a major issue and at cost problem for credit card providers and being able to draw those connections, but there's other things that maybe are more difficult to draw direct ROI conclusions from, but equally as valuable. It's very interesting to hear how different ecosystems and value chains work when it comes to this type of business. When the cost side of the equation, cost of compute, which is theoretically passed through, if you're overhead, are there big buckets of cost that we didn't touch on that would be very important? Generally speaking, Databricks's model is fairly capitalite. They don't very much need to get into the GPU acquisition situation. Yeah, we're all very aware of these days. The reality is a lot of core data processing workloads are CPU-based. It is interesting that when you hear Jensen at Nvidia talk about where he sees a lot of value for GPUs in the future, he does have a view that more of these workloads will transition to GPUs, but the reality today is that Databricks's products are not compute intensive in the same way that you would think about a lot of AI native companies. Which is amazing to me, it just speaks to what training would actually require in terms of compute, but given how much they process, that's actually surprising to me, but interesting. It could evolve. One of the things that Databricks is having success monetizing is called model serving, which is basically when you go back to some of the examples we've talked about is this entire workflow around building the intelligence that underpins an application and you tie that into an LOM and you create what's called an endpoint that then exposes that intelligence to an application. More and more customers are asking Databricks to actually host that endpoint, which is basically like an API on behalf of the customer. In those situations, Databricks will actually have GPU costs underneath that, but again, relative to some of the other examples out there in terms of the scale of cost, it's a very different order magnitude. Not all GPUs are created equal, and that's a whole other topic. But I interrupted you, I think, on the questions of costs. Were there any other things that you would buck it in there that we didn't touch on? So Databricks has run in FreeCasualPositive at the scale that they are at, 4 billion plus an ARR. One could make the argument that FreeCasualPositive is a very low bar that perhaps they should be showing more profitability, but a big chunk of their costs is just around traditional software business model, it's investing in people, it's investing in R&D. It's really, I think, a lot of the scale software companies out there, one of the companies that, in my view, has a strong attract record of any in terms of demonstrating ROI against organic innovation. So from a cost structure perspective, nothing dramatic to call out other than the fact that they are still investing very aggressively behind R&D. And that is a big reason why they've been able to maintain their pace of innovation even as they've expanded to so many new products and areas. I do have to ask, it's capital-like, FreeCasualPositive, they've done a lot of fundraising over the years. Where does that capital go? So this is actually a unique dynamic that is not necessarily specific to Databricks, but it is more of a dynamic that I think more folks are aware of, which is some of these tier one high quality private assets are just staying private longer and longer. There's a dynamic where once you reach a certain level of scale, there is a certain expectation for perhaps early investors or oftentimes more importantly employees to be able to get liquidity for their options or RSUs. In the case of a lot of these companies including Databricks, oftentimes the reason why they need to do these big fund raises is a tax consideration for the fact that once you've provided employees enough opportunities to get liquidity, the IRS starts treating those RSUs and options as taxable. Historically within startups one of the benefits and options was that I was deferred tax compensation, but this is something that I think the industry has started to learn is that there's a real tax bill for basically compensating employees via equity and so to answer your question, the majority of the proceeds from the fund raises of Databricks have effectively been used to offset the employee stock compensation and the corresponding tax bill associated with it. I understand it correctly, it's not necessarily pure secondary in nature where the employees are caching out, maybe that is to some extent, but it is when I get my equity grant, I ultimately only end up getting 66 of the 100 shares because 34 are used to pay the taxes associated with that compensation is exactly that makes sense. And on the point of more private companies are staying private for longer, it's been a very interesting dynamic as a former public markets guy. I can actually understand some of those challenges, but what would you say just from your seat? It's very interesting, you'd sit at the crux of all of this. Do you think it's a when, not if, about going public or what would you expect that catalyst to be to, whether it's specific to Databricks or you can talk about industry as a whole? It's literally a trillion dollar question now with some of these assets and the scale that they've gotten to in terms of both the business and valuation. I think that for, again, that tier one list of privates, there's increasingly a term that we all know of the max seven in the public markets, but there's effectively a max seven in the private markets. You've now seen that there is enough infrastructure in place from a fundraising perspective, both in terms of capital availability, but also just the process of doing these very large scale fund raises at late-stage growth that has basically made it pretty easy for if you're at a certain level of quality to stay private. Going public becomes less about the need to access public markets and more of a discretionary decision of what are the pros and cons. For each business, there's going to be different decision around that. I think in the case of Databricks, I think it's fair to say that because they were private in the 22 cycle and growth tech, they were able to continue to play offense in a way that a lot of their public peers were unable to. It really helped to accelerate their business for a bunch of different reasons. You could imagine, but generally speaking, the ability to continue to invest both behind sales and R&D was something that really benefited Databricks. I can't speak for them, but I would imagine that that was an informative experience where the thinking would be that the reasons to go public would have to be sufficiently high to overcome the benefit that they've already experienced of staying private. It's a very interesting dynamic and I think you put it incredibly well there with a trillion dollar question. So it's going to be one that's interesting to watch a lot of headlines out there this week about what could happen next year, but it's almost like a believe it when you see it type environment now. Totally. I think we talked a lot about what has gone right and the opportunity set ahead, what stands out to you from a risk perspective? I think we've glossed on some of those, but when you think about risks for the business, is there anything that pops out the most to you? It is dynamic enough market where pace of innovation is still very important that it can't be taken for granted that continued R&D execution sustains. We've seen examples in the space of companies that have perhaps taken their eye off of the ball. They've been slow to ship certain products and how that can really show up in the numbers. So I think first and foremost, it's easy to say execution, but I do think in the case of Databricks, that continued execution at scale is important because they are doing so many different things. How they execute around the newer AI products is going to be important over time, maybe not even necessarily over the next two to three years because again of that dynamic that we talked about earlier where the core data processing tailwinds are so strong. What I do think for the ambitions that Databricks has, it will be important to execute on some of the AI products in the same way that they executed on the lake house. We are in a very similar dynamic right now where there's a question of category creation. What exactly does an authentic application product portfolio look like and what do you even call that and how does the industry coalesce around that set of tools and products? I think is not just a question of product execution, but again, that marketing and commercialization DNA. In the case of Databricks, time and time again, what I've seen is that the way that they make decisions has been with a very long-term mentality in mind. Go back to the decision around how they named the company. There are always these trade-offs where if you have a shorter time horizon in mind, you might make a certain decision. You might not open source something. You might be tempted to monetize whatever feature you just came out with. I think the ability for Databricks to continue to maintain that DNA is going to be really important because especially as we enter AI, every single time you make more of a short-term oriented decision that inevitably opens you up to some sort of vulnerability down the road. I think it'll be really important from a culture perspective, which is something that we spend a lot of time focusing on, that they maintain that core culture of being long-term, but also again, that founding academia-based DNA of thinking first principles, it's easier said than done to just say, "Oh, keep doing that." For us, that's an important thing to say on top of. Yeah, it might tailor back to your original answer on the staying private versus public. Having that long-term mentality is a little bit easier to do in the private markets, I think, than the public markets, which is notable and at such an interesting point in time with AI, it's all very interesting. We close these conversations out with lessons and I think you might have tapped into some at the end of your answer there, but what would you say are lessons that you can take away from Databricks and in the spirit of whether it's pattern recognition or anything else? What lessons can you take away from Databricks and investing in that business? Honestly, it would be just reiterating the last point around the long-termism. I think that's something that a lot of investors and founders talk about, but I think in the case of Databricks, it's being able to actually point to so many specific examples where certain decisions are made where there is that clear trade-off. There are certain times when people talk about being long-term where it isn't clear what the trade-off is. I find that the way that the Databricks team is able to talk about the bets that they're making going all the way back to that original founding view of those three bets of where the world is going and operating in accordance with the level of consistency is something that really stands out about Databricks. We talked about how there were ways that they could have monetized sooner, even going back to the cloud example where it wasn't entirely clear that the industry was all in on cloud, but they were so convicted that cloud was a real thing that they never came out with an on-prem version of their product. You can imagine, again, that there were a lot of examples where they could have that ended up leading to less monetization in the near-term than they might have otherwise had. They had this very clear view of where the world was going. If you just go back through the number of examples we've touched on, you can just point to a lot of examples of maintaining that long-termism and recognizing why certain decisions are made. I've found that following Databricks, that's been something that I've started looking for in other companies. I think that's a very interesting point on the long-termism, but understanding what the actual trade-off is, it's so easy to gloss over to that second point. Very interesting. This has been a pleasure, Alan. Thank you for educating me on something that I only knew a very surface level amount of material on. It's been a pleasure. No, this has been great, thank you. This episode is produced in collaboration with WCM Investment Management, this discussion reflects WCM views as of the recording date December 11, 2025, and should not be considered a current investment advice or a recommendation to invest in Databricks or any other security. WCM has a financial interest in Databricks, which creates an inherent bias in this discussion. For additional disclosures, visit WCMinvest.com.
Podcast Summary
Key Points:
Summary:
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.