Go back

Building the Context Flywheel for AI Data Agents

60m 34s

Building the Context Flywheel for AI Data Agents

Prapadba Senkar of ATLIN argues that the key to making AI useful in enterprises is not just intelligence but contextual intelligence—the institutional, semantic, and process knowledge that humans acquire through onboarding and experience. He notes a stark gap: while AI intelligence has grown exponentially, most CEOs see no financial benefit, and few AI projects reach production. This is because AI agents lack the "context" that humans absorb over months—such as how "new customer" is defined differently by sales, marketing, or finance. ATLIN addresses this by building a "context flywheel" that bootstraps knowledge from existing business systems (e.g., Salesforce, warehouses, BI tools) into a context lake house infrastructure. This allows AI to reverse-engineer metrics, ontologies, and relationships from usage patterns and SQL queries. The system achieves high accuracy (89% for column descriptions) and uses "context tree boughs" as portable units for agent workflows. AI does most of the work, asking humans only to resolve ambiguous definitions. This approach dramatically reduces the activation energy needed for metadata systems, overcoming the traditional barrier where such projects fail before reaching critical mass. Senkar believes this will lead to a renaissance in how companies structure and encode organizational knowledge for both humans and AI.

Transcription

9249 Words, 50639 Characters

English
[MUSIC] Hello and welcome to the Data Engineering Podcast, the show about modern data management. Your host is Tobias Macy and today I'm interviewing Prokala Senkar about strategies for building a context flywheel for your data agents. So Prokala, welcome back and for anybody who hasn't heard your previous appearances if you can just give a quick introduction. My name is Prapadba, I'm one of the founders of ATLIN. ATLIN, we are building a context lure for AI. I have been fascinated by the topic of how to build shared context for what I used to call the humans of data for now a better part of a decade. I started as a data practitioner myself than these large scale data projects realize that data engineers and less scientists, machine learning folks, business, business, business, even these people have their own version of context, which made answering a simple question like number on the dashboard is broken, we don't know why really complicated, like seven systems and five people were involved in answering that question. And ATLIN was founded on the mission of building a shared context and collaboration lure for the humans of data, which we have now realized as far more valuable in a world where AI trends are starting to take over some of those tasks, which largely humans used to be the glue for this context inside organizations and so on that journey for now, better part of this decade with our customers. And do you remember how you first got started working in the data space? Yeah, well I have been in the data space since I was 21 years old. I at the time was doing work in finance and data and I had done some work with nonprofits. I said these big problems like healthcare and education and infrastructure, they don't use data and it feels like they should and let's go do something about that. So, microfund over iron and I we founded a company called social cops that was found in the mission of bringing data science to real world critical problems and that meant at the time going in partner with agencies that already had reach and scale. So, it was large governments, it was the UN, the World Bank, the Gates Foundation and we used to be their data team, we used to build data platforms for these problems at scale and that is really well I learned everything that I learned about building and running data teams and how complex and chaotic it can get to actually make a data project come alive. And as you mentioned you started at land with this mission of collecting all of these disparate data systems and the metadata and semantics around that so that humans could make sense of all of the information that was available and how it was being used and how it can be used. Obviously that has had a dramatic shift in terms of the scale and scope and workflow around it with the introduction of all of these agentic capabilities, especially since the beginning of this year. I'm wondering if you can talk through some of the major notable changes that you've seen in terms of the roles and responsibilities for these data catalogs and metadata systems as agents become more of the actual operational substrate. What's interesting about the moment we are in AI on one hand, I learned San Francisco and I walk out of the streets and everyone's like "AGI is here" and you know every day the models keep getting better, the benchmarks keep getting better two years ago. It couldn't pass the bar. Today it's the top 1% of test scorers in the last decade intelligence has compounded 1,000X in the last six months intelligence has compounded 2X. So by any parameter exponential growth in every intelligence benchmark that you can have. On the other hand I work in the real world. I work with real companies and real enterprises and when I ask them how are things going? Is AI useful? Almost always, I drop, you know somebody talks about this one pilot that they're making its way to production. The data on this could not be more stark. 56% of CEOs report zero financial benefit with AI, you know one of five projects make it to production. And so there's a gap, whichever way you argue this AI being valuable, we've probably proven that. AI being useful, we haven't really proven that. And so what's the gap? And this is where I like to always go back to understanding humans and what it takes to make humans work before you figure out what it takes to make AI systems work. And hidden there is in plain sight as the answer. Companies in the real world care about performance. Performance is about driving outcomes in the real world. What is performance a function of? Performance is a function of I like to think intelligence but also context. If you think about human job performance less than 10% of human job performance is explained by cognitive intelligence. Which makes sense, right? Like would you say your best employees, also the person that scored the highest on the SATs? Or would you say they're the person that you know takes feedback the fastest and learns and grows and you know just you know absorbs institutional knowledge and you know works really hard. Like is that what the best person in your team looks like? And almost every time I say that people say yeah of course that's what the best person in my team looks like. For the last 100 years we have built infrastructure that helps humans become greater than jobs. When they join a company they are on boarded. They're on board. They're given a mentor. They go and you know attend and shadow other people in their team. They learn in these conversations. Their manager gives them feedback. They learn that way. They make mistakes on the job and then they learn through that. These are all the things that it takes for humans to get really knowledge and skills and expertise that it takes to get really good at that job. And that's what's missing. Today we have AI agents that are basically the smartest in turn we can ever hire but there is no way to really onboard them onto our context. And this is what I think of as an extra on here. I think the next frontier is contextual intelligence. It's like and I truly mean this. I don't think there is a concept of intelligence without it being contextual. It's like an academic construct. Just measuring cognitive intelligence that's not really useful in a real world setting. It doesn't really mean anything. And the next frontier is going to be about how do we actually take in harnesses and diligence by giving it the context that our companies have to make it useful. And to that point of onboarding and the idea that we've been doing this for a long time I think that the fact that we have to do it at such an accelerated pace and over and over again really stress tests the systems and protocols that we've had in place to do that onboarding where because we're working with humans who are able to do a lot of their own discovery and self learning in self direction or find the right person to go out and ask in physical space. We have been fairly laxed asical in terms of the level of rigor and detail that we put into that onboarding setup. And so there has been a lot of delay and pain for newcomers to a team to be able to actually get up to speed and be effective where it's not unheard of for somebody to say that I don't actually expect anybody to really be fully up to speed for six months after they start which when you really think about it is fairly insane. And so the fact that we have to do this over and over again and we're bootstrapping these agents from scratch every time forces us to be more deliberate in terms of ensuring that we actually have all of these details and protocols and information more explicitly codified for the agents to be able to bootstrap which also helps the humans who are going through that same process. Yeah absolutely. In some ways I think AI has shined the inefficiencies that exist in our human organizations that humans glue together and now we realize that they're missing in a very explicit way and that really I do believe there will be a renaissance of or like a new kind of learning organization you know like human like what does it take to build a really successful company in the AI world will be dependent on how companies find a way to actually build new ways of structuring and building these and encoding these in the real organizations etc. And for the context of these metadata layers and the semantic details that we bake into them I think another piece is that particularly for human analysts they're able to do a lot of that information gathering and there is that implicit knowledge, that institutional knowledge that gets built up over time where because they've had to run the same query 15 times or because they've had to ask the question of the sales team or the marketing team 15 different times to make sure that everybody's an agreement they have the necessary context to make sure that the queries that they are producing are accurate for the organizational context where an LLM might be able to fetch some of that detail and maybe go through query logs of queries that were generated by humans to be able to leave that information if you then persist that going forward, it makes everybody's job easier. So I'm curious how you're seeing the nature and shape of these context layers and the metadata systems evolve to more explicitly incorporate and surface those types of details that maybe we would have bothered investing the resources into engineering into those catalog systems. Yeah, absolutely. So let's unpack what a human analyst does today and what it takes for them to get really good at their jobs. Let's say you get a question and the question is something like tell me what my top 10 new customers are. Sounds like a really simple question. Actually, it's a really hard question to answer because the first question is like, well, who's asking? What decision are they looking to make? If sales are asking, then it might be the top 10 customers by revenue. If it is marketing, it's probably the top 10 customers by brand and logo refrensability. If it's product, it might be by adoption of the product. And so even before we get into in the data world, we talk a lot about semantics and how do we define a customer and how do we, there's actually institutional knowledge that if you're an analyst and someone's asking you a question from the marketing team, you already have some sense over this means. Then I use this word new in my conversation. What are the top 10 new customers? At Atlin, we run our business on a quarterly basis. So if someone says new customer, typically means a customer that signed in the last quarter. This is not documented anywhere. Nobody ever tells anybody this explicitly. You just learn it because you join the downhaul and you join the all hands. This just becomes how you think about the business. Then there is the what does a customer meet? And how do we define customer and how do we make that explicit? It's going to be different in finance and it's going to be different in sales depending on again, are you reporting with the word or are you running an acceleration program or a premium program or so on? And then eventually, we get down to the data itself and how do I actually measure what tables do I use, what descriptions do I use and so on? These are the four layers of first knowledge. This is the first piece that I think of. This is all knowledge. And there are different types of knowledge in here. There is institutional knowledge. There is semantic knowledge. There is process knowledge. There is user and knowledge. This is the first step of knowledge that the human analyst brings together. Then let's make this a little bit more interesting. Typically, analysts has not asked a question saying, what I might adopt a new customer. Typically, the analyst has asked a question that says, why might I have a new drop? Or why might I have a new go up? And this is where, for example, there is expertise that people learn. So for example, you learn that July is a seasonal quarter, especially if you do business in Europe because Europe tends to be on vacation in the summer. So the first thing that an analyst does when a revenue drops is that they first check Europe and then they say, okay, is this coming because of Europe and is there a drop just in Europe? Is it seasonal? Is it, if it's not, if it's seasonal, then this is nothing to worry about. So then they check, okay, what are my other regions doing? Have I had a drop in pipeline in other regions? And then they come back with, and this is expertise that humans learn over time in terms of the business and how the business operates. And that's the second layer of what people need, which is, how do I do a what the analysis in my business? How do I think about understanding the root cause analysis of certain things that happen in my business? How do I think about seasonality trends and so on? And then eventually then we have the tools. This is really where, how do I run, how do I go, run a query against my CRM? How do I bring a text to SQL together in the right way? This is what I think of as context. Context is a function of knowledge, skills and tools that all need to come together. The interesting thing about this problem is that we are now at a point where AI itself can bootstrap a large part of the knowledge inside organizations. So one of the things we have seen a dramatic success in the last roughly, at say three and a half months is we have unleashed essentially our context harness on the foundational data systems inside. There's already a lot of information about the business that's hidden in the business systems. If you connect the Salesforce to the warehouse, to the BI and how people are running queries on your BI and you're able to bring that back, you have end-to-end lineage and then you can start reverse constructing, how does this business run, how does this business work? How do you think about all of these questions that I already talked about? How do you reverse construct a metrics tree? How do you reverse construct a non-tology? Can you start thinking about seasonality trends in the right way? And so the first thing we've seen as an evolution as we think about the context clears, how do you bootstrap? And how do you bootstrap context from all the existing business systems that exist inside the organization? And the second thing that we see is how do we bring this into the development workflow of agents? And so we have now a construct we call a context tree boh. Think of it as a GitHub repo equivalent, but for enterprise context. It becomes the unit of portable context for a use case. And this is where essentially iteration loops, simulation loops, one of our core philosophies is AI itself should go do most of the work and ask humans for help. So can it read through for all of these use cases and say hey, you define your revenue differently in sales and finance? For this use case, which of the two should I pick? Human tell me and the human approves or rejects. So how do you build those kinds of loops and then turn this into a portable unit that then builds builds on the context over time? One of the major barriers to entry for adopting a metadata catalog or doing a master data management project or building a business glossary has been the activation energy required to get to the point where it's actually useful and worthwhile where for the first months or years depending on how the project is managed, you're spending a lot of time and energy and potentially money in the process of populating all of this information, but until you hit a particular critical mass where the majority of the information is present and correct and trustworthy, nobody's going to bother looking at it and a lot of times you never get past that point where you've actually reached that threshold and so it eventually will just quietly die off until the next time around that somebody says oh, we really need this system and then you rinse and repeat and so I'm curious how you're seeing these agentic capabilities reduce the amount of investment and time and activation energy required to get past that threshold where the metadata system or the business glossary or the master data management and golden records are to a level of completeness and quality that the business will actually rely on them. Yeah, absolutely. This has been I feel like a labor of love over the last many years for us at Adlin. So 2022, in the early 2023, we were the first company to launch Adlin AI in the category and we kept iterating on it and it was a human assisted workflow to go generate some of this context and we pushed the boundary as much as we could like, you know, honestly, we were able to get about 75% accuracy on this thing. It was the highest it could go and that's when we realized that infrastructure had never been built for a world where AI was the primary producer of context. Metadata infrastructure itself was not ready for that world. And so what we did was we actually rebuild the foundation of Adlin completely. So now Adlin has built on what we call a context lake house. It's an ice-burbonate of file format. It allows us to leverage the same compute that we would leverage on our data systems. And so all metadata and traversal that an agent on Adlin can do is able to leverage this foundation lake house infrastructure, which means that you can do compute level operations if you do one snowflake and data breaks on your data on your metadata. That was the first big architectural unlock that we needed to do to be able to get to them. And, you know, of course, as you do that, like how do you think about graphs and how do you think about relationships reversal and how do you think about, you know, vectorization and how do you think about, you know, all of these things that it takes for AI to be able to do work on your context landscape. That unlock was massive. Sometimes middle of last year was when we announced the context lake house and that was a massive unlock. The second thing that we did was we have done a very, very big investment into our own harness for context. And this is where we have realized context quality really compounds. So I'll give you an example. Honestly, you and I can today go take a table. We can go give it a cloud code and we can say describe the stable, give me descriptions and cloud code does a kind of like okay job, like decent job, but not really accurate enough for me to push it into production. But imagine you're able to take the best AI that you have an intelligence to have and you get it to go read your system of record. and where your data was actually created in your database. Then you get it to read downstream, how everybody is using it, how they're querying it, and what is the usage business patterns against it. You connect all of this together, and then you have AI write a column description. A column description is extremely accurate. Today we are at a point where our customers tell us, 89% of our customers say that it is as good or better than humans from a accuracy perspective. This was a tipping point, we got this tipping point, this quarter, with our context agents. Now if I have high quality column descriptions, I can do a really good job with domain tag, and I can start seeing, okay, these columns and tables belong to sales, these columns and tables belong to charge. Then I can read through the SQL, and then I can start reverse constructing metrics. And then the accuracy of that's pretty high. Once I have that, and then I have semantics from my BI, I can start reverse constructing relationships. And then I can start reverse creating a new and a new way of thinking about the data, and then I can start to look at the data, and then I can start to look at the data, and then goes and reads all the context about usage, and use cases, and it generates simulation environments. So the AI itself is going and running simulations on. Okay, well, Ravanyu, for the question, let me look at all the business dashboards that are already associated with this. Let me look at how these people are using it. What are the queries that run? Based on this, let me simulate, who are the personas that might be using this? Based on this, let me simulate, all the questions they might possibly ask from this agent. So again, this goes back to the context flywheel, because you have that base context, you're able to do a much better job to generate in these simulations. Then AI does the first job of basically say, okay, what do we already have context for? And what's the accuracy? It runs the text to SQL? What do we not have context for? So then what it does is it brings it to the humans, and then says, hey, here's the 10, 12 context pieces we need to add, that we need to improve context for. And this is how essentially our teams improve and run the simulation context engineering workflow to get it to production. Once you launch into production, one of our core thesis is that this is a living-beginting, this is a discipline. With AI, there is no such thing as GA. You are constantly improving your AI system. And so the last part of this, and we bring this into the context sheet for flow is we pull traces and we pull memory. And then we have AI again, we have a set of curation agents that are sitting on top of traces and e-vals. And basically bringing it back and saying, okay, here's how we should improve context over time for these use cases. So this is really where we have customers who would say things like, when we launch this agentic use case, we started with 50% autonomy. And now we're going to 90% autonomy. And the way we do that is by improving context over time for the system. - With the fact that you have humans who are collaborating and agents who are collaborating, what do you see as the value in identifying the originator of a piece of information? Is it implicitly more trustworthy if it's a human operator who is making a change or do people typically see that because the agent is going to be more tireless and maybe have more attention to detail that I should trust what the agent puts in and not the humans, or just some of the ways that you surface that level of detail as far as who provided this information at what point and what is the grounding for that level of detail. And then how you manage the change process for adding new details or adding new fields or new metadata. - Yeah. So I think the base starts at just identity and tracing of identity and being able to say, okay, this is something that was generated through an AI. This was generated through a human. I am increasingly beginning to believe that we are actually going to go from human in the loop to human on the loop. And what AI is able to do, humans are just going to be the blocker to AI going into production and leveraging the scale. Just like again, you know, we have AI within a week doing context generation for millions of assets. Like there's no way a human's going to be able to like go approved reject all of this, that's just not possible anymore in the real world. So one of my mental models has become how do you have AI do all of the work and bring it to humans for governance and decision making. So for example, one of our most popular agents is an agent we call a metrics conflict resolution agent. And what it does is it basically goes reads through all your code and your SQL and all this context that you have. And then it basically says, okay, hey, these are the metrics that you're literally defining differently in two parts of your organization. And that's a human decision. You know, like AI cannot meet that decision for you. And some human in the company needs to decide that when we are reporting to the streets today, we decide that we are going to report this versus this, right? And so that's a human decision and that needs human governance. But is this the most accurate description of this or is this how we're defining the metric? Is this the most accurate definition of it? I think AI can actually do a really good job in the right way. And like, obviously not if you go and just throw this into cloud and expect it to give you a result. But if you build a kind of harness or a kind of guardrails and checkpoints, or I kind of AI systems do it, I think AI can do it pretty. The second thing is traceability is really important. So one of the reasons actually for building the lake house in many ways was actually that versioning lifecycle management. All of these things that we've actually dealt with from a data world perspective will now start with coming very important from a context world perspective. So for example, agent made a decision. And it was a wrong decision. You will want to reverse state, right? You will want to go back to which version of the context that the agent used to make this decision. And so how do you build these same kinds of systems that we've had in the data management world, in the context management world? And how do you bring back a lot of those paradigms in the right way? I think will become even more critical. Again, it's a discipline. And how do we build the disciplines around who's allowed to approve what inside what organization becomes really critical? And we've been talking about data systems, but give you an example in a lot of our customers now user context platform, not just for data agents, but also for operational agents. And so one common example that we'll hear as a managed knowledge and skills is brand voice. And the brand voice agent is so important in all these downstream agents, like customer support agent, the SDR agent, that's on your website, and so on. And if there's only one person in the company who's allowed to make a change or approve the brand voice skill and the changes that happen in that brand voice skill, right? And so how do you think about these human governance loops inside organizations, ownership maintenance, propagation? If something changes upstream, how does it propagate downstream? Those are all challenges that conflict management was sort of over the next couple of years. And as I was preparing for this conversation, I was looking through some of the materials that you and your team have been publishing. And you make a differentiation between the idea of a context layer and the idea of a semantic layer. And I'm wondering if you can unpack some of the different gradations of what it means to go from. I have a metadata catalog, too. I have a semantic layer, too. I have a context layer. And especially when you're dealing with AI systems where the model itself needs to be something that is reflected in terms of the overall data catalog. And also the fact that a lot of these AI systems may be required specialized data stores, particularly in the form of vector indices, and just how that layers on different technical elements of what needs to be present to have a fully blown context layer. Absolutely. Yeah. So the way I like to layer this, a metadata catalog in its true form was about data context. So it was saying, OK, my data needs to have context associated with it. So I need to be able to describe my table. I need to be able to describe my columns. I need to have lineage that's associated with these data objects. And I need to be able to understand how my data itself, or my data context ecosystem works. And a metadata catalog is the foundation for being able to do that really well. The layer on top of that is what we think of as the semantic layer. And the semantic layer is bringing one layer of meaning to the data itself. Right? And so this is where we say, OK, well, how do we define customer? How do we define new revenue? And that's both the conceptual definition. This is where you would hear people use words like business glossary or metrics catalog and so on. And second, it's the actual, how do I actually calculate this and code this in semantic layer? And eventually, how do I execute this? So what is my execution query, paradigm? This is where you'll see snowflake talks about a concept that they call the semantic view. You have some very specific query providers like cube, for example, that will say, hey, run the query across five or six of my different tools, my BI tools, my warehouse, and so on. And so you would see that as the semantic layer. On top of this, you need to actually add when you think about context. Context is-- you do need data context, and you do need semantic context. But you also need knowledge context, and you need procedural context. This is where your institutional knowledge, your skills associated with it actually start layering on. And that's when you will do true context layer, where you bring data context, semantic context, knowledge context, procedural context together, into that one core overall ecosystem from a context perspective. This is also where-- and you'll you do this-- the primary consumer of the context layer is AI agent. And so technically, this means, well, how does AI traverse context inside my organization? This is where, you know, how do I build hybrid search capabilities? How do I ensure we have vectorization that's associated with it? How do I build the right kind of frameworks? At what point-- at, you know, what's the right protocol? Right? This is where protocols like MCP or agent to agent to A, and what point should AI just actually use SQL as an interface? And it's that-- so how do you think about these interfaces, the protocols, the traversal, the graph relationships, all of these start becoming really critical from a context layer of the spectrum? Does that help? I think in terms of the mental model of thinking through all the layers. Yeah, it absolutely does. And I think it's also worth exploring a little bit more. Some of the details of the maybe data modeling and semantic modeling that needs to be factored in. When you do incorporate the LLMs and the ML models themselves into the overall catalog and the organizational context to say, OK, well, I have this model. It is being deployed in the context of this agent. This agent is using these data sources to produce these decisions and how that maybe changes some of the ways that the catalog itself needs to be able to store and represent the primitives of the information that we would typically have just relied on what are the tables and how are they linked together? Yeah, absolutely. And I mean, I don't know if this is in evolution of the catalog or some anticleror or if it's just the new-- I like the thing of the context. Like sometimes I'm like, I'm like, if you're using a traditional catalog or something, we can just connect it and bring it into the context layer. And the context layer becomes your living compounding system for your AI agent. Your catalog just becomes something your humans are continuing to use for as long as humans continue to be the primary user for data analysis purposes. And so I don't know if the right question is, how much of this needs to go back to the catalog? I think the right question that I like to think of now is, how do you manage context as a living-preaching asset inside your organization? And there, I think these concepts of traces and people start using the word decision traces now are really critical to bring back into the observability loop. So I like to think of it as there is a context repo that is the center of it. That's the central point for this AI agent, which brings knowledge, skills, and tools that that AI agent needs to do its job really well. Outside of the context, all the agent means it's the model and the harness. These are the three things that we can agent. And then this context is a living compounding thing. So it has a main denoer. There are probably governance loops that are associated with this context. The context is changing both from the left because your business is not static. So your business is changing. And the context is also improving over time because with traces and observability, you're able to say, OK, how do I improve this context? So I improve the level of autonomy as I move towards an autonomous system. And that's really the unit of the discipline that needs to get built. As humans go from being the people who do the job, to people who context engineer, these systems-- so these systems can do the job as well as a human can, maybe better. And as you have been working with your customers, and especially as you have been going through this evolution from the first version of the Atlin product, we are going to focus on being the source of truth for all of the data that exists in an organizational context through this disruptive time of AI and agentic capabilities to being this context layer that also has all of that core [BLANK_AUDIO] data context and data semantics. What are some of the most interesting or innovative or unexpected ways that you're seeing teams manage the creation and application of this agentic context and data context to empower their businesses? - The best use cases we've seen are where companies are starting to think about their organizations as how do I buy platforms for autonomous organizations or compounding organizations. So the way I like to think of it is are agents in your company doing most of the work and asking humans for help and how do you get there? And if you think about the paradigm to get there, organizations are at different slopes of maturity, right? The starting point is humans are using AI for improving their productivity. And then you start having your first few truly agentic use cases go to production, right? And you have a customer support agent that's doing more work than the humans that are responding to customers or so on. And then you start moving to like a truly compounding system where you have tens and hundreds of agents that are doing work and then coming to humans for help. The most exciting customers that we have been working with have been really building the foundation of this. You can call it a frontier company, autonomous company. And almost all of them are building this primitive that I like to think of as a company brain. And what I'm most excited about this company brain is that it might allow us to move past the human inefficiencies that existed inside organizations. So we're talking to a large telco customer of ours the other day and they said customer support, if I really want to solve my resolution time, I need to solve my network problem. And these are typically disconnected problems inside companies, right? Like typically you have resolution that's dealt with in contact center, you have network that's dealt with somewhere else, you have discounts, hundreds somewhere else. And most of this is because of the way organizations are created. And organizations have these different silos and siloed ecosystems that are built. AI can actually allow us to break the silo. If you learn something from the ground from a customer perspective, AI can hop from the individual customer use case all the way up to the root cause of that problem and solve that problem. And do you even need, do you even need somebody to make a call anymore to report a problem? Can you actually proactively fix the call? And there is a world in which you can see that world happening and not too far away. And I think that can just dramatically change how organizations function and learn. And so those have been the most exciting use cases where companies are really rethinking their foundational infrastructure, how they operate. And not just slapping on AI on top of existing processes and work that they already do. And in your own work of building and evolving the atlin product and understanding the needs of the ecosystem as it continues to shift and grow and change directions. What are some of the most interesting or unexpected or challenging lessons that you've learned in the process? The most interesting and useful one has been internally transforming into an AI native company. And really pushing the boundary of the true problems that companies face as they work towards building that frontier organization. Our roots at Atlin has been, we were a data team ourselves. We never meant to build a product for anybody else. We tried to buy a product quite honestly and we couldn't find something that's all our own problems. And so we eventually were kind of forced into building atlin for ourselves. But that DNA is really important to us because we've always been very close to the problem one of the reasons why I think we've grown as fast as we have and we're in the pressure and company in our space. And all these other things, these accolades that we get, they're all a function of just, we were closer to the problem. And we cared very deeply about the problem. And so when AI happened, it was really important for us to go back to being really close to the problem. And thankfully we are a company that is at scale. And so we have had the ability to have a front row seat in transforming our company operates. Engineers at atlin don't code anymore. They only teach AI how to code. Marketers at atlin don't run campaigns anymore. They're only allowed to teach AI how to market. And so I think we've been very on the front, we all with that. And that has helped us really, it's a stay grounded in what are real problems. How is context management? We start talking about this context thing. Now well over year and a half ago, for how AI systems are going to need context management systems. And the reason we were able to get there ahead of the curve of where the market is, largely just because we faced the problems of the pain that it takes to productionize these systems. So that has been, I'd say, the coolest part of this journey. Also seen how people can become super humans. I would say when people unlock themselves with AI systems, and there's so many stories that we have in this world right now, around AI, around what is it going to do? Is it going to take away human jobs? And I can't prophesize on what will happen in the world in engineers. But what I do know now is that there are enough problems to be solved in the world that we have not found a cure for in humanity. And humans that allow find a way to turn themselves into super humans with AI will be key to being able to do that. So just seeing humans unlock themselves has been beautiful to see. Seeing our customers have their own cloud code moments when they turn on context agents and they tell us things like this is, I had a customer last week who said, this is a 404,000 percent increase to what we did last year. And so you just see the possibilities of what can happen. And you see people come alive. I think that has been really cool. And the last one, the most unexpected one has been that there's so much education to be done right now. In fact, technology can scale exponentially. They say, organizations scale logarithmically. But humans are slow, but not the same. And right now, there's so much market noise. There's so much hype. And so how do you make sure that people are educated, authentic, are able to trust? I honestly think it's-- I would hate to be a biore of technology right now. It sucks. Everybody is trying to market everything. And it's very hard to really tell the difference between reality and hype. And my favorite push to customers has been just like, are you building? Whoever you are, you could be like CTO, CIO, CDO, are you building? Like, did you build on the weekend? And that just changes. You don't need to believe anyone. You just know yourself on what it takes. And so seeing that also come alive and having a part to play in that has been just amazing. And I think to that while the pace of change and the pace of new products hitting the market has been massively accelerated because of the fact that agent-to-coding speeds your time to delivery, it also, I think, makes it more tenable to do that evaluation of technology selection. So for instance, I was recently looking at, do I want to use Apache Doris or Star Rocks, which originated from the same code base, but have diverged in terms of their overall project goals. And so you can just say, OK, well, here, Copilot or Cloud Code, here are the two code bases, do an analysis of the actual code that's there, the GitHub issues, the pace of change, the types of projects and pull requests that are being opened, and give me a comparative analysis of what these two projects are doing since they have diverged and which one I should be focusing on for this particular use case. And two years ago, I would never even bothered really doing that. I would just maybe read through some of the documentation and make a best guess and hope that I made the right choice. Yeah, absolutely. I mean, it's interesting how much agents might become the primary buyer of most things in the world. And I think we're already seeing this like if you look at the data from some database companies, and you look at, I'd say, the percentage of revenue that is starting to get generated directly from just Cloud Code setting up and using databases. And I'd say that's the first place you would typically see this change isn't a developer ecosystem, right? And I think it will start following in other parts of the world. And in some ways, I think it would be a-- it's a positive change, right? Because it allows you to-- the best product wins, the most useful thing to the consumer wins. And I am really excited by all the new paradigms that that will create in just a way businesses are. And so for people who are watching, please subscribe and like and share this video. people who are interested in figuring out how best to leverage the data that they have, figure out what data they have, what are the cases where you would argue against this very agent-native context layer and instead focus on building out the maybe more traditional data catalog either as a first step or exclusively maybe for a particular level of scale or complexity. I would say if you are a company where you had or an organization where you have not rolled out an eco-pilot to your team and you're predominantly in sort of this era of humans are doing most of the work and at the best case I want to drive human productivity from from this then I would say focus on maybe getting your foundational data catalog right and stay there. I would still say you should not be doing human data stewardship and you should be using an agent-like way of being able to build it but maybe just the consumer being a human is maybe the way I would think about it. However as I say this I will also be having to say I do not believe that any organization should choose this strategy because foundationally the bet of data has always been that you want to democratize data to help business and this has been the dream like for the last maybe 20-30 years this has been the dream and the challenge that you always faced is that in the old world you needed to teach business the language of data and this was you know usually to these data literacy programs and all of these things and for the first time with foundation models you don't need to do that you can actually teach data the language of business and so if you think about yourself and the shoes of the business what does the business care about the business doesn't actually care about data the business cares about running the business and doing that really well and finally from a technology perspective we are there and I believe that is the dream of any data practitioner as to why they do what they do and we are finally at that frontier and so if you're investing in any technology right now there is no downside in just starting at and leapfrogging the generations you don't need to go Gen 1 Gen 2 Gen 3 I think you can just leapfrog to Gen 3 and miss the chaos and the trauma that Gen 1 and Gen 2 brought to you and just move directly to Gen 3 and and roll this out in a way where it will drive like a true AI native kind of adoption inside the organization. And as you continue to build and iterate on your product what are some of the predictions that you have for the next set of architectural shifts that we're going to be seeing in the data ecosystem that are driven by the pressures of AI powered systems. The first prediction I have is there'll be a lot of change and it will be almost impossible to predict the change but change is inevitable and so organizations that build open non-lock-in ecosystems that allow them to evolve very quickly will be very valuable like just let's talk about skills skills as a paradigm did not exist six months ago six months ago MCP has a paradigm record exist a year ago you know and so we're just dealing with the rapid pace of change and evolution harnesses what I thought was a best-in-class AI system in October of last year versus what now is old school way of building AI systems it's been six months and while you could feel some cost bias and say hey you know what like I just spent all this time building this thing now if the rebuild the whole thing or you could say well changes the only reality and I'm going to keep rebuilding and changing along with with times and so the first prediction I have is that teams that change fast and build a culture that allows teams to change quickly and evolve quickly will be teams that win and the second prediction I have is there will be a lot of heterogeneity that you get created in the agentics there will be you know dozens of agentic platforms hundreds of agents inside organizations and it will be a new kind of heterogeneity than we've seen before in that world what I like to tell customers is context is your business IP keep it your own like don't lock it into any individual agent don't lock it into any individual system make sure that it's open it's interoperable you are starting to already see customers tell us about how they've context engineered these two different agentic platforms differently each of these systems have memory so now they're all speaking a slightly different version of the truth this was a classic data problem people used to say you are sales and finance revenue number you will get to different numbers this is just happening at far larger scale for truly autonomous kinds of systems already we're already seeing this happen and so what to want to ensure is that you have an open context ecosystem you're not locked into a single agentic platform or a single model ecosystem yeah I'd say those are my two biggest predictions for the next I don't know six months is that is that a long enough horizon I think that's about as long as anyone can hope to look forward and so are there any other aspects of the work that you're doing at atlin or this overall space of context layers and the shared substrate for agentic and human workloads that we didn't discuss yet do you like to come before he close out the show maybe the one last thing is that I am really excited about the future of the data practitioner data people are the only people inside organizations that have dealt with non-determinism before and I don't think I realized what that means until I truly productionized so much of AI and I realized that you know like software engineering you think in binary right like a zero one and you know if you've worked on a machine learning system you've worked on a data engineering platform you have dealt with non-determinism and there is a foundation where data practitioners have an extremely important role to play in the AI world it will need change it will need learning it will need you know there's open questions like who owns the AI platform inside companies who owns context platforms inside companies how does this play out but I think there's a lot to learn from what we've just gone through in the last decade of data and I'm really excited about just the role that data practitioners can play in the same world all right well for anybody who wants to get in touch with you and follow along with the work that you're doing I'll have you add your preferred contact information to the show notes and as the final question I'd like to get your perspective on what you see is being the biggest gap in the tooling or technology or I guess knowledge that's available for data management today I'd say the biggest gap is the confusion more than anything else there is so much confusion right now in in this space and I mean not to say that there are not technology gaps our technology will evolve pretty dramatically you know in the ecosystem over the next couple of years I'm sure but I think the biggest gap really is how do you get a trusted verified source of context about what to believe and what to not believe the ecosystem and how do you make the right decisions about about tooling technology investments so for example one of the things we've started doing is we've actually been running a WTF is contact is the context layer series and we're trying to just be mystified some of what you and I just did right like you know what is the difference between a subantheck you are in the knowledge gap and the crap database and these are pretty deep concepts to go deeper into and so my hope for the ecosystem is over the next couple of years we end up seeing a lot more real life implementations that create trusted spaces for the community all right well thank you very much for taking the time today to join me and share your thoughts and experience on building these shared context layers for enabling agents and humans to collaborate on organizational data problems it's definitely one of the perennial challenges and definitely one of the faster moving spaces that we have to deal with today so appreciate all the time and energy that you're putting into making that more tractable and I hope you enjoy the rest of your day thank you you thank you for listening and don't forget to check out our other shows podcast.net covers the Python language its community and the innovative ways of it being used. And the AI Engineering Podcast is your guide to the fast-moving world of building AI systems. Visit the site to subscribe to the show, sign up for the mailing list and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email [email protected] with your story. To help other people find the show, please leave your view on Apple Podcasts and tell your friends and co-workers. [Music]

Podcast Summary

Key Points:

  1. Prapadba Senkar, co-founder of ATLIN, discusses building a "context flywheel" for data agents, emphasizing that AI's intelligence is useless without organizational context.
  2. Despite rapid AI intelligence growth (e.g., 1000x in six months), 56% of CEOs report zero financial benefit from AI, with only 20% of projects reaching production.
  3. Human job performance relies more on context (institutional knowledge, skills, tools) than on cognitive intelligence; this "onboarding" process is missing for AI agents.
  4. ATLIN rebuilt its infrastructure on a "context lake house" (based on Iceberg file format) to enable AI-driven metadata generation, achieving 89% accuracy in column descriptions.
  5. The system bootstraps context from existing business systems (CRM, warehouse, BI) and uses "context tree boughs" as portable units for agent workflows, with AI performing most work and requesting human approval for ambiguous definitions.

Summary:

Prapadba Senkar of ATLIN argues that the key to making AI useful in enterprises is not just intelligence but contextual intelligence—the institutional, semantic, and process knowledge that humans acquire through onboarding and experience. He notes a stark gap: while AI intelligence has grown exponentially, most CEOs see no financial benefit, and few AI projects reach production. This is because AI agents lack the "context" that humans absorb over months—such as how "new customer" is defined differently by sales, marketing, or finance.

, Salesforce, warehouses, BI tools) into a context lake house infrastructure. This allows AI to reverse-engineer metrics, ontologies, and relationships from usage patterns and SQL queries. The system achieves high accuracy (89% for column descriptions) and uses "context tree boughs" as portable units for agent workflows.

AI does most of the work, asking humans only to resolve ambiguous definitions. This approach dramatically reduces the activation energy needed for metadata systems, overcoming the traditional barrier where such projects fail before reaching critical mass. Senkar believes this will lead to a renaissance in how companies structure and encode organizational knowledge for both humans and AI.

FAQs

The episode discusses strategies for building a context flywheel for data agents, focusing on how to improve AI usefulness by providing organizational context.

The guest is Prapadba (Prokala Senkar), a co-founder of ATLIN, which builds a context layer for AI to help data agents and humans share institutional knowledge.

ATLIN addresses the challenge of scattered context across systems and people, making it hard to answer simple questions like 'why is the dashboard broken?' by building a shared context and collaboration layer.

Despite rapid AI intelligence gains, 56% of CEOs report no financial benefit from AI because AI agents lack onboarding and context, similar to how humans need institutional knowledge to perform well.

The layers are institutional knowledge (e.g., how 'new customer' is defined), semantic knowledge (e.g., table definitions), process knowledge (e.g., seasonality trends), and user knowledge (e.g., who is asking and why).

ATLIN uses its context harness to connect business systems like Salesforce, warehouses, and BI tools, reverse-constructing metrics, ontologies, and usage patterns to generate accurate metadata without manual effort.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.