From Vector Databases to Knowledge Engines: The Next Layer of AI
46m 28s
The transcription discusses a fundamental shift in data retrieval, driven by the rise of AI agents as primary users. Ash Ashrottosh, CEO of Pinecone, explains that while vector databases were designed for human interaction—where a person queries, evaluates, and acts—agents lack context and brute-force their way through systems, issuing dozens of queries and consuming massive tokens. This leads to slow performance and low task completion rates. To address this, Pinecone launched Nexus, a “knowledge engine” that moves reasoning from the model to the data layer. Unlike a traditional vector database, which acts like a library of unstructured information, Nexus compiles context specifically for agents by curating data into new, structured artifacts based on desired outputs. This process, called “context compiling,” reduces token usage by up to 90% (from 40,000 to 2,000 tokens), cuts response times from minutes to under 500 milliseconds, and improves accuracy from under 50% to over 90%. The key insight is that agents need systems that understand tasks and provide structured, cited answers, not raw chunks of data. Pinecone’s internal use of Nexus for operations and customer support validated these gains, highlighting a new bottleneck in data retrieval rather than model capabilities.
About eight, nine months ago, we started seeing a massive shift to who our users are. It turns out it wasn't a human being anymore, it wasn't a different person, I was an agent. 85% of the agents would work. It's just a retrieving knowledge. Only 15% is the model. The model's under problem. The problem is the underlying system that you're trying to get information from. It brought it down from 40,000 to about 2000. It's under 500 milliseconds from a minute to two minutes. Most importantly, the accuracy dramatically goes up from that in the best case with 168. We have well over 90% accuracy. And that is just worse than what. We finally understood why these things were taking so long. And we're fundamentally running on a system that was designed for human beings. What happens when software is no longer built for humans, but for agents? For years, systems like databases and search were designed around human interaction. A person asks a question, evaluates the response, and decides what to do next. But with the rise of agents, that model starts to break down. Agents don't have context. They brute force their way through systems, issuing dozens of queries, consuming tokens, and often failing to complete tests. This creates a new bottleneck, not in the models themselves, but in how data is retrieved, structured, and understood. In this episode, Peter Levine speaks with Ash Ashrottosh, CEO of Pinecoat, about the shift from vector databases to knowledge engines, and what it takes to build systems that actually work for agents. Hey, Ash. Welcome. Hey, Peter. We're alive? Yeah, it has been a while. Good to see you. We're here to talk about-- Laysam here to talk about Pinecoat's new launch. And I know we-- I'm a board member with you, and we've been on this journey together now for a bit. And we've all been working on this new product called Nexus. And I'd love to hear more about it, and the genesis of it, and then what's happening at the launch and where to from here. Yeah, I think we've been talking about it at our board level for several months now. About eight, nine months ago, we started seeing a massive shift of who our users are. It doesn't seem like a human being anymore. It wasn't a different person. I was an agent. And that shift fundamentally changed how we thought about what's the best way to serve this new user in the world of retrieval. If you think about what we had done for five years, six years, since we first pioneered the vector database market, the idea was you provided an interface to a human being, who did a query, got a response back. And it was the human being who provided the context about whether the response was accurate, whether they had to re-ask the question, and they would finally take the action based on whether they verified information or unfortunately, agents don't have that luxury. The human gives them a task. And agents go there and start trying to perform the task. And they spend a ton of time going through this brute force and loop off clitting, getting some chunks of data back. And when you say, just so I have the context, when you say agents spend a lot of this brute force, what are they actually, let's say right now, before Nexus has launched, what are they actually doing in the background? Like querying what and what's the nature of that whole data flow? So you give it a task to an agent to say, is this product or no warranty? OK, somebody asked them. Right now, let's say without Nexus. Customer service. OK, both comes in and says, can you let me know this product isn't a warranty? Right. There's something about a query expansion, breaks up the queries, and then says, OK, let me go figure out what this product is. And it goes to five or six different systems. But sometimes it might be sales order system, product definition system, things about warranty information. And it seems to have different queries, just like a human being would, because that's the interface to be provided for the database. So here's an agent trying to solve a problem without having any context with a system built for human being. So it goes out, which was a query. And it asks six or seven different queries before it first starts to get an idea about the first-- I think of it as a first line of core effectively. Sometimes it would be 40 different queries. And it might be an internal system or external all over the question, whatever. It could be any kind of stuff. But the idea, most of these guys, most of these agents too, is that do a ton of retrieval, figure out, oh, I don't have enough information. Let me go ask more questions. Oh, I have a conflict here with this information. Got it. So this information goes on until they finally figured out, either, OK, I'm done with the task. Let me report back to my human, but the task is complete. For the human nest to actually examine, because most of the time turns out the task completion rates is less than 50%. So have the task written these agents don't actually complete right? And they take a ton of time. In fact, there's a research study that came out of New Seamurk. And 85% of the agents work. These are just retrieving knowledge. Only 15% is the models. They were built for human beings. You're asking agents to come back and do pretend like-- when we talked about this before, when machines are talking to machines, why do we have an interface that looks like a human being? Right. Yeah, you and I have talked about that for a while. Yeah, this is the same problem. It just happens to be agents performing specific tasks. And that change, you know, user, led to what it means to fundamentally change how retrieval is done by playing-- Yeah. And that's what we're calling Nexus. So maybe to help me and help to put the context here, Pinecon, of course, you know, Bill and define the vector database category. OK. So now we're talking about this Nexus. It's a-- you call it a knowledge engine. Yep. So is this just a marketing term? Like what actually-- like-- Yeah. --vector database. Like, you know, instead of doing this, we'll call it something else. But it's really the same thing. And so kind of helped, you know, for me, just help me to understand, like, one is a mark, you know, one, you know, you just put some lipstick on it. Looks different, right? Or there's a really-- there's a different approach built on, you know, built on vectors, not built on vectors. Like, what's the evolution of Pinecon into this? And maybe a second question on that was, how did you actually bump into this? I mean, what were users doing that informed the company that the shift was occurring and that Pinecon was a viable solution for this? So it's a sorry to break, like, maybe both of those. Yeah. I think the distinction is absolutely real in terms of what a knowledge engine is and what a vector database provides to the knowledge engines. I didn't think of a vector database, like a library. Yeah. There's tons of information out there. A human being asks for some information, appropriate books, and pages, and documents that you wouldn't do when you read through this stuff. And you forget how the knowledge out of it and go back, go and make a task. Now, allow the same vector database to operate with an agent. It has to do the same thing, except it doesn't have the context. So it spreads-- When you say, yeah, the agent-- OK. Go ahead. And it has to go through everything where you read all the pages that are relevant. You synthesize across them. And you hope it got the right answer. And that's the brute force approach, because agents are very, very good at reasoning. They can spin up more queries in a millisecond than you can do in an entire day. Right. Right. And so the brute force they're with, which is why you see a ton of token consumed for even the smallest of the basic application. Now, a knowledge engine is more like an expert, an expert in some task you're performing. You want to get some task done. But let's say you are in the-- you have a medical billing task agent. And a knowledge engine for medical billing, for that specific task, is an expert in figuring out a medical billing part. It may not care about your prescriptions. It may not care about-- Got it. Or just say, billy. Just billy. And then that's saying--
a change in uses the exact same data, which is, you know, it's in a hospital. And may have a very different persona, a very different context when a doctor uses it. Sure. Versus when a hospital administrator uses it. Sure. And that's the difference. I think a Webtoon database treats all data like it's a tool of data like a library. And you need that. That is essential. But you need something else on top that can that literally creates the context. They very specifically. Okay. So so I will get back to the other one. I want to follow up on point B here on that when when so you have I get the library analogy bunch of books and now I get I also understand this knowledge. I think I do the knowledge engine, which is as if you've read the books and it gives you the context. But I'm trying to distinguish between an LLM that I kind of thought did some of that stuff. Yeah. Versus what what added things did pine cone do to turn the library into the knowledge agent. And then, you know, without having an LLM like what is the contextualization in the I mean, we can use the example of the billing service, right? Okay. So now how does pine cone know the context itself? Where does it burn that? Yeah. I guess that's the question. No, I think fundamentally today all of the reasoning is done at the retrieval level, which means once you get the data, you got the LLM, you throw it in. Sure. Yeah. Let me figure out the answer. Yeah. Me on me. Let me try to answer. I don't even know if you have all the data. But all that reason over was based on the data you gave me. Yeah. Okay. When you move the reasoning, closer to where the data is, closer to where the the curation of the data, where the actual processing of the data is happening, you can do a lot more things. So for instance, you can get the right kind of data because now you know what context I'm addressing for. More importantly, you can start citing and attributing. So you actually can say, this is the citation of why and where this answer came from as opposed to a little bit not. It just probably talks to some MCP server, get some information, and through forces, it way to some answer, whether it's right or no. So when you move the reasoning from retrieval to curation, closer to the source, closer to the data, significant differences happen. And what you would do is you would you would tell Nexus, I have this data and typically these are the answers I expect to see. This is my my context. So you give it the appropriate data. When you say you, that's a human, that type of type of thing, you want to set it up. You effectively kind of training, we call it building. Okay. Go ahead. Training the context of the knowledge engine to say, with this data, here's the here is the answers I expect to have. Got it. So based on this test data, this is where the interesting part is very similar to a compiler that before I remember. Sure. You know. I code it. Compile since yeah, to generate some code. This one is a continuous compiler, iterative compiler that says, okay, you gave me this data and you want this output. I want to match it. So I want to keep figuring out how to curate, how to break up this data in a way that create new artifacts. In fact, we actually create completely different artifacts. And this is happening all within the inside system. Yeah. Okay. Reasoning has been moved inside. I said, okay. And that's where you start looking at you gave me this data, but this is the output you want. Let me find the most effective way to just completely break up this data. Got it. Into new artifacts. So for example, in case of billing, you might give it the entire hospital data, but what do you care about is just the patient, the doctor and the bill. Got it. Maybe you don't care about the research part. Got it. After we break that up, that's when we embed that data back into fine code. I say. And so the fundamental shift here is the first bill phase, which is you are now compiling the context very specifically for the knowledge engine. That's one part. So as the new data comes in, it gets converted into this new format. That is very close. Got it. Get cited back to the sources. Yeah. It gets put back into fine codes. Okay. Have to database. That's one part. The second part is on the retrieval site. Now, agent says, I'm not only did I give you the data. I want to get some information. And don't give me a point. Don't give me an image. That's cute for a human being. Give me very structured data. Tell me exactly. I say, you know, very structured for medical. I'm a machine. I understand. Strums. Yeah, machines. Yeah. I understand. You don't care about images or whatever. So that's a second part you define as part of your definition of context. Not only do you do define data and what kind of outputs would you also define the format of the output because the format for billing would be very different on the format for the doctor. Got it. But different from the hospital in the ministry. And how hard would it be for somebody to set this up? Let's say that, you know, you start with the human, they kind of organize things like how what and then we'll get back to help customers actually bumped into this. But what are you, you know, what's the, what's the presentation and complexity that a user has to do literally? In fact, we're working on an internal one for our own contract management stuff. We've done hundreds of contracts. What we did was to say, okay, why don't we take all the contracts we did? Let's on one side talk about the successful contracts. Yeah. Let's look at the input of all the contracts with the red lines. This is your source data. This is your destination. Figure out how I can approve something from here to there. Yeah. And we just loaded it with a print phase. Drone's about three to five turns. Six, a few minutes and you feel like new artifacts. Wow. This is literally, I hate to use the word training a model, but you're training a knowledge in yeah, yeah, they're very different way. Right. Right. It's almost like you're training data to be present. You know, you're training data to you are using data to train a knowledge engine. Exactly. And the data is the foundation. Yeah. Yeah. The output and the format of the output. Yeah. And there are several things and we'll talk about the new protocol that we have this defined to make sure the agents can actually define yeah, how they want to get press boxes back. Yeah. Yeah. Yeah. This is this is literally the massive gap that we've had between models that have spent a ton of time building reasoning capabilities. I think we're completely ignored where the real value is, which is on the data side. Yeah. No, it's fine. And then let's say in this case, the agent now come, you know, let's say we have the knowledge engine agent queries the knowledge engine comes back and you know query understandable yeah, query understand sorry an agent understandable language. Yeah. Would the agent still use an LLM in that case afterwards is that sort of the best is that how this works. And so I mean, my takeaway from that is it will simplify or reduce the number of tokens actually used for the backend LLM system and all that because my data is much more prescriptive when it gets to the edge is that fair as absolutely and three things happen. One, the task completion rate, the success rate of a task has gone up on an average about 50%, maybe 60% in a good day, it goes up well above 90%. You actually have at an agent finishing a task. This is even more important because there's no point giving someone a task even they give it did it for free. Okay, you just did nothing and if it fails, that's the biggest that's even worse. Yeah. So number one is task completion rate goes up dramatically. Number two is a time it takes to come did the task. You used to take if you run today any of the task it takes minutes and part of the reason is spending a ton of time 85% of the time trying to just retrieve knowledge. That dramatically goes down and in our own internal various applications that we've been building on Nexus tokens have gone down depending on how badly how good it was written but in 40, 90% reduction in 20 year model tokens. Wow. And that is a big cost that's a cost savings performance saving the whole thing. Right. And ultimately the ability for you to come back and have quote and quote an expert who gives you precise answers very quickly. Yeah, the lowest cost. That's huge. Yeah, that is huge. I mean, it's really a accuracy performance and cost. It's like all of those benefits come together. Yeah. And that that reason, the problem, the problem for users hasn't been the markets. That's why you get demos really quickly. Right. Right. It takes four hours to put it demo together. Sure. But then yet you understand why is it taking so so long for people and interests? I think the difference here is people have been traditionally using kind of ETL pipelines. Right. The ribbon, taking your data, right? Just like the old database. Yeah. Yeah. This is not an ETL pipeline anymore. Yeah. This is context compiling computing on the mind. Yeah. I love that. I love that concept of context compiling. Yeah. Completely understand that. Yeah. Absolutely. That makes sense to me. Yeah. So Ash, I had asked before in the multi-multi-question, multiple questions like what were customers that we talked about current customers started doing this. And that's how Pinecon recognized that there was an opportunity. So maybe talk about a customer who had Pinecon. And then what were they doing like different? Like how did you know that this was a real opportunity?
based on customers. Yeah, let's take a, maybe the customer zero was, was fine. Go actually. Okay. Because we, we had started building our entire operations, operations agent that allowed us to run our business without dashboards, which is managed the dashboards and moved to a model that kept the entire company's knowledge alive and accessible everywhere. Right. So we had this, we still have this, this agentic backplane called ass data. And every query we put out then would take six to 10 queries to come back with a result, we'll take about 45 seconds or sometimes a couple of minutes. And oftentimes we would come back and actually validate that, that was the right answer. And so, and in the process, we also know, it is to take us about 40,000 tokens. Yeah. And you look at this, I'm saying this is a small application. Now it's bringing data from all kinds of places, our data warehouse, our Slack, our GONG, our clay, all kinds of sources. And then you started looking at what was it doing? It turns out our agentic application and the frontier model just went out and blasted, right, trying to get everything possible. Put it through these agents and keep doing this over and over again. Like you are saying. Yeah. So once we got, we moved that to Nexus, we literally took out 90% of the token usage. We brought it down from 40,000 to about 2000. It's under 500 milliseconds. Yeah. A minute to two minutes. Right. Most importantly, the accuracy dramatically goes up from that in best case with 168. We have a low 90% accuracy. And that is just version one. And that I think is was our first revelation that, okay, we finally understood why. Yeah. These things were taking so long. And then we have a customer support agent that somebody had built about, you know, does act me have, are they in warranty, are they in support? And you would go to three different sources. I could talk about the customer record sales record, the product record. You would watch this whole thing take a lot longer than it should. Yeah. And so, yeah. That was our first principle is to figure out, maybe we need a maybe in your system that actually brings a lot more of the context, much, much more closer to the data, right? Then trying to push push it into a LLM update. Yeah. I mean, again, it's just the compilation of data to provide context and knowledge is super, super important. And with the same dataset, you might have different contexts. Totally. And it's important to make sure that the artifacts that we created, yeah, work related completely on the fly. Like we talked about it's a context compiler, but unlike the regular compiler, it keeps iterating until you've got to the right artifacts and say, yeah, for your context, for your knowledge engine, you want to build for this particular agent, this is the right form. Right. So, you mentioned that there's now this new language that the agent talks to pine cone and all of that. What, yeah. What for Nexus, how does all that work and what was the innovation there? Yeah. So, once we built Nexus and you have an engine where you could have an agent define what its task was, what kind of a knowledge engine it needed, it just didn't have a way to specify that. They didn't, they needed to be a language that both knowledge engine and an agent could actually talk. So, we defined something called no ql, it's a knowledge engine, query language, a knowledge query language. And the intent was to put it into three buckets and six basic parameters. One was in terms of what is the intent of this query. I want to be able to say specifically this particular query has some intent on what what what my ask is, what the scope of the data is. And second is in terms of time, I need this response in 45 milliseconds. Don't take it how to come back. Right. Right. Figure out the best way to get could be the response at certain time. And third one was to really talk about governance, you know, how much of the data is it, am I going to go access? Don't give me the entire data. I need to be able to put a governance across the board. Right. You'll do come back and have explainability. Yeah. This is what we, it's not just the knowledge engine, it's about being a trusted knowledge engine. That makes a big difference. No, no, how you do it in the enterprise. Right. What about the economics of this? And how do we think about that? And how do you, you know, I mean, you mentioned kind of the completion rate and other things. This is it. I mean, if I'm a company, I'm going to go build, let's say build an agent, right? Yeah. Can I quantify this up front? Or do I just wait and see and say, hey, like, you know, I'm going to use pine cone nexus and we'll see what happens. Is there a way to say you're going to get 90% completion? It's going to be, you know, 40,000 to 2000. Yeah. That kind of, how do, is there a certain class of data where we know or you know that that is going to be the outcome? Does that happen on all data? Like what? Yeah. How do you think about it? Yeah. How should customers think about it? I think firstly, if you think about where the cost is today, every word of the application is building that entire knowledge, your tree will stack. It's like, you know, I might go back in time and say every, every database application was writing its own query language, building its own database or even further up saying I'm building my own operating system, my own solution. So one is from a user's perspective, even with our own has data. We saw 85% reduction in our actual code required because that whole part is gone. So that's number one in terms of ROI and TCO. Second is for the same data, how many context engines are you? How many knowledge engines do you want to go provide? So the larger the data set, the bigger it becomes. If it's a small data set by definition, I was pretty constrained. A model you can do fine. In fact, you're up the entire model in a entire data set in a context. Yeah. On the model and they'll be fine. But in this case, this was important for us to go after large data sets with lots of knowledge engines, lots of tasks and agents running across the board and the bigger they are, the exponentially higher overall benefit. Right. Right. Pinecon is an infrastructure company. I mean, just, you know, just step in. In an order for, you know, infrastructure requires applications or agents, stuff to get built on top of the infrastructure. Yeah. So, you know, how do teams think about this? How should they go about thinking about building, you know, apps agents on top of this? And how does one build this in and think about it in terms of the global, you know, sort of stack or basically rewriting the sack here for agents. And so, what should that stack be? And how do I get this? How do I, as an enterprise, actually leverages as quickly as possible? Yeah. It does everyone is saying, oh, we got to go do AI, right? So, you know, everyone's demanding, I mean, you know, the leadership of companies do AI, right? So you get it down the better. So, supposedly, I think if you go back to the DNA of Pinecon, it was started and continues to be a developer centric company. Yeah. You have somewhere between 35 to 40,000 developers who continue to sign up who learn about vector databases. And it is those same developers who are moving and building agent applications. For us, the starting point continues to be making no QL public to these developers. And is that now come with, let's say I do a, you know, for the 40,000 people signing up, it's just built in right up front, or is there added like how does that? Yeah, how do I know about it? No QL. So, one, we have to continue to partner with the agent, harness companies. Yeah. Okay. And we may have to put things like the skill.mds for cloud to define a whole gotter interface. Okay. No different than how we promoted the existing APIs. Got it. We have to start partnering with some of these folks. Okay. So, number one is getting no QL to be adopted by the same development community that adopted Pinecon vector database. Now, as they move up to agent applications, the user
whole new AI API across the board. Second is partnering with we intend to make no QL a open standard. So we're partnering with some of the industry standards at the right time. I think we need to get enough adoption to make sure this thing comes in industry standard. So just like you had SQL for databases, GraphQL for APIs, you expect to have no QL for agent applications. In addition, there's one more part we're also working on is to create a standardized agent stack. What does agent stack look like? No, if you think about your traditional agents or the applications, LLM is a new operating system and pine cone is the disk in between now you have one more thing called knowledge. That becomes a standard stack. And to make it very easy, obviously, we have the core database. Now we have this knowledge engine. Plus we're also opening up something called market place that will be announcing. That makes it very easy for someone to have a free package complete solution. You want time to value. You can go to marketplace and look at either an app that we built or a third party. I said it as a blueprint to see how it's just using my work. Yeah, yeah, or you might want to customize it. The idea is for you to start as the hardcore developer of the database or as an agent to application of the knowledge engine or as an end user with a full flight stack beast. So you can interface on it. And that part is both hours and a third party partners. So let's say just so I'm clear here, we have or pine cone does we. Yeah. 40,000 new people, you know, trying out pine cone vector database. Yeah. Okay. Now, let's just say I want to try Nexus. Yeah. Is that do I add? How do I do that as a developer? Where do I? Is it a new thing that I add on? Is it embedded in? And like, how do I get that? It's just another API service. It's a fully managed service, just like pine cone databases. Okay. So all you have to do is get your agent applications to use no QL. Got it. Completely changed the economics. And the most important part here is once we start working with some of these other partners so that it becomes even easier for these agent to connoisseurs that be agent applications to directly use that. The friction gets even lower. Got it. You know, I have this crazy question. Yeah. Anyway, the crazy question is, is the layer that the the knowledge layer with no QL? Is that dependent on pine cone being there? The vector database? Or can this work with any kind of database? The whole idea is not, no QL is supposed to in industry standard. No, but well, our implementation of it be that way. Yeah. Can work with any underlying. Well, I think Nexus is going to be built on pine cone vector database. Got it. No QL is supported by the axis, but somebody else could build on the list under. Okay. So yeah, that's a good distinction. But Nexus is the full. Yeah, it's an analogy. Nexus has both sides of it. Yeah. The top part then the the the dis part and the knowledge part together. Absolutely. And the dis part and then also in there is also the auto ingest part. Yeah. Being able to connect to all kinds of sources of data. Right. Right. Got it. So you can almost imagine every tomorrow, you can have vertical application. Yeah. Somebody has a great idea. Yeah. You don't go through trying to build your own database, your own operating system. Yeah. You just point us to the data sources. Right. Point us to what context and what knowledge you want to go back and what task you're trying to accomplish. And that's it. After that, you're going to get the word of good application you can focus on. Now let's look out two or three years. Yeah. What is this, you know, what what what does it look like when all this is working? Yeah. And you sort of explain what becomes, you know, maybe what's possible that's not possible today. Yeah. You sort of explain that. But let's say in two or three years from now, how does it solve what? Very similar to the Cambrian explosion that happens every time somebody's standard is the most common layer, whether it is an operating system, it's a SQL interface. Now you'll have an explosion of vertical AI applications or agent applications that now don't have to worry about what kind of a tokenomics you're dealing with, the speed, the accuracy, all you have to do is point us to what data sources you want us to engage with. And certainly you can focus on the real vertical application, the real vertical business case, it is trying to focus on rather than the infrastructure underneath. And like we said before, 85% of the agentic work today is knowledge retrieval. So certainly you're proud of the business of dealing with 85%. Right. You take all that effort put it back into a verticalist. Right. Second part is more importantly, if you truly are deploying in large enterprises, trust becomes important. Yep. So not only do we have a knowledge engine, but you actually have a trusted knowledge engine that gives you the entire trace of how we reason to get to this answer, gives you the citation of where the data came from so that you have an explainably AI. At the same time, you're doing it at a, just not just the economic of using a model, but also you're getting out of the business of building ETL pipelines. You're building knowledge engines completely on the fly. The old model of analytics source, transformative loaded into vector database. One time that's gone. Now your context compiling on the fly as you require. Yeah. And that's a big change in how people go back and deploy it. Today, it's a, today if you think about it, you know, the demo is great. It comes out very quickly. Everybody runs an AI agentic application. Yep. And then they start. They have to go through the ETL pipelines. They're, they want trust. They want security. Right. Right. Right. You should remove all those barriers. Yeah. You just might as simply find and drop down the cost. Yeah. So speaking of cost, how, what does a pricing look like for, you know, and how, how is pine cone thinking about evolving pricing relative to what we're talking about here? We have a first draft of it when we continue to work with several partners to identify what the right pricing is. But it will be more aligned with how knowledge is curated. Knowledge is extracted and tasker completed. Unless about infrastructure. It'll not be about region rights. It'll be at a level that is more about task completions. What kind of knowledge you want to secure it. So we, we continue to evolve that one. Yeah. Yeah. And sometimes we thought about just it could be as simple as how many tokens you're saving you. Yeah. It could be, it could be as simple as that one. But it turns out that itself is not enough good metric because somebody could give you a product for zero dollars, but the trust is terrible. Yeah. All accuracy is terrible. Yeah. And that's useless. So we tried to combine both of us. Yeah. I think one other thing we've done is now that we've been opening up to an entire new interface for agents where you expect a thousand ex more agents than human beings, human users. Probably more. Yeah. It was important for us to also change the economics of the underlying platform itself. The vector database itself needed to enable the economics so that you have a vector database, you have a knowledge engine, you couldn't stack all of them at the same kind of pricing and margins. So we also are announcing the entirely new price point that allows for this entire knowledge engine to be much more successful in terms of adoption. So part of the announcement will be the first of the changing the cost structure for the core database itself. We were doing that openness to the year. So not only are you democratizing the access, but it's also opening up the economics for a lot more use cases. Got it. Got it. Yeah. It's exciting. The fascinating element here. And I'll say that it's hard to believe that this knowledge you know, nexus the knowledge engine here and the compiling of data to make context and all that has such a dramatic impact on the number of tokens used, right? It's astounding. Yeah. And if you just think of like, I mean, this is it's sort of revolutionary in the way I mean, we talk about it, you're like, oh, it's casual. Just put this thing in and you'll say go from 40,000 to 2000. I mean, that's a freaking major major shift. And it's hard for, I mean, just intellectually, it's hard for me to believe that, you know, Pankoan like actually has this. Yeah. And you know, I guess it's you and I have seen this pattern.
before. There was a time when IO interfaces, all of the IO code used to run on CPUs. Yeah. And CPUs are expensive. We worried about the cycles we used. And then you started offboarding that under dedicated processors. Yeah. Like IO boards, IO parts, like the TITABAM part. Yeah. Networking, same thing. Because you started off as another. I mean, all of those different. And you know, based on all the earnings of the specialized functions. That is exactly what we did. It's history repeating itself to say much of this stuff you're putting on very expensive for the models. Yeah. Yeah. You're offboarding that to very specialized things. Right. And allowing applications with. I mean, that. You know, it really strikes me that we're it's, and this is good for the industry and good in general. We're really at the very early innings here of this whole transformation because we think like, okay, it's expensive. There's tokens. Now we're going to optimize. It's kind of like all these industries, you know, like there are past examples at graphics, whatever, networking, all that. They created their whole industries that got created by optimizing the first order, right? So the first order was everything runs on a CPU. Right. And you know, it's oh my god, we got to, you know, have more CPUs and all this. But then it was like, no, no, we're going to take, we're going to offload that CPU and go do other specialized things. And they created during, I mean, of course, like entire industries were created out of that with a lot of the same use case being the fundamental like you got to move bits around on the network or you got to show graphics or whatever. It's just the cost load shifted to a more appropriate area. And that's like what we're seeing here. And it's all I will, I will venture to say no pun intended. There's going to be a lot of this. I mean, whether it's pine cone or other areas of the industry, right? We're like in the first inning of the multi, you know, multi inning game. And then we go into overtime. You know, like it's just, it has been done. I mean, it's been tried. The, the first one was, we looked at this one for some time. We knew the problem. We knew the solution. We also spent a lot of time wondering, are we the right people to do this? Yeah. Yeah. Yeah. First one was, what, what am I just, what do I do this thing? Yeah. And then I'm the do this thing. And you realize, they are too far away from the data. To them is just, right, just data, right, everything just good force you be, right, right. If you're able to far away. Yeah. And not only that each of us uses each agent to make application uses multiple of yeah, yeah, models within a single task. So what am I going to load up all the limbs with the data? That's true. That's unbuilding. So ultimately, it comes back to first order stuff. If you're talking about getting knowledge, it has, and the knowledge is being derived from data. Yeah. You have to be as close to the data as possible. Yeah. Yeah. And we are the 12th point. Yeah. Yeah. Yeah. So it's, I mean, it's awesome. I mean, and I think, yeah, there's going to be a lot of opportunity. I mean, I just think a lot of opportunity, you know, pine cone, pine cone aside to optimize the like AI is, you know, it's incredible. It's magical and all that. But it's, it's a very blunt instrument right now. You know, and like, yeah, we're going to sharpen a lot of things up over the next, you know, the long tail of this is to, you know, optimize the initial edition. A lot of things. The biggest one continues to be that on trust and security. Yeah, for sure. And for sure. That's an opportunity in and of itself. Right. But all these other bits, I mean, you know, and if you look at sort of, you know, the past history of computing, a lot of these things, yeah, repeat themselves in terms of the importance of offloading processes, the importance of security, the importance of data governance, the importance of, you know, applications having the right access. I mean, all of these bits and pieces sort of come together. Yeah, I'll give an example of what else we're doing, which, no, um, MCP interfaces, which have become the de facto way. In fact, I posted this yesterday or day before, as we looked at that, they were the first ways to define access, standard access, model access, any source of data. Yeah. Gen 1, great. Nobody cared. It made it very easy. And now you're finding out each MCP interface, successful. A lot of tokens. Yeah. Because they're not optimized. Right. Now you get to the point where, right. Can I put an MCP interface optimizer? I should be behind access. Yeah. Or maybe somebody else designs around. Yeah. So yeah, for sure. So there are definitely very early innings. Um, I think we find one part of the stack that we think we're focusing on. We're going to continue to have other partners. Well, um, you know, I could, you know, I'm looking forward to seeing how all this evolves. Yeah. We love it. Love it. This is this changes on our daily basis. Yeah. So there's a reason. Yeah. Also, great time to be in the business. All right, actually, thank you, Peter. Okay, brother. Thanks. Thanks for listening to this episode of the A 16 Z podcast. If you like this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family. For more episodes, go to YouTube, Apple podcasts and Spotify, follow us on X, a 16 Z and subscribe to our sub stack at a 16 Z dot sub stack dot call. Thanks again, Phil, listening and I'll see you in the next episode. As a reminder, the content here is for informational purposes only should not be taken as legal business, tax or investment advice or be used to evaluate any investment or security and is not directed at any investors or potential investors in any a 16 Z fund. Please note that a 16 Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see a 16 Z dot com forward slash disclosures.
Podcast Summary
Key Points:
A major shift has occurred
Traditional databases designed for humans force agents to brute-force queries, leading to high token consumption, slow response times (minutes), and low task completion rates (under 50%).
Pinecone’s new product, Nexus, is a “knowledge engine” that moves reasoning closer to data, compiling context specifically for agents by curating data into structured artifacts.
Nexus reduces token usage by up to 90%, cuts response times from minutes to under 500 milliseconds, and boosts task accuracy from ~60% to over 90%.
The system uses a “context compiler” that iteratively matches input data to desired outputs, creating optimized artifacts for retrieval and structured responses for agents.
Summary:
The transcription discusses a fundamental shift in data retrieval, driven by the rise of AI agents as primary users. Ash Ashrottosh, CEO of Pinecone, explains that while vector databases were designed for human interaction—where a person queries, evaluates, and acts—agents lack context and brute-force their way through systems, issuing dozens of queries and consuming massive tokens. This leads to slow performance and low task completion rates.
To address this, Pinecone launched Nexus, a “knowledge engine” that moves reasoning from the model to the data layer. Unlike a traditional vector database, which acts like a library of unstructured information, Nexus compiles context specifically for agents by curating data into new, structured artifacts based on desired outputs. This process, called “context compiling,” reduces token usage by up to 90% (from 40,000 to 2,000 tokens), cuts response times from minutes to under 500 milliseconds, and improves accuracy from under 50% to over 90%.
The key insight is that agents need systems that understand tasks and provide structured, cited answers, not raw chunks of data. Pinecone’s internal use of Nexus for operations and customer support validated these gains, highlighting a new bottleneck in data retrieval rather than model capabilities.
FAQs
Nexus is a knowledge engine that goes beyond a vector database by creating context-specific data artifacts for agents. While a vector database is like a library storing information, Nexus acts as an expert that curates and structures data for specific tasks, improving retrieval efficiency.
Agents lack human context and brute-force their way through systems, issuing dozens of queries and consuming many tokens. This leads to low task completion rates (often under 50%) and long processing times, as 85% of their work involves retrieving knowledge.
Nexus compiles context by breaking data into new artifacts tailored to specific tasks, moving reasoning closer to the data. This reduces token usage by up to 90%, cuts response time from minutes to under 500 milliseconds, and boosts task completion accuracy to over 90%.
Nexus dramatically reduces token consumption (e.g., from 40,000 to 2,000 tokens) and improves accuracy from around 60% to over 90%. It also speeds up task completion from minutes to under half a second.
Nexus allows users to define the context, expected outputs, and format of data (e.g., structured data for billing vs. doctor queries). It compiles data into new artifacts that agents can easily understand, providing precise, cited answers.
Pinecone saw a massive shift in users from humans to agents about eight to nine months ago. Agents were struggling with brute-force retrieval, so Pinecone built Nexus internally for its own operations, achieving dramatic improvements in token usage, speed, and accuracy.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.