LLM Search, UI/UX challenges, Context Engineering and the 80/20 of Eval
52m 36s
The transcription discusses the importance of context engineering in AI agent development, emphasizing the need to understand user context and intent for effective interactions. It explores the significance of search in e-commerce agents, highlighting the blend of keyword and semantic search for diverse user queries. Additionally, the text touches upon the crucial role of user interface design in launching AI agents, stressing the importance of thorough testing and optimization for a user-friendly experience. Moreover, it mentions how LLMs can enhance search functionalities by improving query comprehension and response ranking, contributing to a new paradigm in search technologies.
Transcription
10170 Words, 55037 Characters
People talk a lot about system prompt, tweaking the prompt or
choosing different model. What's the best model out there? One
of our hard earned lessons is I spend so much time on context
engineering with the teams. Let me explain what context means,
right? So in my role, I build AI agents in these different
industries. And these agents are kind of used by two billion
customers around the world. Talk to me about what agents you've
been building. So in general, there are two classes of agents
that I've been working on. One is agents for productivity. So
internally, we built a tool called token, you know about it.
This tool is used by 15,000 users internally at process. And
these are people from finance and designers, product managers,
engineers, and people use it for different things. So we started
building this tool three years back. The technology has evolved
so much. Initially, we did a lot of things to make the LLM
work so that you know, there was an intent detection at the
beginning and so on. It blows my mind that you know, now you
don't need like the model itself is so better that you don't need
these things. So sorry, so one one category is tools for agents
for productivity. So token. Second is agents for ecommerce. And
there, I have experience in building agents for online
shopping assistant. So OLX is one of our portfolio companies. So
with OLX, we built a shopping assistant. The idea is that it
helps understand the user intent. So you can say that, you
know, I want the latest headphone or you know, I'm going for a
hiking trip. And I don't know what to buy. Help me. So it can
take very broad requests from the user understand, connect to
the OLX catalog. So that's an example. And currently, I'm
working on a project with in our food delivery business, where
we are kind of reimagining how people would order food in next
year or two years, where, you know, things like not so usually
if you think about how people order food right now, they go on
the search bar, you enter burger, right? But how about it also,
you know, you can go in, you can say that, you know, I want to
have a romantic dinner with my wife. And then it understands you,
it connects to the catalog and so on.
Okay, so we've got four themes that we want to touch on. First
one, which is on everyone's mind is the context engineering
piece. You've got some takes on that.
Yeah, totally. So Andre Kapati recently published a tweet, where
he said that he likes the word context engineering more than
prompt engineering. And you know, I was reading it and in my
experience in building these agents, I spent so much time
together with the engineering teams working on the context. So
I was completely as soon as the moment he published the post, I
think the context engineering word got viral. And when I think
about context engineering, it reminds me of the day with you
know, when you know, data engineering data scientist, you
know, I was a hands on data scientist at some point in my
career. And you know, we always said that you know, data science
or AI is the shiny thing. But you know, if you're garbage and
garbage out, if you don't work on your data, context engineering
for me is like that, we are doing a podcast, I see people
talking on Twitter on different podcasts, and people talk a lot
about system prompt, tweaking the prompt, or choosing different
model, what's the best model out there? Recently about MCP
tools and so on. But for example, you know, if you take about
the latest models, if you think about the latest models,
for my use case, it's not it's not really about that model. If
I use model A versus model B, and if both of them are state of
the art, they already do a good job. What makes the difference
between A and B is A is with context, B is without context,
then A would be much better than B. And let me explain what
context means, right? So a lot of times. So imagine a food
delivery chatbot, aside from people asking for food. When you
put a chatbot in wild, people ask about many different
things, right? Imagine, you know, you want to order food.
Free food. That's what I would instantly do. Exactly. So
people care different people care about different things. So
there are people who care a lot about free food, cheap food,
promotions. There are people who care about maybe you have a
particular coupon or payment card. So a lot of time people come
to a chatbot, they say that okay, do you accept this payment
card? Show me, show me sushi on promotions right now. Or maybe
it's lunch time and or maybe it's breakfast time. And I am
asking for pizza. And none of the pizza places are open. And
you know, you don't know about my use case as a user, maybe I'm
someone I'm an executive assistant, I'm planning
already for lunch, or maybe I want to have pizza. So in this
case, for example, the opening and closing hour of a restaurant
becomes the context. So if you're LLM, or if your agent
doesn't know about that, then you know, we have tried this
before where if your LL if your agent is just good at searching
for for pizza, and it doesn't know about the opening closing
hour, it doesn't know about the promotion, it doesn't know
about the payment information. If you put it in real world, it's
not useful for people. Because you know, I care about a lot of
these things users care about a lot of these things. So one of
our hard earned lessons is I spend so much time on context
engineering with the team. So, you know, I cannot stop talking
about
Yeah, and I don't want to brush over there's a few different
places where it's a very difficult problem that you were
just saying when it comes to ordering the food and the
promotions, for example, and you were saying that it's not like
there's a database that you can query that has all the most up
to date promotions. Yep. So the data challenge in real time,
knowing what promotions are live from what companies and then
feeding that into the context window is where the real
challenges and it's almost I look at it a little bit like when
you're crafting that prompt, you put in variables and getting
the proper data into those variables is the unsexy job of
the data engineer or it's it's more data engineering like than
anything. But now I guess we're calling it context engineering
thing we all agree that you know, most of the enterprise data,
you know, it might look good from outside, but it's messy. So
unfortunately, you know, if you go in an enterprise, it's not
like there's a database where you can say just give me all
everything on promotion. So a lot of data is real time, maybe
a restaurant is running a real time promotion during lunch
from 12 to three on sushi, right? It's very difficult for a
database to have that. So a lot of that information is in real
time databases, it's scattered over different databases. So the
unsexy work here is that you know, have data engineers in the
team, engineers in the team who spend a lot of time kind of
connecting all this. So imagine now a user request comes in and
they say, show me sushi on promotion. You kick in these
data pipelines, you bring the right context, you make it a
part of the prompt. And then the response that you get is what
the response that the user wants to hear.
And there's not a way to do this with search. That is more
simple.
No, so search is the next topic I want to talk about. But
search solves a different problem. Search is more about for
me. Search is about when you're, when you're searching for
dishes or when you're searching for restaurants. That's what
that's a burger vegetarian burger pizza or McDonald's. But when
you're searching for does McDonald's have promotion? Does
McDonald's accept payment card? Does McDonald is McDonald open
right now? That is context. So that's the difference for me
between search and context context is made up of four
things in my opinions. I spoke about one of them. There's
another one which I want to cover.
One of them being the dirty data getting it the right context
when
Yeah, it's composed of four things. Two of them are simple.
Everyone talks about it. System prompt, user message. And I
think we have heard enough about, you know, you need to have
the best system prompt, user messages, the dynamic message,
the users. And so you put that in the prompt as well. We spoke
already spoke about, you know, bringing the enterprise context,
the dirty data pipeline. And the fourth context, which is very
important, is user history. And this is where, for example,
it's connected to memory. So, you know, there has been a lot of
discussion on long term memory, short term memory. The way I
think about is, there are so many GNI agents, products out
there, imagine, you know, everyone is building a
shopping assistant, everyone is building a food ordering
assistant. If even I think about me as a user, I don't have a
loyalty for a particular product. I'm using chat GPT
today, I'm using another product next day. For me, something
that creates stickiness to a product is if that product
knows about me. So if I'm using chat GPT as an example for last
15 days, and now you give me a new tool, I would use it. Unless
chat GPT knows so much about me that okay, I like the
switching cost is high. The switching cost is high. And that
comes from memory. And that is also part of context where you
put the user history as part of context.
Yes, it's funny you mentioned that because somebody said when
they were joining the community, they were very interested in
the problem set of being able to port that memory and that
context from one provider to the next.
Yeah, that's amazing. If you do that, then then you could
switch, right? Because if you do that, but it's not like you can
just hit like download CSV and upload it to the next
I think there's a lot of data privacy issue with comes with
that because it contains a lot of information about you, which
you might want to share or not. And again, one last point there
is. So I'm sure you have seen these diagrams that people
share about memory, long term memory, short term memory, you
know, episodic memory and so on. So memory can be handled in
many different ways. It's almost like, you know, when you're
having a conversation, maybe there's a model which is running
behind the scene, it's noting facts. And that's a part of
memory. But to be honest, one of the things that we tried, which
is much simpler than any of this and it works is, let's say
you're building a shopping system or a food ordering bot. And
you already have a product, right? So there's already OLX or
iFood. So there's already users who are on OLX, who did not have
conversation yet, who are on iFood, you've got that data,
you got the data, you already know what food they ordered, you
already know what items they browsed, you can easily use that
as a simple cold start context. And I have seen that create
wonders where, you know, when you launch your product, you
don't need to be you know, my advice here would be when you
launch your product, because there's already so many things to
solve. And you know, your obsession should be about
product market fit and you know, not having the best technical
solution out there. You should just use the context that you
have from your existing app, and put that in the memory, put that
in, you know, the context or the system prompt. And that already
does wonder that will give you a runtime of three months. And
then you start, you know, when people start having
conversation more, they start using your product more than you
get some dynamic memory. But I think that's one of the lessons
that I've learned where I think the way people think about it
already from day one, it's complicated. There's a simple
solution for to cold start.
Okay, so that is context engineering. Yeah, that is context
engineering. That's it. Search. Search. Search is such a, you
know, so my experience has been building AI agents in e-commerce,
shopping, food, and search is a fundamental tool when it comes
to that. It doesn't apply for all the agents out there. Let's
say if you're building an AI agent for supplier side, you
know, for, for a car dealer or, you know, for a restaurant,
search is not that important. But when you're building an e
commerce agent, search is the most fundamental thing. It's
actually the start of the user journey. If you don't get the
search right, let's say if I search for burger, and if I get
pizza, I drop at that point, you know, it breaks the trust, I
will maybe you have other tools to manage my car to do other
things, but I will never go further in the journey. So search
is very important. And I spend a lot of time with the engineering
team on, you know, fixing search. Let me talk about few things
that, you know, we have learned. So most of the enterprise search
is still keyword based. And, you know, not to put it down, it's
for the right reason, because keyword based works. If I say
burger, there's a taxonomy defined for burger and you know,
show me burger, you don't need fancy semantic search to do
that. So keyword based works. But if you imagine people
searching on a search bar, versus people talking to an agent,
the chatbot you put out there, the kind of conversation, and
especially when you have voice enabled, the way people express
themselves to an agent or chatbot is so different, the kind of
queries are so different. I would go and I would say that, you
know, I'll already give these examples, right, you know, I want
to have a romantic dinner with my wife, presenter, very, very
broad. Or if I'm building a shopping assistant, we have seen
people just go and say that I'm going for a hiking trip. I am a
beginner, I don't know what to buy, help me, or help me furnish
my house. These are the kind of queries that you get. And you
cannot deliver these queries with keyword search. So we spent a
lot of time thinking about first semantic search. And semantic
search is not new. It has been around. It's a search based on
embeddings. So you, you know, for the search engineers out there,
so you they're already standard way of doing these things.
Semantic search is hard. It's still hard to crack. And keyword
search is still important. So most of the time, the solution is
something hybrid, where you know, you get a query, you look if a
keyword search can answer and can answer it. If not, you go to
semantic search. So it's not new, but it's difficult. And we
spend a lot of time talking about it. So let me explain that
with an example. When I say some vegetarian, when I say
vegetarian pizza, if I have keyword search, then I would get
items that have vegetarian pizza mentioned either in their
title or description. But a pizza Margherita, that's vegetarian.
Maybe, you know, it's not obvious to mention vegetarian in the
title of maybe the restaurant person did not mention it. And
then all of a sudden pizza Margherita will not. So this is
keyword search. This is the limitation of keywords. Now let's
say you move to semantic search. This can be solved with
semantic search, because then you have embeddings and then pizza
Margherita embedding is very close to the embedding of vegetarian.
It's solved. Now let's say, come to the example that I say
romantic dinner, that itself that cannot be solved. So romantic,
you know, what is the embedding for romantic? Romantic is
different for you. Romantic is different for me. That cannot be
solved by semantic. And we see a lot of these queries which are
broad, which are ambiguous. Very fuzzy. So this is where again,
you know, semantic search is the first layer, but we spend a lot
of time building a pipeline of search where when these queries
come in, you have a step before search and you have a step
after search. The step before search, we call it query
understanding, query personalization, query expansion,
call it whatever, where again, you use an LLM. So you already
have an LLM for the agent, but let's say now you have search as
one of the tools within that tool, you have a pipeline where the
first step of the pipeline is already using an LLM to understand
the query. So romantic dinner, it understands and maybe if it has
my user profile, it says, this is what romantic means for Nishi or
romantic in general, and LLM says that, okay, let's break it
down into, I don't know, cupcake is romantic, and so on.
Candle it. Yeah, come on, cupcake. What the hell a kind of
romance are you talking about?
We have different definitions of romance, apparently. I mean, out
of all the ways that I would describe romantic dinner, cupcake
was not one of them. I hope my wife is not listening to this.
But yeah, guilty as charged. So anyway, yeah. So there's a
query, there's a query understanding step, which breaks
down your query, maybe into multiple queries. Then you run
it through your search pipeline, keyword search, semantic
search, then there's a re ranking step where again, you use an
LLM where you know, initially, you get a lot of candidates, then
this is you know, re ranking is again, re ranking is an old
concept from the machine learning world where we had these
algorithms, LTR learning to rank, and so on. But in the new
world of LLM, so those algorithms are still important. But they
also in my experience, I've seen them fail when it comes to
these new kind of queries, where you have a lot of context from
the user. And there we you can have another re ranking layer
where you again use an LLM you say that okay, this was the
query from the user. These are the 1000 responses from the first
two steps. This is the context about the user re rank, you know,
what is and then you get these three or these 10 options that
you present to the user. So that's what our typical pipeline
looks like. And you know, again, I can't stress it enough, it
sounds simple. But search is something that has you know,
haunted me in each of my projects. It's difficult. It's
difficult. It's messy.
It feels like a new paradigm of search to or just like a building
on what was already there and trying to leverage the new
technology of LLMs.
Yeah, totally, totally. Now with the all the advancements in
LLM, generative AI agents, like like agentic search, and LLM
helping in search, and LLM augmenting search, you're not
not replacing. But semantic search is still important. But
you know, doing some steps things before things after so
that's the new paradigm. And so you're right, like search is
really being reinvented with LLMs. And I think, I think very few
people talk about it. But if you're building an e-commerce
agent, search is fundamental. And it's very important to crack
that. Yeah, let's talk UI.
Let's talk you I first talk UI. Okay. So again, I have launched
many agents, many journey experiences. And typically the
way we launch them is we AB test. And I have burned my hands so
many times. You know, it's again, it's, it's there's a lot of
pain there, that I'll give you an example, you know, we want to
build online shopping assistant. The first thing we do is, you
know, let's build a chat GPT for everything, right? So let's build
a conversational experience. We build that we tested internally,
I tested it, I spend a lot of time, you know, long hours at
night, it's super interesting. It knows you, it's connected to
the catalog, you can you say that, you know, I want to furnish
my house, it, it shows you furniture, it shows you, you
know, chairs so far separate sections and so on. It's beautiful,
it works. You're so excited that you're also so benevolent when
you're using it, right? Yeah, you aren't thinking like, how can I
prompt this to give me free furniture? Yeah, how can I prompt
this to get the right furniture? You also know what prompting is.
So you probably, yeah, explain it differently. Totally. So I'm
already seeing how this is going to fail. Yeah, totally. So we are
so excited. I'm so excited. We launch it. We think we build the
best product out there. We launch it, we AB tested our test
like it really falls flat on our face. We are like really like
I'm refreshing the results. And it's, it's, it's terrible. It's
not even you know, it's not even bad. It's terrible. Then you
think, you know, it's terrible because nobody's using it or
no, people are using it, but they don't like it. And they're
giving you that feedback right away or you don't immediately
get feedback, but you can you can measure the conversion
numbers. So A and B, so that's the magic of AB testing, right?
So you have a conversion number on A, you have a conversion
number of on B. And you can see the conversion number is broken
on this data. You look at the yeah, you look at the agent and
everything is amazing. It's still amazing. So why does it not
work? Oh, so and this has happened, you know, again, you
know, I'm, you know, talking about failures, it's also
important to you know, embrace failures. And this has happened
many times. And you know, my learning here and you know, again,
I'm a technologist, you know, I'm not a designer, I'm not a user
researcher. So I also learned it the hard way is the way I think
about is it is so there are two dimensions. One is technology,
second is user adoption. Technology, like, I feel that
technology today is more advanced than the use cases. It was not
always like that. So four years, four years back, I tried to do a
project where I felt that I'm so ambitious with my idea, but the
technology is not ready. These days, whenever I try to do a
project, I feel our technology is, you know, few steps ahead.
How can I use it? That's the problem. So that's technology. But
user adoption is not going at the same pace as technology. Think
about it. We all use chat GPT, you know, chat GPT, it's been three
years now. And I think it's been internalized, people are now
understanding, we're learning new ways to use it every day,
we're practicing. Oh, this actually, I should ask chat
GPT instead of just doing what I normally would do, like ask a
friend or a doctor. Yeah. And you and me, we are probably maybe
biased sample that, you know, we are both in this field, and you
know, maybe we are in a bubble. But, you know, even my daughter
uses it, I see my, my mom does something with it when we go in,
you know, social circle, people who are not in technology, I see,
so it has already kind of penetrated outside. So I won't
say that, you know, it's a hype or it's in a Vienna bubble,
like, it's going. So chatbots for productivity, chatbots for
these kinds of use cases, general use cases, it's, it's a
thing in the world where we live in today. And people, more
people know about it. But think about it. Do you use a chatbot
to buy stuff? I don't. Do you use a chatbot to order food? I
don't. And imagine it's been three years. It's surprising,
it's surprising, you know, that, and we all there's so many
demos, people are talking about it. We are at Ray Summit, the
hackathon, people build this stuff booking.com, there are
these assistants, travel planner, all these ideas in our, our
head. But as users, I eat food, I shop for things, I go for I
book flights, I book hotel. But I so far, and I'm in this field,
I'm obsessed with this technology. But I haven't made a
single of this, like, that's a UI problem. I think it's a user
adopt, it's a UI and user adoption problem. It feels like
the UI inadvertently introduces so much friction. And it's not
the way that we're used to doing things on the internet for our
shopping experience, that we say, you know what, it's easier to
do it the way that I know how to do it.
Yeah, exactly, exactly. So, so after we failed, we did a bunch
of user research, you know, ever, by the way, we have a new
found respect for designers and user researchers, I think, and
now I appreciate. So I work closely with designers, user
researchers in the last few years. And I think I understand
this field much better now, and a lot of respect. So in fact,
these days, you know, when we put together a project, you need
a designer and user researcher, because technology is one side,
you need to understand in the end, you want to solve user
problem, right? So after the test fail, we did a user research,
we spoke to we did surveys, we, you know, we, we called in some
users, and some of our learnings, one. So if you're the user,
and if I, if I give you, okay, now I give you a new way of
doing so you're familiar with the UI, you, you go use you use it
every day. And now I give you a new UI. It's friction. You will
not use it naturally, you will use it. If it's really solving
fundamental problem for you. It's really it's something
fundamentally different, maybe something you used to struggle
a lot that I don't know, maybe looking for house, and you had a
lot of constraints. And a search bar was, you know, you were not
able to define with the filters and search bar. And now if you
could, I give you a voice experience, and you just spoke
to it and it just understands you, then you use it. But if it's
just incremental, if it's a better way of doing search,
yeah, and it's a whole new interface. Yeah, that's a pain
in the ass. Yes, I as a user. Yeah, I would be pretty pissed to
that you changed the interface on me. Yeah, exactly. So that's
what one of our first learning is that when you give people
something new, when you change the UI, it has to be something
you the value for them has to be ready. And they should know it
immediately in the first 30 seconds that okay, this is the
value for me. Because if they feel that okay, why am I even
doing this? Is it why don't I just go to search bar and use my
actually trying to introduce AI into it because it feels like
all these guys just want to be able to say to their stakeholders
or their stock? Yeah, yeah, what do they call it? The
shareholders. Yeah, these guys just want to be able to say to
their shareholders. Yep, they're using AI. So the stock price
goes up. It's so important to handhold the users. Often we
build tool and we just say, you know, there's the saying, right?
So you build, they will come doesn't happen. So you build,
then you need to handhold. So onboarding guiding the users. A
great example that I always like is, you know, I have Alexa in
my house. It's a black box that sits there. It's so inviting.
It says talk to me about anything. Then you talk to Alexa.
And out of 10 things, I talked to Alexa about eight, eight, it
feels two works. But then that's a design problem. So if and
that's the case also with many conversational chatbots, we say
it's a plain screen, right? And plain screen is nice, neat, I
like it. But at the same time, if I'm the user, and you now give
me this thing. And behind the scene, there's an agent which
can do a lot of things. Maybe there are 20 tools connected to
that. I don't know. I didn't build it. So you need to onboard
me, you need to guide me. And over the last few years, I've
been working with designers, and there are some kind of excellent
ways to, you know, some and these are not new in the world of
design. There are you can, you can, you know, maybe when I
enter, you can already have some boxes that I can interact with,
you can have tooltips as I go. So it can be a guide, it should be
a guided journey for the user. So that's a second learning. So
I've tried both a chatbot. So you know, immediately go for
chatbot, it fails. Then you try UI, you say that okay, this is
the UI, then let me try a new UI powered by Jenny. I make it
different, you know, there are a few things happening on the
screen, it's much more dynamic. That has better results than
chatbot. But you know, that's also that that is constrained,
that's limiting chatbot has more flexibility, right? So okay,
that also doesn't work. So I'm kind of coming to conclusion that
the best interface is a mix of UI and chat. So these days, so
there's this word that, you know, we discuss a lot in our, you
know, and with my colleagues is the concept of generative UI.
Even the UI is generative, you know, let me first tell you what
this means. It's like, I'm talking to an agent, I'm having a
conversation. Instead of just replying to the conversation, you
always present me some UI elements. Sometimes it might be,
you know, show me some items, some carousels, sometimes, you
know, related item and so on. And these UI components can be
dynamic. So the agent needs to decide based on my user based on
my previous message, maybe, you know, with the design team, I
build 10 UI components. And then based on the user request,
sometimes you get similar dishes, or you know, similar items,
or, you know, sometimes item carousel, sometimes something
else. And what we are seeing, we still need to test it more.
But what we are seeing is that this experience, because, you
know, we buying is a very visual experience, you know,
conversation is very limiting, you know, I just want to
sometimes, you know, even me as a user, I want to scroll, I want
to click, I want to swipe left, right. I make decision about what
I want to eat based on the image, you know, this image makes me
feel hungry. So I was so if you're
dynamically creating these different widgets on the fly,
only depending on the input that I give you from chat, or also if
I click on something, then you can show me more like that. But
okay, that's a great point. It feels like that is very mix of
traditional machine learning and new agents, in a way, because
if I'm clicking on something, that's a recommender system
problem. Yeah, yeah, totally. If you're clicking on something,
that's a recommender system problem. You know, the traditional
world still applies. And I think that helps in the, you know,
that's how TikTok, for example, you know, your TikToks, you're
looking at these feeds, you'll swipe left, right, and their
recommendation algorithm gets better. But if you think about,
you know, again, if you think about this movie, Iron Man, Jarvis.
Jarvis is an assistant who's watching you what you're doing
in that environment. It feels so natural. So we have built, you
know, it's still not live, but you know, it's at a proof of
concept where we build an agent who was watching the screen. So
it's not just answering you, but it's also watching your actions
on the screen. And then it talks back to you, depending on
those actions. So it whenever you add an item to the cart, it
says, Oh, so a good choice. Or, you know, I knew that you would
like this item or something like that. And that interaction feels
so much more natural than you saying everything based on,
you know, typing chat where, you know, it's just watching you.
Yeah, it's almost like you having to suck that idea out of
your mind and put it into the interface is a lot of friction.
Yep. And if the agent can just watch you scroll and click, then
it can be there with you. And it's much more of a copilot
experience. Yeah. And I've seen that in action. We don't have a
product live yet. But this is something this is an idea we are
playing with right now. But also the user adoption problem
there is I imagine you're going to get a lot of folks that are
like, I don't want you watching everything that I look at.
Yeah, yeah. Yeah. So there's that trade off. Yeah. So we have
we have not cracked it yet. Totally. You're right. So this is
something which we are exploding and every user is
different. So you need to I don't think there's a silver bullet
that I can share.
Yeah, there are people that are fully okay with that. It's like
I'm fully okay with that. My, you know, the way I think about,
you know, data privacy, of course, it's important.
But, you know, my mentality is that take my data if you can
give me value. So if I see value and return, I'm happy to give
my data. So and different people think differently. So last
learning on UI, before we move to next topic is contextual. So
being contextual. So often we, we try to build a chatbot, which
is like the solution for everything. Right. So we give
that as an interface. But one thing that we have seen much more
useful is, and this creates a lot of friction because it's a new
interface, you don't know the capabilities and so on, we are
used to regular interface. So one thing that we work better is
that imagine your regular UI, and imagine, you know, there's a
maybe there's a floating button there. And depending on what
you're doing, maybe you spend five minutes looking for things
and you know, looking for items, it pops up at the right time,
contextual. And it helps you with a very narrow task, very
micro task. I'll give you an example. Maybe I want to buy
a headphone. And I'm looking at a headphone, and it pops up, it
says, do you want to compare this headphone with the latest
headphone from Apple? It's amazing. That's an aha moment for
me. If that happens, I would click on that. And you don't need
an entire conversation. This opens and the comparison happens
using an LLM. And by the way, that's a very simple, it's you
don't need a agent for that. It's a simple LLM call connected
with tool. Well, I guess the hard part is, if you want to put
it into a table, dynamically creating that table and what
yeah, this checks this box, this doesn't check this box,
especially on a mobile screen, that doesn't look good. But, you
know, you know, we have designers and you know, there are
different design ways to solve it. Yeah, there are different
designs ways to solve it. But the point is, instead of having
you know, just a chatbot that does everything and user has no
idea about its capability. If you define some micro job to be
done. And you help to the help the users at the right time. We
in our experiments, we find that much more effective.
Maybe real fast, we can detour into how you make sure it's
coming up at the right time and giving you the right
suggestions. This is where evals comes in.
Ah, perfect segue.
Yeah, this is where evals come in. This, this is again this, you
can compare it with the traditional world of push
notifications. So all of us get these notifications from
different apps at different times. You know, you know, I use
Domino's app like pizza and Domino sends me notification. And a
lot of times the notifications are bad. So and if the first few
notifications are bad, then you stop either either you silence
them, or you know, even if you receive it, your brain doesn't
process it because you know it's bad. If you get a good
notification, if you start with on a good, if you have a good
start where the notification actually maybe it's 130pm and
I'm in a meeting and I'm really hungry. And you know, the app
knows that I'm vegetarian and it pops up that you know, it knows
that I like falafel. That's the context. So if it pops up at
130pm that you know, Nishi, you you haven't had your lunch. I
noticed that you haven't had your lunch. I don't know how it
will know it but you didn't order anything from our app. So
we're assuming you haven't had your lunch. Yeah. So but that's
that's the thing. That's the notification which I would love.
So again, I think this has been this is an old problem. It
applies to push notifications, it applies to other things. And
now it's still true in the world of LLM where if we pop up
contextually, and again, the idea here is to A/B test and try.
So the idea again, the idea is that don't do too much, don't
overdo it, don't send 10 messages. But you can so if we
take food as a context, so there's lunch, dinner. So
breakfast, lunch and dinner. Maybe you optimize you say that
you know, at the right time, when it's lunchtime, you pop up
something and then you use the profile of the person to
whatever information you have about the person to write
something which the person can relate to. And again, this is
where LLM you know, we have done some experiments with the LLM
and it does well. But again, you know, no silver bullet. These
are a few ideas which we are trying that don't be don't
overdo it. And whenever you do it, make it personal. You know,
people say don't make it personal, I say make it personal.
I could see how you get people that are in my case, for
example, I'm checking out on one of the apps, looking at some
food. And at this particular restaurant, I've already added a
few things to my basket. And it says, Oh, hey, I know you like
vanilla milkshakes. These guys make an incredible vanilla milkshake
and that pops up at just the right time. And do you want to add
that to your basket? So again, it just reminds me of like
traditional recommender system. Yep, in a new way. And so it's
almost like, both with search and with this, totally, you're
doing these old tasks that we did in predictive ML, you're just
adding an extra layer on top of it to make it even more
personalized. Yeah, and even better, ideally. We talk a lot
about this internally that the problems are still the same,
right? So we are still recommending Netflix, organize
this recommendation system competition, I don't know, 20
years from now, 10 years from now. So we still talk about
recommendation, we still talk about search. And recommendation is
not completely solved. Search is not completely solved, sending
notification. So it's the same problems. But now the toys are
different, the tools are different. And those tools are
helping solve the same problem in a better way. Yeah. Okay, going
back to evals. What is your take on those? Yeah. So when I
think about evals. So we talk, you know, process is also an
investment firm. So I speak to a lot of founders, and we talk a
lot about more, you know, what is the mode of the product, how
to differentiate your product. And you know, a lot of time, you
know, you see people say that, you know, people don't share their
system prompt, you know, system prompt is more than you know, I
understand that you spend a lot of time to engineer that. But I
believe in evals so much that I think, you know, when I'm not a
founder, but if I would be a founder, and if my system prompt
leaks, I would not be worried. But if my evals leak, I believe so
much in evals, then I if I would be a founder someday, I would
say that evals is the real mode of your product and not your
system prompt. A lot of times when we launch these products,
before we launch them, how do you know if it's good enough,
right? And again, this is the same problem that you can think
of in an regular software engineering world where you
build software. And software development lifecycle is much
mature now where you know, there are this quality assurance,
there's a field in the entire field quality assurance,
testing, testing in production, regression test and so on. So
there's this entire thing. And now with AI applications, and now
when the entire product is about AI, we come to this new world
where and especially now LLMs are non deterministic. So it's not
your regular AI. So how do you how do you know if something is
good enough? And we spend a lot of time, we debate a lot. So
there are you know, people in the team who disagree and you
know, everyone is passionate. And you know, someone says that
okay, we did not think about this case. It doesn't work. Other
people say that okay, maybe you're being too particular. It
works. It generally works 80% of the time. So the answer to that
is evals. So you can do that in a systematic way through evals,
where the idea is that you build a system to check evals is
nothing for me but a system to check if what you build is good
enough. That's the first case, do you launch it? And second, once
you launch it, once you start getting real user queries,
because it can also degrade users can also ask different
things which you did we are not prepared for evals is that the
traffic light that keeps you informed, so that you can so
there's a part of it which is pre development during
development and there's a part of it, which happens like in
production. I've heard it explained as again, going back
to like traditional predictive ML, you have offline online.
There's like the offline training batch jobs type thing. And
then there's the online like, boom, real time type of
predictions. So few mistakes, which I have made myself and I
see a lot of people make. So I'll talk about that in evals. One,
I think it's a mistake to wait for your product to be launched
to write your evals, because you know, you want real user data,
right? So evals need data. So imagine you have a chatbot and
you need to see what how people are using it in the wild. But
it's too late already because you launch it. Maybe your chatbot
sucks and you will know it much later. So so this is where we do
a lot of simulation synthetic data generation, where before it
goes live. And it doesn't have to be again, it can be simple.
It's like, you know, you get together in the team, and
everyone is playing with the chatbot. And you create evals
based on that. Or you use another LLM where you know, maybe
you are the the lone engineer in the team, you give it 10
queries, and then you use an LLM that you know, generate 100
more queries. And that's your synthetic data, or you give a
different personas. So a lot of people they wait for they wait
a lot before evals. But I think first step of evals is
simple. It's even a person in the room just, you know, firing
queries, you create a data set of 2050 100. And then that's a
data set. And then you see how your agent responds. And maybe
you manually go through it. And if you go through it manually,
you know, in half an hour, you'll realize that, okay, these are
the scenarios where it does okay, these are the scenarios where
it doesn't do good. That influences your thinking that you
know, what metrics should I what metrics are important for me?
How should I use an LLM as a judge? LLM as a judge is a
popular concept, right? For evals, because human labels are
expensive. I'll assume the audience knows about it. Another
thing that you know, which I've noticed is that people
immediately run to LLM as a judge. But again, there are a lot
of low hanging fruits. I'll give you an example. It's so LLM.
So these days, you know, we are a bit spoiled. So LLM as a
judge, let's just, you know, users and so you you get the
input for synthetic data, you get the output from the agent.
Now you ask an LLM, you prompt it a bit that okay, did this
input satisfy the user intent? And it's easy, you can do that.
And immediately people run to that because it's easy. But
there's something which is even easier, which a lot of people
miss is a lot of time you have more deterministic because you
never know that maybe the LLM itself is making mistake, right?
And then you have the final label, which is a mistake. A lot of
time, there are more deterministic metrics. So I'll give
you an example from food ordering. So imagine a food ordering
experience, user comes in, you have the agent conversation. In
the end, if the user generates a cart, and usually we send a link
to the cart, then that's that means that conversation was
amazing. It was positive.
They bought something they bought something. So the final
metric that you're looking at is did they convert or not? Did
they convert? But even you know, it's getting a little bit into
the details. But you know, sometimes people come to the cart
stage, but they still don't order. So there's a so not always
here, but hey, but I'm saying that was there a card if you
link up if you think about a funnel that you know, a user
making search a user going couple of steps. And if you think
about a later stage in the funnel could be about making
order could be about adding to the cart could be about
something else depending on what the use cases. But those are
positive signals that this conversation actually went through
different stages of funnel. And it's a deterministic metric
because I can see that okay, I have these hundred conversations
which of these conversations went to through this stage of
funnel. And that's a positive conversation. So grab that for
the evils grab that for the evils. And that is so much better
than using because LLM as a judge will make mistakes. Look for
those metrics, whatever those metrics are for and don't don't
run immediately to LLM as a judge, there's a step before that.
Another thing I want to talk about is so if you again, if you
look on internet and if you see advice from people out there on
evals, a lot of good advice, but it very soon it gets complex
where people talk about, you know, conversation level
analysis, singleton, multiple turn analysis,
the silent failures on the silent failure, the eval frameworks,
yeah, tool calling, you know, did we call the right tool? Did we
call the did we call the tool with the right parameters? Did if
you have multiple tools, did you go if you can even do a state
management, because maybe a user query calls five tools. So you
can even map that okay, from tool one to tool two did the
maximum mistakes so it can very easily blow up and you'll be kind
of paralysis analysis, you'll be scratching your hair. Often, you
know, we we present these evals to business folks and imagine,
you know, you have a meeting and you know, you have these 12
different levels, people get lost and you know, you lose the
audience, you lose everything. So for me, you know, for us, what
we have seen when we build products that process, it's
like the first level of eval is very simple. It's like, this is
the query from the user. This is the entire conversation, you
know, multiple queries, maybe there are 10 messages exchanged,
you use and maybe you tried the usual metrics. Now you come to
LLM as a judge. Before you go to turn by turn analysis, before
you even look at what tools were called, just give the entire
conversation to an LLM and ask some basic questions. Ask it
that did it satisfy the user intent. Did the conversation
went in a direction where you were trying to the agent was
trying to close the order because you can in the end, you
care about making a sale. And these are so simple and this
business people, they relate to it. And then an LLM will make a
judgment, you know, true, not true, partially true and so on.
And there's already so much information there. If you now see
that, okay, these are the things that work, these are the things
that don't work. To be honest, a lot of time, our evals stop
there. We don't because there's already so much information
that there's already so many things for us to fix based on what
doesn't work, that we never go to level two, because level two is
you know, tools, did you call it?
So there's there's almost like a hierarchy. Yep. And in your eyes,
the hierarchy is first and foremost, there's like, looking
at the metrics of, did they try and actually make the sale?
Yeah, because that's our top level metric that we're trying to
affect the needle on. Yeah. That's really all that matters at
the end of the day. So if the agent isn't closing, yeah, then
totally, you got to be who was it in the color of money? What was
that? Who's the actor that was in it? You got to always be
closing and you got to teach your agent put that in the prompt.
Yeah. So yeah, that's one way to look at it. But of course, you
know, on a lighter side, does the agent follow answer? So if
you were the user on the other side, did the conversation went
in a positive direction? Or you know, the agent asked for
burger and you give them pizza? So did it satisfy the user
intent? Did it go in the right direction?
So so there's those are like, very strong evals that you need
to be focusing on first. And then if you need to then you can
start peeling back the layers, like did it call the right
tools? Did it use the right parameters when it called the
right tools? Yeah, yeah, I can see that. Yeah, it's almost like
there's an 8020 here. Yeah, that preos principle. Yeah, we've
got these evals, they're going to give us 80% of the important
stuff. Yeah. And it's only these two evals that we got to look
at. Yeah, totally. So that has been our experience. So we
also made this mistake where we already got into complicated
evals. It doesn't resonate with business. It leads to analysis
paralysis. So these days, you know, whenever we do evals, we
ask it practical business questions first. And then there's
already a lot of insights there to act on and only then there is
value in going deeper.
Yeah. But it's not as actionable, I imagine. Yeah, like you say,
analysis paralysis where you're like, there's so much that we
need to do here. So coming back and saying, is it answering the
user's question? Is the user experience nice? Yeah, is the
metric being moved? Like the metric we care about? Is it trying
to move that? Yep. Yep. Those are very, very basic. Yeah, very
useful. Yeah. There are two more lessons when it comes to
evals. One is about so at process, something we do. This is a
recipe that we have tried many times. And it has been very
successful is labeling party. So we call it labeling party,
where, you know, we invite, you invite your team, you invite
some stakeholders, it's important to invite, you know,
business folks as well, you get 15 people in the room, you order
some pizza, that's why it's a party, it's important. You book
one and a half hours, you spend time with these people to
actually go through real conversation. And then this is
what the agent responded. Now, if I imagine and then the question
you asked them is that imagine you are the user on the other
side. And then you ask a couple of these questions that we did
the conversation go in the right direction. Did it try to close
the order? And a couple of other things, did it kind of break
any guardrails or whatever, it depends on you know, what the
agent is trying to do. So you define a couple of questions,
easy questions. So don't talk, don't talk about technical
stuff like tool calling and so on, business related questions.
And what what would come from that is that you spend one and a
half hours and at the end of one and a half hour, you'll have
I don't know, every person will label 10 data points, you'll
have 150 data points. And when you look at that 150 data points
and now run your LLM as a judge, the prompt that you wrote and
ask it to label these 150 data points, you'll see that there's
a difference between what LLM says as human obvious, they'll
always be different and label data and the LLM is a judge.
And then you will see that depending on your use case,
sometimes the LLM will be more lenient. Maybe the LLM always
said that your agent is perfect. Or maybe the sometime the
LLM would be stricter. So depending on the use case, I've
seen both. But now when you see the difference where the
discrepancy is between human and LLM label, you look at it
again manually, all this is manual. And then the way to use
this data is you use this discrepancy to inform yourself
that okay, this is where the LLM is making mistake, you know,
human is right here. Now go back to your LLM as a judge prompt.
Give it examples of few short labeling, give it example that
when this question comes in, this is what you do. So I think
often we also use we take think LLM as a judge very easy, we
just write a prompt, we expect magic to happen and get
something but our experiences you need to iterate on that
prompt and the going full circle is context engineering.
Yeah, yeah, it's context engineering. All these things
are related. But you need to improve your LLM as a judge. So
we spend a lot of so we do these labeling parties, we improve
the prompt, we give it few short examples based on what the
this and now your LLM as a judge, you run it again, we see
that okay, now it's closer. Earlier it, they mismatched, I
don't know, 30% of times. Now the mismatch is 15%. And again,
you need to keep doing it every two weeks, do a labeling party.
So that is a recipe that we have tried and it's been very
successful. So yeah, that's an actionable insight. Yeah. And
it's a manual process. And but it's building. It's yeah, it's
rewarding. It's, it's nice as the builders of the system to you
know, sometimes look into the data yourself and
and hear from the business side of the house. I imagine, hey,
this makes no sense. Yeah. Why would we hear about this? Or
why, why would the LLM say this? Yeah, you discovered a lot of
other things which you know, you were not even maybe hoping to
hear. But that happens. Last point there. It's actually kind of
it's a little it's related with labeling party. So when you do a
labeling party, how do you offer? It's very important how you
offer people how they can look at the conversation. And this is
where for example, we use a lot of these external tools, Lang
Smith, Langfuse. So there are these observability tools where
it captures the interaction. But at the same time, if you're
looking at Lang Smith trace, it's very user unfriendly. It's
like, you know, a bunch of Jason's if you have multiple tool
calls. It's like, you know, I love looking at it, you know, I'm
a technical person, but the business person gets lost. So one
thing again, which we do, which works very well is with now with
very low effort by using v zero versus and so on, you create
custom annotation tool. There are standard tools as well in
market for labeling data like label studio, there are the tools
out there, they're good. But you know, our learning is that every
use case is different. You need to visualize different things for
the user. I'll give you an example. If I'm looking at food
conversation or shopping assistant conversation, if I'm the
user who's evaluating it aside from the conversation, I also
need to see that Okay, these are the items which were shown to
the user at that point. These items and I need to show it seed
visually. Otherwise, how will I judge these items are from these
restaurants, maybe these restaurants are open or closed at
that time. Because if you don't give me that information. So
annotation tool plays a big role. And what we find ourselves
doing a lot is, you know, with couple of it's like hacking, we
just quickly put together a tool and half a day using v zero
reversal. And people love using it. And then you get the audience
is much more involved than going through a trace and lang smith.
So again, it's small point. But it makes a big difference user
experience. Yeah, user experience.
Podcast Summary
Key Points:
Context engineering is crucial for building effective AI agents, focusing on understanding user intent and context.
Search plays a vital role in e-commerce agents, requiring a mix of keyword and semantic search for user queries.
User interface design is essential in launching AI agents, involving extensive testing and optimization for a seamless user experience.
Leveraging LLMs (Large Language Models) enhances search capabilities, enabling better query understanding and response ranking.
Summary:
The transcription discusses the importance of context engineering in AI agent development, emphasizing the need to understand user context and intent for effective interactions. It explores the significance of search in e-commerce agents, highlighting the blend of keyword and semantic search for diverse user queries. Additionally, the text touches upon the crucial role of user interface design in launching AI agents, stressing the importance of thorough testing and optimization for a user-friendly experience.
Moreover, it mentions how LLMs can enhance search functionalities by improving query comprehension and response ranking, contributing to a new paradigm in search technologies.
FAQs
Agents for productivity and agents for ecommerce.
Context engineering helps agents understand user intent and provide relevant responses.
Search focuses on retrieving specific information, while context involves understanding user queries in a broader sense.
System prompt, user message, enterprise context, and user history.
Using existing user data for a cold start context can provide valuable insights and improve user experience.
Semantic search helps understand user queries beyond keywords, enabling better responses to broad and ambiguous requests.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.