10. Observability in the Age of AI with Jeetu Patel, Atin Sanyal, and Yash Sheth
0m 0s
Yash and Atin, former Google and Uber AI leaders, founded Galileo to address the critical gaps in trust, reliability, and cost control in modern AI systems. Their decade-long experience in building conversational AI and AI infrastructure shaped a vision focused on AI observability across three key pillars: infrastructure resilience, agent behavior (including evaluation and runtime enforcement), and token economics. Galileo’s integration with Splunk and Cisco’s digital resiliency stack enables enterprises to monitor, analyze, and secure agentic workflows in real time—critical for industries like healthcare where misbehavior can lead to severe consequences. The solution provides deep visibility into token consumption, detects anomalies at machine speed, and supports intelligent routing to use cheaper, more efficient models without sacrificing accuracy. Galileo’s small language models reduce inference costs by 95%, making observability scalable and cost-effective. Customers gain confidence in deploying large fleets of agents—like moving from one to 75 in months—by having full visibility into spend, performance, and compliance. With capabilities for forecasting, runtime interception, and actionable insights, Galileo transforms AI from a black box into a measurable, trustworthy, and economically viable asset. This full-stack approach ensures enterprises can confidently scale agentic workloads while maintaining security, governance, and financial accountability.
Meet Yash and Atin, Founders of Galileo
Hello again and welcome to the Inside Products Podcast.
I have with me today two very special guests from a recent acquisition that we made of a company called Galileo, Yash and Athan.
And we have Vikram, who is the 3rd Musketeer who is not here today because he has had a new addition to his family.
So he's enjoying well deserved time with his family.
So we'll, we'll make sure that we have him on the next pod.
But welcome to both of you.
Thank you, thank you.
Before I get started, I actually, you folks have had an amazing background.
You were at Google, you worked with deep mind.
Talk to us a little bit about that.
What did you do over there and what learnings did you have?
Before we go into, how did you think about founding Galileo and what happened from there?
Building Conversational AI at Google for a Decade
Yeah, so, you know, I've, before we started Galileo, I spent about a decade working at Google building conversational AI systems.
You know, we started with very rudimentary speech to text systems and you know, there was a time when we were able to convince Larry, Sergey and the business that conversational AI is going to be extremely, extremely important.
That's going to be the main way by which the.
Speaker 1
World consumes.
Speaker 2
Technology, yeah.
And so we've gone over the a whole host of models over the last, you know, decade, whether it's like right from caution mixture models to RNNS to DNNS to Transformers and then to large Transformers.
I think the, the big unlock that LLMS, which is the the large transformer architectures have had is we didn't have to really hard code the natural language understanding parts into like a, a hard coded workflow that initially these assistants were built on and on top of that, these models became multimodal.
And so, you know, we when we, you know, we are part of this, you know, early team that was building out these models and scaling them.
You know, we worked with, you know, demos and team as well.
And on top of that, the, you know, our team powered about 20 plus products at Google, whether it's the Google Assistant, but you know, Google Cloud and Google meets and the cars and the home devices and the phones.
So you have.
Speaker 1
The same voice engine that you had, actually.
Speaker 2
The same conversational AI engine.
So it's the voice and natural language understanding trending along with, you know, actions that you can do in different modalities.
So I think that's.
Speaker 1
Were you in charge?
Also?
Did you also think deeply about like, tone and tonality?
Speaker 2
Tone, punctuation and then even quality of the generated voice.
Like how can we?
Speaker 1
Make it sound.
Speaker 2
Natural exactly and so that was that was adopted by contact center businesses and you know like the meeting you know in COVID we saw a huge spike in you know these Google meets and Webex and all of these products and then conversation AI transcripts and now we take that all for granted now but like that was a tough time for us to scale these this technology to long form speech and long form conversations so that's where I come from and you know towards the end of my.
Speaker 1
Do you currently run research and core engineering and this team and then often you run the product in the engineering side?
Speaker 3
That's correct, yeah.
From Siri to Uber: Building AI Infrastructure Foundations
And so often.
What was your background?
Speaker 3
So I started my career in the earliest days of Siri, worked on some of the first versions.
Speaker 1
Tom Gruber.
Speaker 3
Tom Gruber yes, and this is still in the Imagenet Alex Net days where deep learning is still kind of in the labs, but really learned how to build end to end productionized language systems before deep learning models were a thing.
Spent a few years there and then really joined a very interesting team at Uber.
Back when Uber was honoured there, they were growing and they had barely any AI footprint.
It was a team of six or seven people.
It's known as Michelangelo, and we literally built out their entire.
Speaker 1
AI.
Was this before Monica or during Monica's time?
Speaker 3
This was before and during.
Speaker 1
OK.
Speaker 3
Yes.
And but we really were a few of the folks who built out end to end AI infra and the foundations and blueprints which have been adopted pretty much across the globe.
We were one of the first to build feature stores, which are data systems for machine learning models as well as monitoring systems, where I really saw the battle scars of how do you make AI models do the right thing and how do you measure them.
And then met Yash and Vikram who were solving the same problems at Google.
And we, we sought out together and we said, how do you solve the trust problem in AI?
Speaker 2
For unstructured data and language models in particular like.
Speaker 3
That, yeah.
So unstructured was one of our key pillars.
And we said that we want to build for the next 10 years of AI, not the past 10 years, right, because we had seen deep learning models get cheaper and cheaper to productionize.
So we saw that coming.
We didn't, of course, foresee exactly how it would manifest, but we kind of feel like we hit the nail on the head on the on the unstructured vision.
And not only did it, you know, was it exciting for us because it was future facing, but we feel like we spent the first couple of years before a chat GPD happened, really building the foundations of understanding AI models.
Jeetu's Three Pillars of AI Observability Explained
Yeah, sure.
Nothing.
What I'm going to do is I'm going to take a step back and say if if I were to think about how observability for AI is broken down, in my mind, the three distinct components to it #1 is the observability for the infrastructure itself.
So how are you going out and measuring resilience of the infrastructure, GPU utilization, uptime, all those pieces networks?
The second area then is agent behaviour, observability, where you can actually see how the agent is behaving, seeing if they're they're behaving in a way that you don't want it to behave.
And when that is happening, can you dynamically at runtime intercept to kind of alter that that that that behaviour?
And then the third area is around tokenomics, you know, are we is the consumption pattern of tokens for these agents within the boundary conditions that you want to actually have or is it actually going out completely in, in a different stratosphere which could then create a very, very different economic impact for you?
Am I missing anything in that?
And then what are the things, how are you folks thinking about how Splunk fits into that entire picture?
Speaker 2
I mean, I think this Jitu that these three things, you know, literally cover the entire egenic stack.
You know, we while infrastructure seems like, you know, we sometimes ignore infrastructure and think that it's commoditized, but like you know, as agents are skating, we are seeing customers hit infrastructure bottlenecks very quickly.
You know, as we all know, just the networking impact of agents is just 450 times that of a human.
And so our customers are actually asking us for an end to end observability across not just the agent behavior which we talked about in in quite depth already with evals and enforcement, but how you know, how are the infrastructure components dealing with agents?
You know what, what network accesses do these agents have?
What are the security layers that we can put on because without observability, you can't have security.
And so if you want to run agents, whether it's in a desk side computer or it's in a server rack or it's in the cloud without end to end observability, we can't enforce the right guardrails around, around the agentic processes.
The next thing is the, and the, and the final piece is tokenomics where, you know, I, you know, tokenomics is actually bubbling up quite a bit now as a, as something that's of importance.
But you know, we've, we've always tied the end outcome, the ROI of agents to the spend that that we're making.
You know, it's when we talk to customers in the past in the first couple of years of, you know, post chat GPD, we've had so many dinner table conversations with CI OS and saying, hey, I spent $15 million on LLMS.
What's the ROI like?
Are we shipping more product?
Are we, you know, are we actually moving the needle on our revenue?
And the answer has been no in the in the past, but now I'm starting to to see that change where there is measurable revenue, there is measurable impact, there's measurable productivity.
But the thing that's missing is the infrastructure to measure that itself.
How can we quantify and tie the agent's behavior and the tokens spent to the business impact?
Speaker 1
Value ROI.
Speaker 2
Because what leaders want to actually do is we don't want to spend less on AI.
We want to spend more on the right pieces that are giving the right ROI.
Identifying Key Users for AI Observability Solutions
There are 1000 things going on with AI.
There's a lot of tinkering going on.
Speaker 1
So give me a concrete example that a customer can say if I have these class of problems I should be thinking about observability for AI with Splunk.
Speaker 2
So there are three classes of problems when we think about who needs observability for AI.
Most enterprises today are having some form of agentic behavior in their workflows.
These are tools and and Productivity Tools and coding agents for their developers.
That's that's the biggest spend today and that's the the biggest amount of agentic activity.
That being said, you know the first party and 3rd party agents, those are the the other two categories where first party agents are the ones that that enterprises build themselves and that could be around their internal workflows, but also consumer facing products.
And third party agents are the the entire agent sprawl we are seeing in every SAS provider having their own agents.
Speaker 1
And you want to make sure that you provide observability for the entirety of the estate of all agents it.
Speaker 2
Has to be.
We cannot just have like blind spots around.
Oh, I measure my first party agents, but what about the rest?
Like you know, but the point I was trying to make is the, the agentic behavior in products and services is only growing and today is just the early days.
As more and more, you know, workflows in our tools in our in our products become agentic, the the spend on and the observity needs for agentic on 1st party, third party will increase.
But we today as as we stand, we cannot ignore the coding agents and the tools that that vertical.
Speaker 1
As well.
So I'll, I'll come to you, Athan, on the, on the product side of the house, because there's so many products that actually use these terms that sometimes it's hard for a customer to say, I can't, I can't tell the difference.
Concrete AI Observability in Healthcare: Prior Authorizations
So here's my question to you.
If I'm a customer at a bank or an insurance company or a pharmaceutical company, pick a, pick an industry or a vertical.
What is a concrete example where what you do and what we do with Splunk Observability for AI can meaningfully change the game on the kind of value they harness from AI?
Speaker 3
Absolutely.
I'll give you a concrete example, say, of a healthcare company, right?
They're building workflows to solve critical, to make critical decisions on, say, prior authorizations of expensive drugs, right?
They're using Jenny I to solve it.
And they're building agents to automate this whole gnarly process.
And it's ripe for disruption because these workflows are just filled with unstructured data, PDFs, documents, and LLMS are for the first time, we have technology that can make sense of it and take actions for it.
So we have the ingredients.
Once you build an agentic workflow, a whole host of things can go wrong at many layers, including the the three layers that you mentioned.
I'll talk about the first layer, which is tokenomics and second layer which is the the behavior.
Starting with the behavior.
One wrong mistake that the agent makes because of a loose prompt or a wrong instruction can lead to a bad action, which in this case would be the denial of a drug for a patient in need.
So the results can be catastrophic because of wrong behavior.
Then there's the cost aspect to this where tokenomics comes in and I'll tell you how layer 2 and layer three are related, right?
There's the absolute measurement of tokens that will that will be needed to make this take this action, including parsing all the PDFs, the docs, the insurance docs.
All this requires tokens and we can't escape that.
But how?
The question is, how do you do it efficiently?
Because these models are probabilistic, they can go wrong and they can spin into loops.
So it can be very expensive to get this action done.
So the problem to solve is 1 measure the total number of tokens it took to take the action and two whether the action was taken or not, which is also key part of tokenomics.
But the in insight and the signal for whether the action was done right or wrong comes from the layer below.
Speaker 1
So that's a OK.
So I, I, I measure whether or not the, the, the key actions taken with the use of tokens.
What do I do if I find out that like what's a good scenario and a bad scenario in that particular instance?
And what does our our technology help with on that front?
Speaker 3
A good scenario is that there's action completion happening, which is the agents are doing what the instructor action is, which is to approve or deny the drug and do it with high saliency or high accuracy.
The bad scenario is of course doing the wrong action, but the bad scenario could also be very.
Speaker 1
Consumption on tokens while.
Speaker 3
Extremely slow because you know, whether it's a healthcare workflow or an e-commerce application that can, you know, deliver goods to you and issue refunds, the architecture of agents behind the scenes is the same and they have the same proclivity to, you know, the, the errors that they can make.
Speaker 1
Why?
Galileo's Vision: Full-Stack AI Resiliency with Splunk
Why did you folks feel like being part of Splunk would actually be additive to your mission?
Like what was it that that Splunk brings to the table that you felt like you didn't have?
Speaker 2
I think, you know, to, to my, you know, to my point earlier, you know, the, the thing that we're super excited about is that Cisco is building the critical infrastructure for the AI era.
Splunk as a business overall is looking at solving the digital resiliency problem and agent take or AI resiliency is the next chapter in, you know, in that in this entire business.
We also heard from our largest customers, some of the largest Fortune 100 customers that they have an entire digital resiliency stack, whether for, for application monitoring, for, for enterprise security.
And then as more and more services start becoming agentic, they asked us like, how is your platform going to tackle all of it?
Because we want the entire digital resiliency to be in one end to end flow, in one end to end solution.
And and so from a product and technology perspective, what we've built is extremely additive as well as strategic from Splunk's customers as well as our as well as Galileo's customers.
Speaker 1
This is a truly a full stack observability solution from infrastructure observability, application observability and and then going into agent behavior tokenomics and pulling it all together in a way that's one cohesive interface in Cisco Cloud Control.
Speaker 2
Absolutely.
And there's the, you know, apart from Splunk's capabilities, also Cisco, Cisco has the AI defense aspects that you know, the, the defense claw and AI defense as you know, pieces that actually, you know, fit in very well with the story as well as thousand Eyes and security solutions.
With, with all of these capabilities, Cisco can truly provide the, the solution for critical infrastructure for AI agents.
Like it's, it's not, I'm not saying that lightly because you know, this is, you know, 1 bespoke point in time solution is not going to solve.
And the, the, the overall problem to build and, and all of this allows us to build the, the fabric for digital resiliency, agent resiliency across agent software like.
Speaker 3
You know, whether you take the example that took Fable down, right or the healthcare example I just gave, these are anecdotal examples of singular workflows that of course have the same same failure patterns.
But really, when you imagine an agentic fleet running in production that happens across multiple devices, multiple data centers, multiple availability zones in the cloud, and to detect a security issue in an agentic flow, the data can come from anywhere.
So providing #1 the unified data fabric to be able to a unified interface across all your data, you need that to solve the security problem.
You need observability on all this data to solve the security problem.
It all blends into one another.
So having this full stack really are is saying that I have the key ingredients to solve agent reliability.
And Cisco, from what I know, is the only company in the world that has that.
Speaker 2
And to your point, like Cloud Control stitches it all together and we're super excited about where we're going in the direction every product in Cisco is, can, can be.
Speaker 1
Operated integrated into cloud control in like two or three days or something like that.
Speaker 2
Yes.
And you know, we literally like, you know, our, we joined Cisco and in a week at Cisco Live, we demoed, you know, cloud control with agent observability and tokenomics.
And so that's, you know, that's been an incredible journey as well because, you know, it's so easy to build cloud control apps now that every product in Cisco can be operated through one interface.
Ensuring Trusted AI Delegation and Agent Protection
So let me try to summarize and see if we can actually and, and, and correct me if I'm wrong on any of this stuff.
In order for us to have agents do work on our behalf, what is what do agents need?
They need to make sure that they're treated as digital Co workers and have tools access in order to have a pattern where we can delegate to the agents in a trusted way rather than an untrusted way.
What you have to do is protect the agent from the world so that if there's prompt injection attacks, there's data poisoning, all of that, that stuff's happening.
We can figure out a way to intercept it with things like AI defence, and #2 we can also make sure that we protect the world from agents that are misbehaving.
And all of that has to be done with detection that happens at machine speed and scale.
Absolutely.
So what you folks have built is the full stack and the observability apparatus to ensure that all of this gets stitched together.
But the key components that Galileo brings are agent behavior observability, which includes evals and then tokenomics so that you can continue to keep refining how the agents are behaving in the consumption pattern of tokens and marry that with the infrastructure observability and the application observability that we have with Ollie Cloud and App D And and the reason this is really pertinent to be part of Splunk is because Splunk has time series machine data.
And observability is fundamentally a time series problem.
And so you want to make sure that you can actually go out and utilize that data in a way that's effective.
Speaker 3
Even security, right?
Security is now a series of events, right?
Like the adversarial attacks I mentioned, it's all time series.
Speaker 1
Right.
Quantifying Customer ROI and Cost Savings with SLMs
And if a customer does this, what have we seen as the primary ROI that the customer gets?
Like what?
What does a customer go to bed feeling Peace of Mind with?
Speaker 2
I'll just give out an actual data point like you know one, one of the one of our largest enterprise customers went from one agent in production to 75 agents in production in a matter of six to seven months.
You know, if you have the right agent resiliency stack, then the enterprise can feel confident all the teams with building different agents and, and even an onboarding agents from different solutions feel confident that they have the right playbook in place.
And, and that's extremely important.
If we can not only deliver the right products, but the right solution is the most important thing.
You know, when we when we go to customers, you can say now you have the end to end playbook of what every team needs to do What, what is your governance and compliance strategy?
You know, and, and our products support all of that.
That becomes magic, right?
That that becomes magic to our customers, yours, because we're coming with this end to end solution that will help them accelerate.
Speaker 1
And by the way, you folks have build your small language models SLMS so that we can actually lower the cost of inference as well for the observability stack that we're using.
It's like down by like we can reduce it by 95% that's.
Speaker 2
That's huge to the point that like our, our, the, you know, the price that our customers pay, they actually, the, the savings that they get from the SLMS actually make up for the price as well and and actually end up saving more so.
Speaker 1
Your inference cost is de minimis in this thing?
Exactly.
OK.
Deep Dive into Tokenomics: Forecasting and Interception
So given all of this, are there questions I should have asked you folks that I didn't that you want to convey to our customers?
Speaker 2
Yeah, I think so.
One of the things you do that we're hearing all the time these days is how are we solving the tokenomics problem?
Our customers have, you know, have been seeing a huge spend, you know, unexpected spend on tokens and then they don't have the the ability to measure the the right ROI from.
Speaker 1
Their and they don't have the ability to fund all of that either because it's hundreds of millions of dollars and people don't just have that lying around.
Speaker 2
And we're always competing on, on what to, what to fund, what to invest in.
And so I think there are a few things that that we're looking at deeply from, from the from building the right product and and then also helping customers with insights.
I think the key problems that we're trying to solve for with tokenomics is firstly, how do we give customers the deep observability into where is the spend?
Speaker 1
Happening.
Happening by which agents?
By which?
For what activities and what tasks?
Speaker 2
What activities, what projects, what tasks which, which personnel, you know, what are the budgets that are going off, you know, all of the measurement part.
But that's not enough.
What customers have been asking us is like, how can you help us determine what's the ROI from these agents and, and how do we map the spend to the ROI?
And then second is can we actually do some powerful forecasting on these spends?
And, and Splunk with its time series data, the Cisco data fabric getting all of the time series information.
We have our our time series foundation model that was built from scratch on on time series data.
We're leveraging these capabilities to better give give customers better forecasting on their token spend that is music to people, I think.
Speaker 1
There's going to be a third area, which I'm not sure if we have fully cracked the code on, but we're going to need to, which is the dynamic interception, yes, when agents start to go rogue so that we can actually stop.
Speaker 2
That actually there is a rudimentary solution today.
Speaker 1
But it's a very blunt instrument, right?
Speaker 2
Now it's a very blunt instrument like we we, you know, any single prompt cannot be more than these tokens.
If an agent that's right session goes beyond these many tokens.
Speaker 1
But over time, if you were to think about from a product hygiene standpoint, like how would you want to build this?
You would want to make sure that you almost in runtime, start to instruct the agent to pull back and not have that level of consumptive pattern being exhibited for the job that they're doing.
Speaker 3
Absolutely.
I mean, runtime is really the only way to curb cost, right?
Is to detect rogue agent behaviour or anomalous behaviour at runtime and.
Speaker 1
Prevent not just intercept, but also influence and redirect the.
Speaker 3
Agent redirect.
Exactly.
That's the actionability.
Speaker 1
Bit yeah, exactly.
Speaker 3
That's the key part of tokenomics, which is detecting the issue and taking the action and curbing the wrong action, which at scale would lead to millions of dollars of wastage.
Speaker 2
And you know, Speaking of blunt instruments, there's a fourth aspect to tokenomics is today, Speaking of blunt instruments, most, most users are just using the best, the most powerful model for every single thing that they're doing.
And this is another ask that keeps coming up is like you.
Speaker 1
Have intelligent routing.
Speaker 2
Using our eval capabilities, can we measure that you could do the same job with a much with a 10X cheaper model and give that proactive insight?
Speaker 1
Post facto.
Speaker 2
Post facto, but still.
Speaker 1
Proactive so that you can then go back and actually reapply.
Speaker 2
It because most teams are not investing the time to to to measure which.
Can there be a cheaper model that can do an equivalent amount of job?
Speaker 1
Sure, sure, sure, sure.
So that actually is a very interesting point, which is like over time be able to say that this is how much this task cost you, but this was the distribution of the task between a Frontier model and an SLM.
But if you had done it with this level of routing intelligence, you would have actually seen a fraction of the cost.
And here's what the cost would have been.
And so that's what we're going to incorporate the next time we do something.
Speaker 2
While not sacrificing accuracy because that is the most yeah, yeah.
Speaker 1
Gentlemen, thank you for being at Cisco.
Thank you for actually doing all the amazing stuff that you're doing.
It's a pleasure.
And I'm sure we'll have many, many more of these conversations.
Speaker 2
Thank you for bringing us here.
And it's it's an exciting time, this is.
Speaker 1
Exciting time.
Podcast Summary
Key Points:
Yash and Atin built extensive experience in conversational AI at Google and early AI infrastructure at Uber, emphasizing the need for trust, observability, and scalable agentic systems.
Galileo’s core mission focuses on full-stack AI observability—covering infrastructure, agent behavior, and tokenomics—to ensure safety, security, and measurable ROI in AI deployments.
By integrating with Splunk and Cisco’s ecosystem, Galileo provides a unified, end-to-end observability platform that enables real-time monitoring, dynamic interception of rogue agent behavior, and intelligent token cost optimization.
Summary:
Yash and Atin, former Google and Uber AI leaders, founded Galileo to address the critical gaps in trust, reliability, and cost control in modern AI systems. Their decade-long experience in building conversational AI and AI infrastructure shaped a vision focused on AI observability across three key pillars: infrastructure resilience, agent behavior (including evaluation and runtime enforcement), and token economics. Galileo’s integration with Splunk and Cisco’s digital resiliency stack enables enterprises to monitor, analyze, and secure agentic workflows in real time—critical for industries like healthcare where misbehavior can lead to severe consequences.
The solution provides deep visibility into token consumption, detects anomalies at machine speed, and supports intelligent routing to use cheaper, more efficient models without sacrificing accuracy. Galileo’s small language models reduce inference costs by 95%, making observability scalable and cost-effective. Customers gain confidence in deploying large fleets of agents—like moving from one to 75 in months—by having full visibility into spend, performance, and compliance.
With capabilities for forecasting, runtime interception, and actionable insights, Galileo transforms AI from a black box into a measurable, trustworthy, and economically viable asset. This full-stack approach ensures enterprises can confidently scale agentic workloads while maintaining security, governance, and financial accountability.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.