Cultivating product thinking, cross-functional leadership & the future of AI agent infrastructure w/ Jaikumar Ganesh #247
49m 43s
The transcript features a discussion on engineering leadership and product thinking in the AI era, hosted by Patrick Gallagher with guest Jai Kumar Ganesh, Head of Engineering at Anyscale. Key insights include the shift from "how fast can we build" to "what is worth building," requiring product thinking across all engineering layers. Scaling patterns are context-dependent, varying by company culture and business model, as seen with Uber's "move fast" approach versus Anyscale's thoughtful scaling for ML infrastructure. Ganesh emphasizes that leaders must align with strategic goals, such as revenue projections, and prioritize P0 initiatives without "peanut buttering" resources. For infrastructure directors, core responsibilities like reliability and scalability require proactive planning, including educating executives on trade-offs. The conversation also covers infrastructure trends: the need for auditability in non-human workflows and managing complexity from autonomous agents. Modern tools like Kubernetes reduce team size needs, but strategic decisions (e.g., cloud dependency) remain critical. Ganesh advises leaders to ask the right questions, automate where possible, and focus on customer pain points to drive meaningful innovation. The episode highlights cross-functional leadership to break silos between engineering, sales, and other functions, aligning technical teams with revenue and customer success.
This episode is brought to you by SADARO. If your team runs Kubernetes, chances are upgrades can feel risky. That's what the TELUS platform is for. It's built on a multiple, minimal OS designed for Kubernetes. A minimal means clusters cannot drift, so they stay identical and upgrades are a non-event. Minimal means there is almost nothing to attack. 50 binaries, no shell, no SSs, so it's secured by default. Upgrades gets boring, security gets boring, and that's exactly the point. Check it out at SADARO.com. That's S-I-D-E-R-O-Labs.com. It all comes back to truly understanding your customers and asking the right questions. I do believe all engineers will need to have product thinking. And this can mean different things. The layer of the stack that you actually built, why are you building this? Who's my customer? What is a pain point I'm throwing to address? I'm just not going to listen just because some product managers told me he built this kind of thing. The fact that it takes the amount of time to build is going to keep it reducing. This becomes even more important. One person can launch three agents to iterate on three prototypes to see that, which one has legs kind of stuff. Instead of having debates in a room, you can build some foundation building blocks. Either that will clarify the questions you have any overhead. Is this a writing to build? Is this going to be to complicate? Or it'll give you enough data points to ask the questions with your customers' kind of stuff. Is this the right angle kind of thing? Hello and welcome to the Engineering Leadership Podcast brought to you by ELC, the Engineering Leadership Community, and generally founder of ELC, and I'm Patrick Gallagher and we're your host. Our show shares the most critical perspectives, habits and examples of great software engineering leaders to help evolve leadership in the tech industry. We've been exciting and expansive episode all about cultivating deep product thinking across all levels of your engineering organization, near term infrastructure trends for enterprise-ready AI agents, cross-functional leadership and culture change coming your way right now. With Jai Kumar Ganesh, Head of Engineering at any scale. In our conversation, we discuss a bunch of different topics. Like how in the AI era, the primary constraint for CTOs has shifted from how fast can we build to what is actually worth building? And we get into how to develop product thinking across every layer of the engineering organization to meet this challenge. JK shares two critical infrastructure concepts for the next 12 months to help you ensure auditability in non-human workflows and prevent organizational madness as thousands of independent agents begin to operate within the enterprise. Plus, we deconstruct leadership lessons from JK's unique experience as a GM, tackling topics like breaking down silos between engineering and sales and other functions. Culture change and cross-functional leadership practices to align technical teams with revenue goals and customer success. Let me introduce you to JK. JK has a deep background in engineering and customer facing roles with a proven track record of building and scaling engineering organizations. He previously co-started and co-led Uber's AI group, which was the central ML group at Uber. It was also in the early team in Android at Google. Enjoy our conversation with Jai Kumar Ganesh. Well, all to say, JK, just wanted to first off welcome you to the show. Thanks so much for joining us on this Wednesday. How are things going on in your world? Yeah, thanks Patrick. Excited to spend some good time talking. Stuff is fun. We know lots of things happening in the video and the industry and you're right in the middle of it. Right in the middle of it, powering a lot of it. And it's a special time to be building that way. I wanted to set up some context for why I've been excited for our conversation and then wanted to dive in and get some of your thoughts and reflections on how things are shifting. So you've scaled engineering orgs at multiple major companies. And not only that, but you've also scaled multiple AI ML products at some of those companies. And now at any scale are leading engineering plus several other functions within the company. And when I think about any scale, like you're also not only scaling the organization, but any scale itself powers the infrastructure for how modern companies are scaling right now with a lot of their AI ML products and otherwise. So you are you are at the intersection of all these things going on. And so with sort of those dual perspectives, both on how infrastructure is changing and how modern engineering orgs are scaling in your past experiences. Like I just want to open up with your perspective and reflection on the scaling patterns of different errors. And what's similar, what's different? What have you been reflecting on and what are you sort of seeing in this space? And my thought is we'll kind of will span organizational and infrastructure. So we'll definitely dive into both. But yeah, I just want to open up with your reflection, JK. Yeah. So a lot of this obviously is going to be a big idea. Have you perspective geography matters a lot? I did start my career in India when I was working at Motorola during the heat. Days of Motorola. So scaling organizations has changed a lot. And I actually don't have a crisp formula because what I've learned is that it all depends on the company culture, the domain, where are you located, what kind of people you have. So people who say about, oh, this was phase one of scaling. This is phase two of scaling. This is phase three of scaling. I actually call it bullshit to be honest, because it the nuance, the details, all these things vary from company to company. But at the same time, there was some general trends in the 2000s. A lot of that was like engineering ICs who were driving a lot of the decisions. The craft really mattered. So for example, when I was a nice seem, my manager used to have like 50 to 70 director boards and hardly had anyone on once and stuff like that, right? And then afterwards that slowly changed where there was a lot more material and good quality engineering management coming into practice and engineering management as a key profession about taking care of your people, etc. and the industry itself honed in a lot. Now with AI, etc. that's changing a bit more again. So I would actually say that's a phase three. You've seen all the chatter around founder more. You've seen the chatter about like, you know, hey, do we need so many layers of management anymore? The more people are a lot more productive using AI is how that is actually going to change. So I see these three clear phases of scaling organizations which has changed over in the last 20 years. A lot of it depends upon the company and how what the company optimizes for. So for example, at Uber, it was about move fast at all costs because whoever got into a city first had an enormous marketplace advantage. So it's scaled too rapidly for its own good. That system worked for Uber at that point of time. We just took that system and put it in another company's, it would have totally failed. This is where copy-pasting of scaling does not really work. That is all very situation dependent. I think maybe this is a two-part question because I think you're so right like being able to kind of define that theme of like what is our sort of thesis around scaling. So like for Uber, move fast at all costs because being first is a marketplace advantage. There's like a there's like a defined thesis around there. How would you describe that for any scale? And then maybe the second part of the question is how to help people either formulate that or optimize around whatever whatever that thesis is for them. I think for any scale, it's much more, I wouldn't say to Uber level scaling. It's more around more, more thoughtful scaling because we are in the ML in fras space. So ML in fras space generally moves a bit more slower than a B2C company. So we need to make sure our core components are rock solid and we have two parts. We have open source rate and we need to make sure open source rate adoption keeps growing. And there is a lot of technical depth in the various components because it's a genetic system. It can do LLM inference. It can do non-LM workloads. It can do batch processing of data. So it can do training of models. And so we need to it's like an operating system that you're actually creating. So when you're creating an operating system, you just can't throw 100 people at the problem and think that in three months it's going to get done. Right. It's like laying brick by brick by brick by brick and making sure the foundations are actually solid. Right. And the same thing with any scale, which is a commercial arm of Rayway customers who want the full platform come to it. It's also we are powering their system. So we have to be very, very thoughtful. So it reflects back on organization scaling, right? You have to be very, very thoughtful on where your components are kind of stuff. And then the strategy also has to play in with the ecosystem that you are actually playing with. Right. So for example, ML and for companies are not going to scale 0 to 100 million ARR in one year. Because ML and Friesmo, it has a very strong mood. And then it builds up much more slowly like a snowball. And then after some time it just becomes too big. And you're like, oh my goodness, this is an outline. Right. And so you have to scale in sync with your company with your company's core product and the needs of your customers. One CVE patch can put a whole fleet in doubt. What should be routine takes too much time and it often ends with your best engineers fighting fires instead of shaping neural map. The root cause is the enemies, a general purpose operating system that was never built for this. The moment someone runs a package manager or a box at 2 a.m., you have a drifting note. And worse, the OS comes with a large attack surface that you never need it. That's what the tells platform is for. It's built on a multiple minimal OS designed for Kubernetes. A multiple means clusters cannot drift so they stay identical and upgrades are a non-event. Minima means there is almost nothing to attack 50 binaries, no shell, no SSS, so it's secured by default. It handles the fleet's lifecycle for you, provisioning, upgrading and retiring machines automatically. That design matters most when the sticks are heist, including the edge with no one on site and AI clusters where drift is expensive. Obgoise gets boring, security gets boring and that's exactly the point. That's s-i-d-r-o-labs.com. For maybe folks that are, let's use a hypothetical example. Maybe somebody is at an organization and they're kind of grasping blindly at what their scaling thesis is. How much do you recommend them to define, here is what we're optimizing for. Our version of we need to be thoughtful about this and scale with our customers and then optimize their engineering organization or company operations around that.
So let's talk from not a big company, but maybe a 500 or a thousand or a hundred percent company. Obviously it starts from the CEO, the executive leadership team saying like, you know, what is it that the company is going towards next year? What is their revenue projections? What needs to be built? Should the sales team be scaled more because now we have a product that you can actually sell or should you know should our indeed team be scaling fast? So these are all strategic decisions that need to be taken, right? So everything else flows from that when there is clarity and that aspect of it. And the TC's force scaling comes once from that because if you are still dependent on VC money, then you need to make sure your burn rate is under control and stuff like that. If you are free cash flow positive, then you have a bit more flexibility on where you want to invest in, which are the places where you know, you can start off here. I need to start off this new group because I want to create this new product that doesn't exist in the market. I'm going to start off with a team of 10 people on that and you definitely have more flexibility when there is a printing press. So best example I give is Google. Google had a huge search plus ads printing press. And that's the reason they could do things like Android and Chrome and now Waymo and all these things, right? I mean, almost 99% of the companies can't do that today, right? So your scaling is completely dependent upon your business model, how much money you're making. So the scaling TC just goes from that. So to give some tactical advice that it comes back to, you're knowing where your P0 priorities are, they actually funded well. And whatever system you use for thinking like, you know, to define big rocks around them to make sure they are funded well. And the mistake many organizations may or many individuals may just treat everything as equal priorities. Even I have made this mistake. To do everything as equal priorities and then you know, peanut buttered it and nothing really gets done, kind of stuff, right? So I wonder if there's another question around this because let's say some of these things are unclear. And I'm thinking of maybe the person of somebody who's like a director of engineering and maybe it's kind of like reporting up to a head of edge VPV, CTO. And maybe some of these things are unclear. Like what might be like essential questions for that person to ask to gain that type of clarity around like where are we going next year? How do we make money? Like what do we need to be conscious about in terms of how we're funding our organizations? Like, yeah, so I think, so let's take an infrastructure, uh, director of stuff, right? And there is a reason I'm picking that position because the main job is reliability and scalability, right? So they need to know what's the scale numbers that they need to hit. I'd actually put it the other way around. If they're actually asking the question, hey, tell me what's the scale they're already not doing their job. They already know the limits that their system is hitting. And what they should ask is like, hey, what's our revenue growth projection going to be? Is it 2x or is it 5x? And that's the whether it's a finance person or the salesperson can give them that number and then they can back calculate saying, hey, this is a scale that is needed. And for this scale and this reliability, these are the systems that are undeliable and I need to be able to staff of these projects. And so for that, I need to add five people and usually you need to be at least six months ahead of the process. So if you're doing this for January, you need to have this conversation and June off the preceding year so that you need to hire up and build up that team, right? So no, again, it comes back to what your core competency is, therefore infrastructure person. So for a director of infrastructure, it's about reliability and scalability. And so then knowing which components are reliable, which components are scalable. And if your head of engineering comes from a different background, sometimes they come from a product engineering background or your CEO may not necessarily come from an engineering background. So it's your job to educate them saying, hey, this is the impact on customers. And some CEOs may say, that's fine. Given the company state, I'm okay to take that tech debt for next one year. We can fix it year later. So it's your job to educate to make sure, okay, but you need to understand the consequences of that and it's your decision. So that's how I see if I'm a director of engineering for infrastructure to say that what are the two things I'm most responsible for? Indian scalability. Okay. So then think about like my team's position to that and my 20% group and they're asking me to do 10 different things. Hey, I just cannot do that. Right? So in today's world, what are stuff I can automate with the agents? What is something I really need people for? That's how I would think about it. Well, I think this is a great sort of segue to kind of dive into some more of the patterns around infrastructure and how that's in how that's changing. Because I think like this director of infrastructure example, as you were sort of sharing some of the possible solutions that they could explore given maybe capability or headcount limitations. There's a lot of ways that this is shifting here. So maybe we can talk about like some reflections on infrastructure and maybe some of the changes in the space like from your time at Android or Uber, how are things shifting? What were some of those patterns and then maybe where some of the shifts happening right now? Shifts have happened in multiple ways. So for example, let Android, Android is creating an operating system on the phones, right? So it does not like infrastructure as we think about it in terms of service and services, but there was also a big back in component. So a drone teacher that Android had was that Google is way ahead of the rest of the industry in terms of like, you know, the scalability of their back end services. So you have your Gmail app on the phone, but the back end was the scalable infrastructure that the Google already had built up and was probably the best or if not one of the best in the planet at that point of time. This is like 15, 20 years background. And so there are a lot of scalability of Android was about just how to make good reliable phone software and how to get the level of abstraction correct between what goes into the operating system and what goes into the apps because you had to ecosystem of app developers and you have to give them the right APIs and the right abstraction layer. Should we build this in the operating system? If you decide to build an operating system, you can kill a business and have business. And you've seen that happen. Say with Apple music and stuff like that, right? So there are like strategic decisions to be made, which affected the rest of the scaling of the organization. And at Uber, it was different. Uber deliberately chose not to take a dependence on any cloud provided at that point of time. So they had on on-prem data centers and everything was built in house and abstractions like Uber Nadeez, etc. hadn't come in there. So so everything was built in house that necessarily meant that you needed a big infrastructure team, right? To handle data centers to basically own the entire text act because we had to move fast. Uber created this concept of platforms and programs. So it's like, you know, there's horizontal platforms and programs were like teams which were working with the product team. Hey, write a growth. Make sure you're you optimize, write a sign up or drive a growth. You make sure you optimize driver sign up. These are all program teams that were built on top of the platform and teams, right? So the people for each one of these teams was different and the scaling requirements for each of these teams were required different too. So because the infrastructure space was different when the company was started and some of this was also like, you know, the CEO decisions, for example, Netflix, famously done so AWS, but the CEO of Uber decided I'm not going to take any dependence on any cloud provider kind of stuff. So that has an impact on the infrastructure space. So today infrastructure has changed a lot, right? You know, you have popular paradigms like Kubernetes and then you've got things like terraform and you've got containers and a lot of the stuff is automated. So you don't need as big infrastructure teams anymore as you needed from the decisions Uber actually took. I also seen some signals of people moving away from cloud providers saying, hey, I'm going to bring this stuff in house because the cost is too much, right? And so then again, that changes the calculation paradigm. Now with AI agents, there are a lot more decisions you have to make versus on build versus buy. What is a component that I need to build versus buy? So let's take specific things like our back, which is role-based access control infrastructure teams have to take care of it. There are solutions out there. Let's take about authentication. You have all zero octal, these solutions out there, right? Including say Sari. There are agent to get Sari solutions out there, right? So the leader of the infrastructure team has to now make a lot of decisions based on build versus buy. What are my core competencies that I actually need to build in house versus buy? Now, if you're a Google or a Facebook, you can build everything in house because it's slightly coupled in and you have big teams. But if you're not one of those companies, you have to make these decisions on the time. Where do I build versus where do I buy? And now the same thing has moved on to ML infrastructure. ML infrastructure kind of falls in the bit in between pure infrastructure and product teams. And the landscape of that has also changed. So like one Uber when it created Michelangelo, which is a popular ML platform. So it was one of the first to create a platform for its own internal use cases and then companies like Pinterest, Spotify, etc. Have also at BNB, etc. Have created their own ML platform. But now with LLMs, this question has opened up beyond for the ML in front teams. Your previous choices were either go with SageMaker or Rotex AI or build one something in house. And now with LLMs, they're saying, hey, do we need a platform? Do we need point solutions? Should we use something like based on our fireworks for LLMs influence or should we actually build in house? So again, these questions have opened up and you need a depth of expertise that most companies don't have because that's not their core competency. And so that has changed a lot and you will see a lot more third party solutions being adopted. Then the owners is on third party solutions or other companies and any scale obviously is one of them now to make sure you are reliable. You do the job that the user is wanting you to do and for the price that they're willing to pay, right? So the infrastructure landscape has changed and that has affected the scaling landscape and the world of AI is an AI agents is fundamentally ordering it. We had yet to see the full impact. Just beginning on that aspect. My next question was going to be sort of, it said, where do things go from here? Where are the active solutions or changing patterns emerging right now? Or the big challenge is that people are facing with some of the implementation right now. And then maybe the next challenge on the horizon or the next opportunity on the horizon that people are shifting to or converging around. Let me take that question from an MLA. I see this and go better than that. And the MLA AI space like I mentioned, there are companies which want point solutions for there. I want to build a customer chat board and I want to either call into OpenAI or I want to host our open source model and there are solutions to it.
actually do that. There are some other forward looking ML companies which are looking like, hey, I just don't want to point solution. I have my recommendation model that the marketing team is using. At the same time on something like LLMs to be served well also, I need one solution to rule them all. And that's where Ray comes in. And Ray being open source really, helpful because there's no vendor locking to that. So we are seeing a lot more emergence of this pattern of enterprises adopting LLMs and thinking more about what should their ML strategy be. They're putting a lot more thought into this than they were doing two years back, for example. Then what happens in the scaling phase is I do believe in like, you'll see a billion dollar companies with just five engineers. I think that's going to happen because so much of the stuff is automated with the agents. It's so easy to get your website up and running and you can get a full stack application done in two days. And now with the models of getting better and engineers are also getting better of understanding how to use these models. So that is not just AI Slop, but knowing what kind of specifications to give the model, how to use multiple LLMs to do code reviews, etc. kind of stuff. So the scaling challenge is completely changes. Now it's not just adding engineers. The key aspect that it goes to is like knowing what to build because previously it was like the amount of time it required to build was a long pull. Even if someone had an idea, it's like, oh my goodness, I have to hire not 13 engineers to do this. How do I sell them on the vision and stuff? Right. Now, that's changing a lot at changing fast. And you will see in the next one or two years, a lot more of smaller teams, smaller companies producing a lot more. And we are just starting to see this change. In fact, all the agents are more adopted by startups and entrepreneurs, not by existing companies as such. And over the next couple of years, you will see more of that actually happening. So that changes the infrastructure landscape too. There will be so much more stuff produced by agents where do they get hosted? What happens to the reliability of it? What happens to the scalability of it? Right. So there are come. So you will see it because the output will increase so much. This dynamic, I think, is so interesting. The big scaling challenge being knowing what to build. And I think like particularly when you're talking about when you're mentioning sort of early, like the lead time inherent in infrastructure teams trying to understand like the revenue, how that scaling, and then how to work that backwards towards what you're doing scaling from a director of infrastructure perspective. How does somebody who maybe is like owning infrastructure as a capability start to wrestle with that the internal question that's going on around knowing what to build maybe the bottleneck in that being the right thing? Like how does somebody kind of wrestle with that particular problem? But obviously, you know, you could say, I would know what to build. I just have to scale this kind of stuff. And I'll just do my job and go home and I'll be okay. Great stuff. I think now the you can continue doing that. There's nothing wrong with it. But looking from the angle of like, you know, person who'd be more productive or more effective is like, who is able to sharpen their skill of knowing what to build? I do believe all engineers will need to become like product managers or not become in the sense. What I mean is that need to have product thinking. The layer of the stack that you actually build. Why are you building this? Who's my customer? What is a pain point I'm drawing to address? I'm just not going to listen just because some product managers told me he built this kind of thing, right? Just it was true all this. But the fact that it takes the amount of time to build is going to keep it at using. I think we all need to train our minds to be able to put ourselves in the shoes of the customer and understand the product dynamics and understand where the company is going for and try to push for that clarity. Well, I think that point is so it's interesting because I've seen that pop up in a couple different scenarios. So we had the CTO of Brex. They did this thing. They were doing Brex 3.0 and they essentially isolated their team. Did a hacker house for three months to rewrite Brex from an AI native perspective. Brex was an AI native product. What would they build? And one of the things they had talked about was how many times they had thrown away the entire code base and just rewrote it from scratch. And I've also sort of heard patterns where people will pursue three solutions in parallel and just toss it out from scratch and go and pursue multiple ways to get to the same outcome. And I think that parallel build model is really interesting. Yeah. And I think that the coding agents just make it much, much more easier to do that. One person can launch three agents to a trade on three prototypes or three initial steps to see that you know, hey, let's see which one has legs kind of stuff. Right? Previously, there was a lot more times spent in design thinking, hey, how do I design this center for whiteboard? Now this time, it'd be like, code has become cheap, especially prototype code has become cheap. So you can, instead of having debates in a room, you can like build some foundation building blocks and that either that will clarify the questions you have any overhead. Hey, is this the right thing to build? Is this going to build? Is this going to be too complicated? Or it'll give you enough data points to ask the questions with your customers kind of stuff. Hey, is this the right angle kind of thing? Right? So that face has shrunk. And that's actually exciting. So I want to kind of capture a few insights here because I think like this is describing everything a really interesting build environment for people. And then my question is going to be like, what's the infrastructure strategy or patterns to kind of navigate this environment or like any scale's point of view on how to navigate this bill? Like so far, like what you've kind of shared is that like the scaling challenge is knowing what to build more stuff is going to be produced than ever. So there's all these other downstream questions that have to be asked a lot faster rate. And it's complete. It's so much easier rebuilt rebuild the system from scratch. So you can do that and that changes the patterns for for infrastructure and your prototype code is it's cheap and engineering managers and people in infrastructure need to like understand product thinking and being engaged with that to like be able to anticipate all of the different downstream needs. So I'm just thinking about like that environment. So like well fast forward nine months now like what so then what are sort of the infrastructure strategy or patterns to navigate this type of environment? Or I guess your perspective at any scale or maybe like your bet for how people will be dealing with this in the future. Follow me. The best analogy is Android and iOS. So before that, you know, people used to create this job on mid-lit kind of stuff. You know, you had this thyline web browser kind of stuff and you know, the carriers used to get key when you could download ringtones and stuff like that. And then the smartphones came around for every app. I remember Google search and Google search had to be tested on like 50 different phones kind of stuff. So kind of repeating the same problem. And then a platform got created and Android and iOS which allowed app developers to focus on one thing. And then now this I saw the same patterns at Uber with machine learning for every single machine learning problem. You had to go and solve the ETA problem. But you had to go solve the pricing problem and you're repeating the same things, right? And that's where ML platforms came up. I do see that as more and more agents come in, there will be similar patterns that will come out and you would need platforms to solve that problem. Obviously, observability is one of them and there are a few companies doing that. The key thing for me is more around decision making context. Now, you will have an agent that is deployed in your enterprise and it's making some decisions without human involvement. And it has led to some outcome say in nine out of 10 cases, it's great outcome. In the 10th case, it did not. So you need to have a log of like, you know, a context graph of why these decisions were being made by that agent. Now, suppose you have multiple such agents running. Why did the agent make such a decision? What was the inputs to that agent point that particular point of time? Right? Because one is just audit, but one for like previously humans made those decisions. So you had something in your head, some return rules, then you also used your own context or whatever you're, you know, if you think of your head as an LLM, your brain as an LLM, you had some context and you made a decision based on that. And now agents are doing that. And so the CEO's, if the CEO asked, "Hey, why did we do this? This customer is complaining. Why did we do this?" And you should be able to explain it internally. Now, let's take it. That's from a product facing perspective. Now, let's take agents internally. You have an SRE agent that does it and it does magically did that. And people are like, "Oh, happy, but then you also want to understand why it happened so that you don't change something and that you can no longer make it." So I totally see complete platform development building over here with, and some of these are already getting built. Like, you know, companies like Gleene, have it, and the you can give some auditability. I'm sure Microsoft also does, but I can totally see a platform coming and over here. So agents will expose this new pattern that some are happening with startup. Some we don't know because they have not been adopted at scale. Obviously, enterprises will require compliance and all those things and enterprise readiness that we need to provide, but there will be newer patterns that are actually emerging just because a lot of the work was previously done by humans. And we could go ask a person, "Hey, why did you do that?" And the person will explain this is what happened. And then you could put systems and processes for the person to follow. Some folks use context graphs as the dumb for this. I'm that I mentioned, but that's a good example of some infrastructure stuff that needs to be built in an agent world. I think that is such an incredible example because I'm now starting to try to imagine a world in which say 10,000 agents are simultaneously working in all different types of this system. And if they're all sort of operating within a style, and I think about then like, you know, some of the patterns of like general operating patterns of like documentation is oftentimes like the biggest bottleneck for people to move fast. I'm taking sort of like those human patterns of like documentation and communication as oftentimes bottlenecks like you only move as fast as like the organization can learn and communicate and document those types of things. And I'm taking that like the agent levels like if you don't have like whatever that sort of input is, then I could just see so many collisions like oftentimes like a bad organization operating without context, without communication, without all those things, how that's ultimately slowed it down. And then if you have all these sort of like infinitely applied agents all colliding together tweaking things madness. And so that here's an interesting part in my thing. The other thing that'll change is that data, the source of truth of data, like who has it? Where does it build template? We build platforms around that so that agents can act on it. So let's take an example of rippling. Rippling is pretty popular HR and finance software. Their thesis was that employee
data is a source of truth. So employee data means when did a person join? What is a salary everything? And before that, and even now this is usually spread out. There's one spreadsheet finance team will have one spreadsheet HR will have one spreadsheet the managers will have one spreadsheet the product managers will have kind of stuff and none of these spreadsheets will ever match. What is the source of truth? The core source of truth is the name of the person when did they join salary for HR purposes? What are the projects they are working on etc. And so they built a black com around it. So when you have these multiple agents working on it, they all need one source of truth of data. And so there will be platforms emerging so which will allow these agents to actually work on it. I actually see that also key and some of the organizations that it's internal external start building platforms around that. The context graph and then a core source of truth that multiple agents can work off of super awful. That's great. JK, I want to make sure we get into some of the cross-functional leadership elements because there's been a couple phrases that you've shared that I think illustrate some of the dynamics of where I think engineering leaders need to continue to build capabilities and skills. So for example, when you're talking about mapping revenue projections to your infrastructure strategy, I think that represents sort of the broad competency that engineering leaders are being pushed into. You're being increasingly called to be responsible beyond just software delivery, but you have to have a broader sense of the business. You have to have a broader sense of product. You have to have a broader sense of the connection to sales. And all of this are things that you have done in different capacities at any scale as serving in a GM capacity and leading a broad cross-functional group. So I think I wanted to dive into that experience a little bit more and start to pull off some of the lessons there. So what was it like to sort of lead this more broad cross-functional group of things like product engineering sales, solutions, architecture, or what was that like? It was hectic, but it was also fun. It also gave me a full picture of the business in the sense that taking the first sales customer call and then working with the field engineers, this is ML infrastructure. So you have to go deep into the code of the customer kind of stuff to make it work and then bringing it back into product engineering. Obviously early stage founders, early first few employees, everyone does that. But at any scale, it was a bit different because we already had rain open source and we were getting a bit siloed between the field team and the engineering team kind of thing. So I deliberately took on this role to break down those silos. This was more around many times engineering teams built in an ivory tower. They will say are the customers stupid? They are or our field team is stupid or these are sales team is not selling our product. Well, look is this beautiful product we have actually built kind of stuff. And the sales team will always come to know our engineering team is in their ivory tower. They don't know what's happening on the ground kind of stuff. I need to meet my sales code as well. I just want to get my customer to say yes to this kind of thing and get a lot of money. Many organizations fail in that stage. Obviously I'm not saying this is not this is usually not a problem in your team of 10 people. Usually starts happening when you have like 50 hundred people. Small groups have been created. You have passed the initial stage of revenue. And that is critical critical stage. And this gave me two things. One is obviously an end to end perspective. And it was not just like me doing on the ground. This was a company level thing. I was in our executive ship meetings presenting their end to end stuff. I was do like you know, hey, this is the customer. These are the problems are doing. And we had a scorecard that I was talking at the all hands of every week to make sure that our company was aware of this ring. Culturally it cost two things. One is the silos got broken down. So engineering really understood it and sales really understood that hey, there's an engineering leader who hears us. That was super important. And the second thing was that you know, given our product, our customers were ML researchers, engineers and trap people. They also know I can talk to a person who knows my pain who can feel my pain and this will get fixed. It also changed our engineering culture where engineers felt that you know customer success is key. And it to this day it happens to this day. It happens where if there's a customer problems engineers jump on it and I don't have to ever ping any one kind of thing. That's because it's kind of built into the culture of this company. But the meta point it we are under all this is just that the engineering leaders, especially once you grow beyond every level. And in fact, at every level, if you do this, you'll just become stronger is just understanding the end of the business and not just being siloed in your area kind of stuff and understanding the why behind the feature is being used. How do you speak the sales language? How do you explain to the field engineers what the constraints are on engineering side and just helping break down the silos kind of stuff. So many times I said, here our job is to serve the sales team, right? Because at that stage, we needed it, right? And so you know, that really helped. I wanted to have into some of the mechanisms that created that outcome. Because I think like that cultural shift of to have the organization understand that customer success is key. And in one of the outcomes being that you need to understand the end of the business, like what were like the interventions or the tactics that helped help make those things true. Yeah. So we made it to the game. So it was basically, you know, we call this thing called six 20. So in six months, make 20 customers successful kind of stuff. So we had a scorecard and there was like, you know, monitors across all the office area of the scorecard and stuff like that. It made it front end center and making it and, you know, the metrics associated with what success means and talking about the customer. And even when a customer, we talked about a customer win the all hands. We brought the salesperson, the field engineer and engineer together to talk about their aspects. So it was not like just sales talking with the customer engineer, talking about it was the trio, which was actually talking in the all hands. So that was extremely important because it was like, okay, we are doing this as a team to make this customer successful kind of stuff. Well, that's incredibly because that shows like the whole cross functional workflow. Exactly. Can you have a, do you have a story of a moment that I really stood out to you like, oh, like this is shifting the culture is changing. Yeah. Yeah. So we had this customer called Suarez. They use the robots, computer vision robots to see where this sewer pipes are actually blocked and they were using the way any skill kind of stuff. And they were adopting the any skill. And one of our engineers literally wrote the sewer code for them. Hey, this is how you should get started. And that became a meme and we started doing and calling it doing a sewer AI and a Slack channel got started doing a sewer AI. So that pattern got established. So every time it's over, we are going to do a sewer AI kind of thing. And so that was the thing when I say, okay, this has caught on on the during grounds of itself. There was like a timeline. So usually these projects, etc. would you do cross functional stuff. There has to be a rallying cry. There is a, it's either a name of the project or clear, Chris goals or like, you know, something that the people can latch on to emotionally. So I'll give you a couple of other examples. We had the same thing and this is purely engineering stuff, which was about like, you know, we had a bunch of test failures, etc. And it was like just grinding work. People are like, oh my goodness, I have to do that. Fix some other code against us. So we kind of made it to leaderboard. We basically gave me five time. I had another thing called project oxygen, which is we had to again pump up hiring across the whole company. So I said, hey, this is project oxygen. This is our goal because this is pumping oxygen into the company. Same thing from Uber days, even after all the backlash that happened in Uber in 2017, 2018, we had this thing called project 180. So in 180 days, we are going to switch driver sentiment. And there was specific projects, six chapters created kind of thing. So those, these are all cross functional projects. And just giving a theme that people can resonate with because otherwise you'll just lose them in the day to day work and some other stuff will come in and this will get ground out. These are some patterns, which really work. Well, it's going to ask like are these initiatives like were they self-organized of like we identified this thing come up and then we went after it or like, I think usually the cross functional stuff are not self-organized because it is cross functional for a reason, right? You know, people will follow the status code by default. So cross functional stuff needs to come from some leader. And it can be leader from anywhere. But usually cross functional stuff needs to also make sure that the CEO and fully agreement with the priority kind of stuff. But it's not the CEO, the VP or who it is kind of the leader has to be fully aligned. Otherwise, you know, it is this moment. So yes, you have to get the buy-in of who is the executive, depending on the size of the company, the size of the team that you're actually doing. So then so I'll go back to the sort of the labeling of six 20. So in six months, make 20 customers happy. So mechanically then was then like the process of this being like, we have these customers cross functional group go make make though these specific customers happy or like I guess what was I guess what was the initiation of this? Yeah. So it was it was more like, you know, hey, it is not we didn't have the 20 customers to to be honest. So there was a little bit of building up the pipeline also for this. It was about making these customers happy. So I was running the stand-ups initially. The leader who's doing a hast to do the hands-on work, then you kind of really know where the bottlenecks are and then you can let your managers or other people take care of it kind of thing. So yes, it was daily stand-ups getting the right people in the room chasing down every single issue. Actually, this is a lot of block and tackle work and that's what it's needed for scaling the hey, they said this thing is not working when it's going to get done by and the engineer will like, oh yeah, he get done by two weeks. No, no, this is needed because this particular customer is blocked by this and you have to do it tomorrow. And because this is higher priority, right? And that usually comes when you have this top level initiative like six 20. Otherwise, people are a lot drawn in the work. That clarity, the sense that clarity, this is important because this is the top level initiative right now. Your other work can actually wait. Engineers on the ground or anyone on that ground does not usually that clarity gets lost in the communication kind of stuff. And then hey, then you could get back to the customer saying it's done and you can automatically see that usage go up. So we had metrics on that usage and whether it's going up or not kind of stuff and then whenever customer stuff is always during the bell and the real tarol and then share the learning so that and then we had actually a doc which said that what are they using it for? What is a pattern we can actually do so that when we get from 20 to 50, we cannot be this hands on and we need to make sure we have some systems which actually skip. I'm I'm extracting a couple of the patterns here because I think, you know, having this this initiative be like a like a priority initiative 626, it's been 20 customers happy as the leader and kind of core decision maker who has kind of the authority or in this case, like the decision making authority to unblock
or a really drive initiative there, like it be really hands on with this type of initiative. And then from there, like the cross function work because it thinks and get lost, like you need to play a role in organizing and making that happen. And it helps when you're in that sort of role where you sort of like have responsibility over a couple of different areas. And then the rally cry to me almost seems like an artifact of like the culture shifting. Is it like as soon as people start to have labels and names and branding around these initiatives, it's like, okay, cool. That's like when you're starting to see things shift, that's correct. That's absolutely spot on. So those are good to take out of this. Yeah. So I'm just mentioning sort of like the collaborative communication piece and there's a few other tactics or leadership practices that I wanted to get into here. So a few things that you and I were kind of talking about in a different conversation was this perspective that strategy or execution failures can trace back to decision making and judgment failures. I was wondering if we would unpack that a little bit. So like, what does it mean? What does that look like and bring us into that a little bit? Yeah. So usually what happens is that there's a roadmap meeting that actually happens. You know, so what is it? It's once a lot of work producing the roadmap. It's a document. 10 people are there. They discuss one smart person in the room, comment something. The person say, yeah, I good point. I'll add that and everyone gives a thumbs up, moves on. First two weeks, things are fine. Four weeks, some struggles start showing up, some project slips. They report, oh yeah, this project is slipping, but we are making the changes everything is good. Week number six or seven, most things start slipping. And then something else comes in, some other project comes in and this kind of project just loses momentum. And by month to month three, this kind of, either this project is on life support or not going as well as needed, etc. kind of stuff, right? This is a pattern we see in every company. I've seen in every company in a pretty sure many people will resonate with this example. I feel like so many people listen to this, we feel like yeah, they're probably in a project right now or that's the exact pattern. And then we will say, oh yeah, this was some people will say, oh, this was an execution failure. The team did not execute properly kind of stuff, right? And some folks will say, no, this is not something changed. The environment's changed, person X, left, person Y, or some priorities change kind of stuff. But I would actually trace all of this to the decisions that were actually taken at the start of the meeting during the roadmap planning or during the strategy discussion. In fact, I actually saw someone describing this very same thing on LinkedIn, the truly last week, and it's like a perfect kind of thing. I'm pretty sure your audience will also say, yes, this is, I totally didn't need to do this. So that's the decision making point, I was referring to. And you have clarity on why you're doing something. And that means asking the hard questions, why should we not do this thing? What happens if we don't do it? What is it that we are trying to solve? So if you start a project, suppose you're improving the funnel of some sign up funnel, and people will say, oh, I do project X project, I'll ship these three features. And they'll ship these three features and nothing happens. Because they say project was successful. You ask me to three, ship three features. I shipped it great. Here's my pro, I want my pro, why am I not getting my pro motion? But the intent was forgotten. The intent was like, I have to change the funnel dynamics. So do I continue? So that basically means you ship those three projects, and the funnel did not move, is either you did not understand the users, or these are the wrong features to build, or something else happened, which completely destroyed the advantages that these three features gave in. So this is the intent. Why are you doing it? So that decision making clarity around, why are you doing it? What is the advantage your product is going to get? What is the customer pain point you're actually solving for? How is your product differentiated? And what is the market map? That depth of clarity and decision making, is absolutely key. And the people who are good at it, or the companies who are good at it, will make better decisions, and will have higher chance of success in completion of projects. Well, we're all smart. No one says a project failed because I was not smart. I have never heard someone say this, including me, I am also in the same camp. It's not that you or me are special over here. We are all human beings. We'll never say, oh, because I was not smart. And my team was smart. Just things happened outside our control. It could very well be that the team was not involved in the decision making, and they were just asked to do it. But the key point is that it's around how you think about products, right? You want to build what not to build? Where is the market going? Is this the right thing to build? Those, you need to spend, have those uncomfortable discussions in that first kickoff meeting, it's a, and that is absolutely key to execution. In my opinion, executions, people, it's easy. The strategists who are involved in this strategy will always point execution. Our strategy is great. Execution was problem. Execution people are the people who generally don't have voice because the strategy people are that, and those CEOs and the VPs of the execs kind of stuff, right? And they will feel, yeah, but there are very few people who push back and say, no, you're strategist, you're wrong. This was wrong. We should not have done this kind of stuff. They know it, but they don't have the full context to be able to clearly articulate and debate that kind of stuff. This is where your engineers, product managers, come into the picture to have those debates and have that clarity. I think it's great perspective. The other question I wanted to ask here is I wanted to deconstruct this practice that you were talking to me about, where you use, I guess, an internal individual leadership practice of using AI tools to analyze strategic documents and identify failure points in the organization. Bring us into that practice. When do you do it? How do you do it? And how are you using this to sort of analyze some of the decision-making challenges or systemic failures that you're identifying in an organization? I think I use, yeah, this is a risk of it here. I use AI more like an assistant because the risk is, when we talked about clarity, decision-making kind of stuff, the risk is like, oh, I'll just ask Chad G.P. what to do. Hey, it's my strategy dog. Ask Chad G.P. what to do. And that does not give you the clarity of decision-making that is needed, right? You need to have a point of view that you have strongly thought about why, et cetera. And this does not come in one day or one hour. It is like deeply understanding the business and your customer's kind of stuff. And then obviously you can use AI as your assistant to saying, hey, what am I missing? Is there a gap? Kind of thing. Now, that's, I do want to call out because, especially in today's world people will say, oh, yeah, that's easy to do. I'll just ask AI to do it. And the second part is how I personally use AI is like, many times I get a dog soft system, I do not know. Whether it's even engineering, design, document, kind of stuff, right? And then I can use AI to say, hey, analyze this code base, give me an understanding and I can ask questions about this. Does this make sense? And here's this design dog. I have this thing explained to me because I haven't gone deeper into it. The same thing can be done from a product management perspective, which is like, if you have a strategy dog, you can ask it questions to help you better understand. If I give me a data point from some customer here, which allows me to either refute this argument or not kind of stuff. So this comes back to what we talked about initially as to how with agents, right? You know, this is where the internal, for your AI agent running it enterprise to be able to give you good answers, it has to have the knowledge base of your context of your customer data kind of thing. And that's where they actually help you. It's like a tool to give you clarity. I'll be put in this way. You want to unscrew something. You are using a screwdriver to unscrew it, but you know you want to unscrew this. I use AI tools in that way. I know I want to do this, rather than saying, I do not do, "Hey, you do it and give me the answer." Kind of stuff. There's an intention. There's a very clear intent. There's clarity in my head as to what I want to do. And I need the best tool for this job. So that's the best way I would say this with AI agents for analyzing strategy dogs or filling in the missing pieces, use it as a sidekick, not as your primary mechanism. I think it's really powerful distinctions there. It turns out like, to me, I can immediately extrapolate that. So the tactical questions that I'm asking is like, it's a bunch of different frame of frame of reference. I think it's great. Take it. We've got some ratifier questions. If you're ready to jump in. First question. What are you reading or listening to right now? I'm reading a book called Explorers' Gene. And this is by Alan Hutchinson. And it just talks about why do humans explore, why don't other humans explore kind of stuff. It talks a little bit about human civilization, day-to-day decision-making. Like, why do we want to get out of that comfort zone? I haven't finished yet yet, but it's fascinating to me. I love that. Number two, what is a tool or methodology that's had a big impact on you? It's more about being intentional about your calendar kind of stuff. For example, I, this is just me. I sometimes overthink. I ruminate a lot. But like, hey, this is always constant background thing that keeps happening, right? So the method and some days, days like commuting into San Francisco on the train are the days I've blocked off on my calendar, like, you know, right to your docs kind of stuff there. So that has kind of helped me saying, okay, now it's not the time to think. We can do a deep dive on, like, Tuesday evening, then during commute time, kind of thing. I love that. Okay. What is a trend you're seeing or following that's interesting or hasn't hit the mainstream yet? The trend is like, you know, we talked about agents, etc. for coding. So this is the thing. This is already in mainstream and Bay Area, but it's not hit mainstream in rest of the world. And I spent four years in New Zealand. I have seen that it is not hit the mainstream there at all. So how all the stuff we talked about agents, kind of stuff will change how we do things. Bay Area, Bay Area, that's a solid mainstream. No, it's not at all mainstream kind of stuff. So the whole software engineering, engineering field and general technology field and all the associated fields are going to change a lot. And so that is something, yeah, like I said, I'm a little bit uncomfortably excited. There's new, the good things, bad things. It's quite a big societal impact. Yeah, it's just when you extend it to like, say, every single software engineer in the world in every geographical location, running 15 parallel agents simultaneously. It's like, it's one of the things where it's like, you know, they see humans are horrible at estimating like long time horizons and like exponential curves and stuff. That's one of those things that breaks my brain and the more I try to get to that point. So I love that final question, JK. Is there a quote or mantra you live by or a quote that's been resonating with you right now? I think it's more like, you know, all this stuff and especially working with the M and M kind of stuff. At least this is me play long term games with long term people don't chase the new hotness. Whatever it is like, you know, company X is hot today. A month later, company Y is hot and there are some people, especially in barriers, very easy to fall into this four more trap in the technology industry kind of stuff. And the four years I spent in New Zealand was really, really helpful in that sense because I used to pick up my daughter from school for four years, not once. I was asked what I do for work, not once.
by anyone other parents. And I totally like, you know, you're defined by your work in the media. So extremely hard what I said, I don't think, but I truly practice it enough as much, but that's that's something that's always in the back of the band. Play a long-term game, so it's long-term people. I feel like that is such a powerful thing that people need to hear right now. I'm going to make a sticky note, I'm going to travel that would forever. Jakey, there's been an incredible conversation. I mean, being able to deconstruct the patterns of scaling, the patterns of infrastructure and how things are shifting, all the way down to cross-functional leadership in some of the different ways that you're helping make better decisions or drive better decision-making in your organization, just a ton of fun. So thank you. Thank you so much for having us. Fun, let's make that new Mexico trip and I want to try out your cooking. I love it. I love it. I'll send you some recipes. I'll send you some favorite stuff. If you're listening to this and you're wondering, how can I connect with other engineering leaders in my city? Pull up your phone right now and go to elc.community, click our chapters page. You can see that on the menu on the left. Find your local chapter and click join. We're hosting virtual and in-person events all the time and this is the best way to help you get involved. Expand your network in your city and support your leadership and career growth. So pull up your phone, head to elc.community, join your local chapter and get involved. A huge thank you to all of our local leaders who make community happen and thank you for listening to the engineering leadership podcast. (upbeat music)
Podcast Summary
Key Points:
The primary constraint for CTOs has shifted from speed of building to deciding what is worth building, emphasizing product thinking across all engineering levels.
Scaling patterns are highly context-dependent, varying by company culture, domain, and business model; there is no one-size-fits-all formula.
In the AI era, infrastructure trends focus on auditability for non-human workflows and managing organizational complexity from thousands of independent agents.
Engineering leaders must align technical decisions with business goals, such as revenue projections and customer needs, and prioritize funding for P0 initiatives.
Modern infrastructure (e.g., Kubernetes, containers) reduces the need for large teams, but decisions like cloud dependency are strategic and impact scaling.
Effective leadership involves educating executives on technical trade-offs and consequences, especially for reliability and scalability.
Summary:
The transcript features a discussion on engineering leadership and product thinking in the AI era, hosted by Patrick Gallagher with guest Jai Kumar Ganesh, Head of Engineering at Anyscale. Key insights include the shift from "how fast can we build" to "what is worth building," requiring product thinking across all engineering layers. Scaling patterns are context-dependent, varying by company culture and business model, as seen with Uber's "move fast" approach versus Anyscale's thoughtful scaling for ML infrastructure.
Ganesh emphasizes that leaders must align with strategic goals, such as revenue projections, and prioritize P0 initiatives without "peanut buttering" resources. For infrastructure directors, core responsibilities like reliability and scalability require proactive planning, including educating executives on trade-offs. The conversation also covers infrastructure trends: the need for auditability in non-human workflows and managing complexity from autonomous agents.
, cloud dependency) remain critical. Ganesh advises leaders to ask the right questions, automate where possible, and focus on customer pain points to drive meaningful innovation. The episode highlights cross-functional leadership to break silos between engineering, sales, and other functions, aligning technical teams with revenue and customer success.
FAQs
The primary constraint has shifted from how fast can we build to what is actually worth building.
He emphasizes understanding your customers and asking the right questions, such as why you are building something and who the customer is, rather than just following product manager instructions.
Ensuring auditability in non-human workflows and preventing organizational madness as thousands of independent agents operate within the enterprise.
It prevents cluster drift, making upgrades non-events, and secures systems by default with only 50 binaries, no shell, and no SSH.
It requires thoughtful scaling with a focus on rock-solid core components, as ML infrastructure moves slower and needs careful foundation building.
They should ask about revenue growth projections to back-calculate scaling needs, and educate leadership on the consequences of technical debt.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.