Go back

FRAMEWORK FOCUS: Building AI Products with Rami Sayar

from FRAMEWORK

0m 0s

FRAMEWORK FOCUS: Building AI Products with Rami Sayar

In this episode of the Framework podcast, host Speaker 2 interviews Remy, a venture partner and engineering leader at Microsoft AI, about the evolving landscape of AI, search, and agentic applications. Remy shares his journey from software engineering at McGill to leading major AI features like Copilot Search and Bing Chat. He explains that while business fundamentals remain unchanged, the LLM era has raised the bar for user expectations, requiring founders to solve specific problems exceptionally well rather than relying on generalist models. A central theme is context engineering, where domain expertise helps systems fetch the right data to meet user intent. Remy emphasizes that startups need their own evaluation loops and proprietary data to build defensibility, since open source models are widely accessible. He also discusses the balance between human-in-the-loop and autonomous systems, noting it depends on the vertical and latency constraints. Founders must translate technical value into customer language, especially in non-technical industries. Finally, Remy highlights the productivity gains from running models locally and predicts that edge devices will increasingly host AI capabilities by 2026.

Transcription

7100 Words, 38885 Characters

English
0:00 Speaker 1 The business fundamentals probably haven't changed since the LLM era, but I would say that the LLM error has accentuated or has increased the bar for what the user can do what the user can do. And the the the companies that are had this nailed down free LM era are even more valuable today, especially if they're able to leverage the technology. 0:23 Speaker 2 Yeah, and this is kind of an interesting thought. It just has. It's almost like both the floor and the ceiling have jumped. It's just the floor has jumped a little bit more than the ceiling. So it's like that middle gap. Yeah, gets that's sort of where you have to play and that's why you have to show show value. 0:41 Hey everyone, welcome to another episode of the Framework podcast. I'm joined today by our venture partner and advisor, Remy. I'm glad you can do this in person given our, our Seattle ties. But today we'll, we'll want to run through sort of Rammy's experience and kind of this thought leadership around search and AI agenti applications, things like that and, and sort of where he sees the future in the space. 0:59 And so with that, welcome Remy and, and for any new listeners, maybe if you want to give a bit of background on, on yourself and, and kind of your career journey. 1:06 Speaker 1 Yeah, absolutely. So thank you for very much for having me. Thinking about my background, I'm a software engineer by training. So I I went to McGill University in Montreal, hence the connection to to Canada and to framework. And for the last 15 years, my career has really been focused on the intersection of AI design, startups and the web. 1:26 I currently lead an engineering organization within Microsoft AI's generative AI group. I oversee engineering managers and technical leads responsible for building some of the major features that you may or may not have heard of on Bing Com and Microsoft Copilot. And so some of the projects that I've had over the last six years included deep Search, generative search, Bing Chat, Ding stories, and of course our most recent project, Copilot Search. 1:51 Well before this, I worked quite a bit on machine learning and IoT solutions for Microsoft's global customers, spent some time in New York doing that. And before that, I spent four years as a senior technical evangelist at Microsoft Canada, speaking internationally, helping developers to kind of work with open source on Azure and really working with Startup Thunder is to help them build and scale up as they grew on Microsoft platforms. 2:14 Speaker 2 Amazing. And maybe starting on that search side, I mean you've sort of shipped like you mentioned sort of different experiences on a massive scale, like both Geo and user type. And so just curious how you seen kind of the front end user experience evolve and kind of that willingness to engage with AI, whether it's sort of abstracted behind the scenes or or sort of in their face from a, a search perspective. 2:35 Just curious on on how things have evolved. 2:37 Speaker 1 Yeah. I mean, the way to think about it is that we've kind of left from keyword search to more of this kind of chat interactive experience that you're seeing more and more of on search today. It's been an interesting challenge because one of the things that happens in a very large scale search system is that you're really interfacing with an AI system at the end of the day. 2:58 Like a user is mostly giving in a query and input. And then there is a very large distributed AI system that runs in the background and tries to understand that query intent, tries to get the right sources, and really just builds an answer or provides a set of results to your query. 3:13 So there's a lot of jobs to be done in this kind of user experience. And that's where AI has been really interesting in helping advance us from a user experience perspective, but also from a quality perspective of the results as well. Search is a great example. So maybe what we can do is kind of deep dive into how that's relevant for founders. 3:33 What is the hardest part of making AI maybe actually useful for end user? I think is a question that I'd love to tap on based on my experience for the founders that are building AI products. Ultimately, in my point of view, you should be serving multiple jobs that users need to do with your product. 3:52 Users do come with an expectation of what your product should be solving for them. I mean, that's why they sign up, that's why they try it. In our case, users are coming to bing.com because they they are expecting to get an answer to a question or to find a resource. So it's a great example of a complicated product where there's lots of user expectations and in general, users don't come a search engine with a single intent. 4:14 They in this span of like maybe one session of searching, they might want to try to get a recise factual answer. Maybe sometimes they want to do some research or explore a topic or even sometimes be entertained or discover something that they didn't know to ask yet. And what you look at search and you look at a search page in particular, we're often trying to satisfy a lot of these jobs that they have with different parts of the page. 4:37 And sometimes these jobs do conflict with one another. And it's always a balancing act, especially when you're thinking of building products that are useful for end users to balance the multitude of jobs that they might have when they come to your page. And being able to kind of still build a cohesive product that lets all of these users that have different jobs that they're trying to complete actually get to the end. 4:59 And this makes it really challenging for founders because general purpose ML models that everybody's excited to talk about today, and I'm sure we're going to talk some more about this today. They are optimized for kind of the average. If you look at all those foundation models, they're, they're called foundation models because they're really solving a multitude of problems. 5:21 They're not specialized, They're not fine-tuned for a particular task necessarily. I mean, you can get fine-tuned versions of these models that solve something specific, but in general, at its base, it's solving for the average in the issue is that if you're trying to build a great product, so in my mind, a great product is solving a user's problem at that moment really, really well, right. 5:42 So if you're solving that problem really, really well and you're trying to do that with just a generalist model, oftentimes you're going to fall short. You're really going to fall short of their expectations, in which case they'll be very disappointed and leave. Or in some cases, you'll just fall short, but not so short that they abandon the page. 5:58 They just kind of are OK with whatever output or quality that they got. So a model is pretty good at solving a ton of different jobs. You can still frustrate the user. They may still miss their current job. So for example, in the search space, like maybe we over summarize when somebody wants sources and that kind of causes some issues or we go too deep instead, like and and they just really wanted a quick cancer. 6:23 So these are examples of like where we're adding friction by accident because we may not be like balancing all the user intents as well as we should. So the other thing that you got to keep in mind though, is that these systems are kind of dynamic. Users are still learning what these systems are for in most cases. 6:38 So especially when we're talking about AI today and these models and these startups are building AI products, users are coming in with a set of expectations, but they're also still learning from what your product can do for them. And as a result of that, I would say they're mental models are fluid, means the product has to constantly adapt without breaking their user expectations too much. 6:56 You have to remember that also some users are coming in with some previous training or in a certain space or maybe some previous experience. So like in search for example, the classic thing that we're always trying to balance is should we just show more blue links for this feature? 7:10 Speaker 2 Let them take the. 7:11 Speaker 1 Journey exactly or should we like actually show the the more richer experience, but that that, that problem and that kind of previous training that a user may have had also extends to like vertical SaaS products. Like if you're if you're trying to build this new AI CRM SaaS product, you have to remember some users are coming in with previous experience using Salesforce or other cream tools. 7:34 So balancing that is really hard. So you're always going to have to fight that urge. And in our case, it's fighting the urge to just add more blue links and that we know that blue links help with our metrics, but it doesn't necessarily help you build for the future. And as a founder, what you're trying to build is for the future vision that you believe is going to happen in the. 7:50 Speaker 2 Marketplace, that makes sense. And yeah, I mean, like I think you brought up an interesting point on kind of search, but more in a vertical AI standpoint. And I think context is also a big thing. I mean, I think we've seen a lot of products where you now have this on the front end sort of simple natural language interface for any user regardless of sort of technical ability can sort of ask and inquiry whether it's their CRM data or sort of an aggregation of data. 8:14 But I'm curious sort of the performance gap in terms of the sort of the top tier products that can do search sort of across your tech stack. How are you, how have you seen them sort of add sort of almost dynamically add contexts when a user searches? So if I have a natural language prompt and I ask sort of let's say sales context, how did I perform on these past 10 deals in the future, there might be contexts within my organization or with another salesperson or somewhere else in my data that I'm not including in that prompt. 8:43 So I'm just curious sort of how our how important is kind of that prompting piece both from either the users perspective or doing it in the back end, so you kind of keep that natural language at the front? 8:53 Speaker 1 That's a really great question. You're basically asking us to like, how do we read people's minds? Yeah, that's actively right. 8:59 Speaker 2 Because they don't always know the amount and the level of context you need to put in, but then also how to structure that's that's the best process by them. 9:06 Speaker 1 Yeah. And there's a whole effort and research that's happening both in the open source community and in the academic spaces. And even us, when we're building our systems generally, you can kind of categorize this as context engineering. It's like figuring out from all of the data that we have access to all of the grounding that we get from the public Internet. 9:27 Like what is the most relevant pieces of context to provide the model so that when we answer a user's query, it is a giving the answer that they're expecting without necessarily missing a piece of really important context. O if you're, if you're thinking about in the search space, like oftentimes we do have an, we do have like these experiences that are built in where users can follow up with the second question. 9:50 And from there, like if you look at how Copilot works, Copilot is very much a conversational product. You have the inbuilt user experience and expectation from the user that they can follow up to provide either more context or refine their initial query, right. 10:06 But when you're a vertical SaaS company and you're building a tool that lets users at a company complete a task, it gets a little bit more complicated, right? Because the task is typically part of a giant workflow and missing context in one part of the workflow fails the whole workflow or missing it at the end kind of causes people to have to go back and restart, right? 10:27 So this, it's, it's very, very challenging. And that's where I would say the most important piece is having problem domain expertise. As a founder, you know the problem space really well. You've know what pieces of information are critical of what steps and you're able to build the system. 10:44 So go fetch them and provide them as context for your model or for your experience so that the right pieces of information are available to create the output that is actually the most helpful. And that's where evaluation come into play. You should have your own metrics. And one of the things that I like to share in my newsletter, for example, is that yes, there are global benchmarks that, you know, are kind of showing how all of these open source or foundation models are improving against a leaderboard. 11:11 But ultimately, a company that is building a vertical ass AI product should be running their own benchmarking against these models and should be running their own evaluation loop for their product as well. And it should just be a systematic automated thing that's part of their build dev test cycle. 11:29 Speaker 2 That makes sense. I think that also brings up a key question on kind of the human in the loop versus sort of autonomous execution, because I think the the natural sentiment is I want to offer my clients fully autonomous, they never have to touch it, things like that. But a lot of the times and human in the loop will actually perform better, but you still get sort of that efficiency gain. 11:49 And so I'm just curious sort of do you have any thoughts on kind of how founders or or products people should sort of be thinking about not necessarily forcing fully autonomous, but also understanding with their vertical, with their workflow, with their end user is a human in the loop sort of agentic process a little bit makes a little bit more sense. 12:06 Speaker 1 Yeah, it's a great question. So like, what's the future of what's the future UX for these AI systems is effectively the question. The human in the loop is a fully autonomous, it's hard for me to say with absolute certainty like which one of these paradigms will be the dominant 1. And I know that that's not what you're asking me, so I'm going to give you an answer anyway. 12:26 I do think that it really depends on the problem space. Like you have some problem spaces and some businesses where latency is super critical. If you're in a factory space, for example, you're generally don't want to stop the production line, right? Like that is money, every minute that the production line is stopped, it's not being generated. 12:45 There's no revenue being generated, right? So for those situations, what you want to do is you want to perhaps maybe automate and have alerts when things are kind of outside of the bounds of what you expect. And then in that case is like a human loop is more of a supervisory monitoring role. 13:01 And that's actually how most factories work today. Like they have a lot of automated systems. They'll these automated systems when things are out of bounds will start triggering alert and that's when you'll see a human intervention. But in another problem domains you shouldn't be doing that. You don't have all latency constraint. 13:16 You're not expected to have something be complete like a workflow be complete in seconds like yeah, it's not a continuous cycle, right. It's not like a fully automated system anyway. And in there, in my mind, a human in the loop UX paradigm makes so much more sense. 13:34 And you should be building a software or system that lets folks jump in and steer these AI systems to the output that they believe should be there. I'll give you an example. Like I know a lot of engineers are very excited about the idea of vibe coding and they can kind of set it and forget it and have these coding agents just go do the work, right? 13:55 Like that's, but that's realistically not how things will actually be built for production. Let's be honest here. Like it's a fun project to be able to vibe code everything and have this AI model go and, and do these things. But ultimately the engineer is accountable for the code that gets produced and gets shipped to production, right? 14:11 It's something that is critical for the success of the business. So the idea that one of these UX paradigms for AI systems will be dominant in the future, I would say is not true. I think it will be specific to the vertical. So the problem domain, so the place that you're to the thing that you're building. 14:27 So we, we, we shouldn't as a as a community decide that like all of these startups need to have be overly prescriptive, we have to be open. So specifically the vertical and the product domain. 14:37 Speaker 2 That you're in, that makes sense. I mean also each vertical will have inherently different forms of raw data. Like if you're still at accounting or fintech examples, you have more sort of binary quantitative data. But then if you look at something else like sales or marketing, you're looking a lot of conversational data and a lot of sentiment and stuff like that. 14:55 And I'm sure that also sort of is a key consideration into that sort of human in the loop versus autonomous. 15:00 Speaker 1 Absolutely. There's a ton of examples of startups that have done really well, both automating the arts that are not, that are not worth having a lot of emphasis on. But then there's other parts of their product that very much is augmenting the existing users as part of their workflow. 15:18 So 100% agree with. 15:19 Speaker 2 That makes sense. And maybe taking it a step further into kind of that how founders are thinking about product differentiation, especially when they're competing against other tools that do similar sort of workflow automation. I'm just curious sort of where do you see the biggest pieces are the strongest pieces of product differentiation in more sort of agentic platforms or whether it's fully autonomous or human in the loop? 15:41 How would you say from a a product tech standpoint, you can be defensible and and the answer can also be that you can't and it's just the the feature of product shipping velocity makes it difficult to to actually truly differentiate. Just curious on on your. 15:55 Speaker 1 Take there. Yeah, this is such a tough question because if it's a rephrase your question, it's basically like how should early stage founders think about building defensibility into their business plan into their, into their, their effectively their startup idea, especially when the core models are like this AI capability is pretty much open source at this point. 16:14 Like anyone can pick up a fairly performant open source model, fine tune it for their use case and have a very strong system for their target market, right? It's interesting. It's a really tough question. And the other thing that also consider as part of this is that the cost of software development is also going down pretty fast. 16:31 We're all getting so much more productive with these coding and ALMS engineers are able to accomplish so much more in a shorter timeframe than timeframe than ever before, especially if you're a very strong senior engineer. What I see from our senior engineers with AI is that my God there's dev test loops are getting shorter and shorter and faster and faster. 16:49 I think ultimately what makes a startup successful today is still the same that what would have made a startup successful pre AI, well, pre this generation of AI could be specific, like AI has been around for a while. Yeah, yeah, yeah, yeah. 17:06 And ultimately it's the same thing. So it's like you need to have a deep understanding of the customer problem, deep understanding of the industry that you're in and not just having that understanding, you also need to have a core data piece that nobody else has. When I think of companies that are building in a certain vertical stats, I'm like OK, well what makes like what data are you using to build your AI system? 17:28 Speaker 2 Anybody else enriching it? 17:29 Speaker 1 And how are you enriching it? How are you building on top of it? And that's super critical, even more so today than it was in the past. Like in the past, you could have just competed on like the feature sets that you offered and how fast you were shipping, sure. But today that's not really as relevant as like how you're helping your customers with the AI tools that you have. 17:48 And for that, owning your own data, having your own fine tuning loops for even if you're taking open source models in, it's super important. And then there's the other side of this is like, can you close deals? Can you sell? There's lots of examples of fantastic startups with great technology leaders that just couldn't close the deal. 18:04 And ultimately the startups one in that space were the ones that could even if they're technology was inferior, right. So in today's day and age, when we when I think about defensibility for startups, ultimately it goes down to are you able to leverage open source models with your own data, fine tune them and take that value that you get from these from this AI system to your customers and get them to effectively adopt it fast enough to generate revenue so that you are ahead of the curve when it looks like when you look at your peers. 18:36 And that's really what defensibility is today. It's do you understand the customer problem so well that you're able to enrich a lot of the open source models with your own data and get that value delivered to customers fast enough that they're constantly willing to pay you for the next version to keep their subscription active? 18:53 Speaker 2 Yeah. No, I think it's a great point. And like that almost 2 step is 1, there is that enrichment piece, but then also it's then being able to communicate how is your enriched data, how is your enriched processes actually providing value to the to the end customer. And I think that's also might be sort of a bit of a struggle as being able to sort of if you're in an industry or a vertical that's not highly technical, like let's take insurance for example. 19:17 That doesn't necessarily mean that you have to build a non-technical product. It's just how well are you at abstracting that technicality away into the back end? So that if your end user is a broker, for example, who's not necessarily going to go under the hood or or be in the weeds on that stuff, how do you sort of abstract that so that they can actually feel the value and they don't get caught up in sort of the technical jargon or things like that? 19:41 And so I'm just curious sort of any advice to founders who are sort of selling into primarily sort of non-technical or historically sort of heavy industry even like your energy, oil, mining, things like that on how to sort of communicate the value of your products? 19:54 Speaker 1 You have to use the language of the customer. I think that's the key core thing. Like you have to understand your customer. You have to know how they, how they like, what language they use, how they describe the problems. Because when you're going in and trying to say, hey, listen, we've just upgraded to the latest. I don't know, let's say the latest. 20:10 Quinn model or the latest open AIG GPO access model? Yeah, that means nothing to them unless you can translate that to specifically what it means for in your case, like the example we gave of energy, like what it means for energy. That's that's the key thing. You have to speak their language vertically translate. 20:25 Yeah, yes, it is awesome that you been able to take a new open source model that came out, fine tune it for your, for your sector and your industry, your target market. But like, unless you can explain what that actually means from a value perspective to your customers, you're just not going to it's not going to make this slash that you expect. 20:44 Yeah. And it's hard for us. I'm going to technologist, right? Like I'm an engineer. I love the tech, right? I love the fact that like, Oh yeah, we're getting point wise improvements on a certain metric, right? Like I love that. And I do that all day in my current. 20:57 Speaker 2 It's not inherently have to go in and explain what does. 21:00 Speaker 1 Yes, I think that's the thing. That's the challenge, right? Unless you've built this habit of being able to translate that to what your customers are looking for, it's going to be hard for you to to have the slash that you expect from a new model or a new update to the system. 21:15 Speaker 2 And I'm also just curious like from a user perspective and whether you're selling B to B or B to C, do you think that kind of the average technical capability of the user is also increased or the technical understanding of the user has increased? 21:27 Speaker 1 100% of our users are definitely using the products in ways that we didn't expect it. But it's not universal, right? Like it's a distribution, right? There's, there's always going to be a set of users that are still going to follow the same query patterns that we expect her, the same usage pattern that we expect them in the past. 21:44 But there is a larger growing set of our users that are, we would have considered to be pro users like 3 years ago. The input that they're putting in are longer to the point where they're understanding how these AI systems work and they're providing input in a way that they know will get better results out of our systems. 22:03 So it's like it's interesting to see. And I would say that that trend will just continue over time as users expectations grow as their experience heroes. You're going to see that translate beyond just search beyond just like the the chat consumer applications. So pretty much any vertical SAS where users are coming in with this prebuilt expectations rebuilt experience and they are going to be putting things into your system that you probably wouldn't have expected five years ago. 22:29 Just think about how many of our users and how many of us are used to just saying. Now summarize this for me. 22:35 Speaker 2 Like clean this up clean. 22:36 Speaker 1 This up like refine this like, you know, like we're just putting that into pretty much any copilot that's built into any of our products that we buy, right? Like it's, it's just gonna keep growing. 22:46 Speaker 2 Yeah. And do you think that that needs to be sort of a a critical consideration for founders that are especially if they've got a wedge product that is automating a certain workflow in a certain industry? That now that the average technical expertise and the ability of a user is, is increased a little bit more, that decision between build versus buy becomes a little bit more nuanced than it did before. 23:07 Like the, the ability for a non-technical user to build something for themselves, whether it's automation or, or whatever. You used to be a lot more difficult in the past than it is today. Yeah. And so I'm just curious sort of how, how do you think founders should think about that in terms of their end user and, and how they they pitched their product and, and is it a consideration now for users to just say, hey, I know you're automating this one workflow for me, but I also can probably just try and build this myself. 23:31 Speaker 1 Yeah, I would say that's 100% happening today. Like the buy bar is getting so much higher and especially because the cost of development in general is going down. And that's kind of where we go back to the fundamentals of of a startup or business. 23:46 Like do you understand your customer so well? You can communicate that you can communicate the value property. Do you understand your customer so well that you've enriched your own data that you have proprietary access to it? Nobody else has with the models that you can get from the public marketplace, like from open source or from any of the model providers. 24:04 So that you can communicate that value specifically to the customers in your sector in such a way that the buy decision is obvious because they just can't replicate that in a short enough time in their own stack. And I think that's ultimately where we're going back to like the business fundamentals probably haven't changed since the LLM era. 24:21 But I would say that the LLM error has accentuated or has increased the bar for what the user can do, what the user can do. And the the the companies that are had this nailed down three LLM era are even more valuable today, especially if they're able to leverage the technology. 24:40 Speaker 2 Yeah. And this is kind of an interesting thought exercise. It's almost like both the floor and the ceiling have jumped. It's just the floor has jumped a little bit more than the ceiling. So it's like that middle gap, yeah, gets that's sort of where you have to play and that's you have to show show value there. So it's interesting, I know we touched on a little bit ago sort of some productivity and efficiency gains that founders can get by using AI internally in their internal stack. 25:02 So not necessarily just in their product. Um, so I'm just curious sort of any thoughts or considerations that the founders should take into account as they're sort of starting their business in terms of building out their stack, whether it's from a sales and go to market execution standpoint, whether it's from thinking about compute costs, infrastructure, things like that. 25:21 Just curious on what are kind of the new age considerations for founders when when building a company? 25:26 Speaker 1 Yeah. So I would tackle this question in, in two different ways. Like 1 is founders really need understand what I call the variables of an AI system. So this is something that's critical. I mean, you have to understand what capacity really means, like how much GPUs you actually need? 25:42 Are your GPUs being underutilized? The different models have different capacity needs. Can you make a judgement call on like, hey, you know what, maybe this model gives us marginally better quality. This other one is significantly cheaper to run. So there's lots of variables to this new AI systems world, right? 25:58 Like there's capacity, there's cost, there's like. 26:00 Speaker 2 Latency. 26:02 Speaker 1 So you have to be able to optimize all these levers to like make it so that you're maximizing the value that you can see your customers. And the only way to understand these variables and to be able to do this on a recurring basis is to actually have automated systems in place that can evaluate new technologies and new systems and new processes and techniques that you're applying in your in your business. 26:21 So that's like Part 1. And then Part 2, you have to think about things in a systematic way. And it goes beyond just the tech stack, like how do you systematize, go to market with the evolution of your product? How do you build automated pipeline? 26:37 So like communicate this value, this new value that you have in a fast and cheap way. To kind of summarize your, the answer that I'm giving you here, it's automation is beyond just technology. It's now all throughout the entire company operations. And you have to think holistically as a founder, how do I automate most of the parts of my business so that I'm able to take the value that I get from the tech stack and the technology evolution and quickly distribute that to my customer should they see the value in our business? 27:06 And they continue to be our loyal customers that pay more and more every, every single yeah. 27:12 Speaker 2 Yeah, that makes sense. Like we've also seen like on the investor side, just the standard from a growth perspective, but also from a metric stand, sorry, from an efficiency standpoint, the bar is become a lot higher. And now what before was considered tier one or tier A growth and efficiency and unit economics and things like that. 27:30 Now that once again, that benchmark, I sort of jumped a little bit. And so you're seeing companies scale to 1 to 3 million of error with sub eight full time employees. And so it's like that almost R to FTE ratio has been that's just and we've just seen it on a massive scale in terms of the amount of businesses that are growing so efficiently. 27:48 And I think a lot of that attributes to that piece. So it's like like you said, it's you kind of have to balance the internal efficiency piece so you can keep up with that benchmark, but then also not forget the fundamentals of of still communicating that that value prop to the customer. So I think that's interesting. Maybe just sort of on a more curious that's for for yourself in terms of your productivity stack and different tools that you're using. 28:09 Just curious on on kind of what, what tools have you been experimenting with that have made kind of your daily productivity boost and things like that? 28:17 Speaker 1 Yeah. So I love LM Studio and Alarma. I run a lot of machine learning models locally on my own machine. So that has enabled me to effectively build custom tools for a lot of the things I used to have to do manually. And it's been really a big productivity boost. 28:32 I don't have to think about token costs. You don't have to think about, oh, whoops, did I accidentally just charge myself a million, $1,000,000 a month? Because I just run everything locally. I can see like when things are going haywire. So it's like, it's interesting because the ability to run models that are pretty good, like they're not obviously the tip of the spear here, but they're good enough. 28:56 And, and I mean that like they're good enough locally on my machine has been a huge productivity boost. The the I used to have all these side projects that, you know, would I would, would take me so long to get through because I just didn't have that much time. I was yeah. So we're focused on there weren't a huge priority. 29:13 But suddenly I have these coding LMS that I run locally. Sometimes I use the ones that are available online as well. I'm not going to say which ones, but we'll keep that a mystery. But locally on my machine, like it's just been helping me get through so much more of the side projects that I had that before it used to take me forever because they like the contact switch is obviously very costly and getting up to speed on the latest improvements has been very costly. 29:42 Like for example, I can today, I can quickly jump to a Python project that I haven't touched in like 6 months, run, updated to the latest libraries, update the APIs within hours. And that used to take me days. So I think for me, it's really just like having these models locally on my machine has been a true game changer because a lot of the projects that I don't want that are just kind of internal to me. 30:04 I don't want to share out to the world I'm now able to get through more of and get more things into it that are improving my productivity that makes in a way that's like wasn't possible before. Yeah. 30:15 Speaker 2 Yeah, I know for sure. And then maybe just taking kind of a a bird's eye view. And I know we haven't really talked about all sort of the content that you push out through through Remus readings on on a sub stack, which everyone should definitely go and take a look at. But just curious sort of in at the start of this year, you and some of your colleagues from MIT had put out kind of some of your your AI predictions for for 2025 and justice. 30:38 Curious sort of looking back, sort of what are the ones that you felt sort of played out? Was there any that you really had high conviction that were going to play out that didn't sort of come to to fruition? Just curious from a reflection standpoint, maybe we have to wait for for a year end refuse readings to get this, but just curious if you give us a sneak peek. 30:55 Speaker 1 Yeah, I mean, I'll give you a sneak peek. So yes, I am gonna have a follow up so you will have to subscribe and and wait for it, but I'll give you a sneak peek. So I'll, I'll just focus on the predictions that I made because I feel it's like that's fair. Like it's these are, I'm grading myself. 31:10 So I think that what I said and I'll read it out. Throughout 2024, I catalog through my newsletter, Rammy's Readings, a series of increasingly powerful open source LMS optimized to run on devices with consumer grade hardware. Thanks to engineering prowess, my own desktop equipped with an NVIDIA RTX 4090 is now overkill for running most cutting edge models. 31:30 At C2025, Ji was all the buzz with nearly every OEM marketing their local AI solution designed to run on their own hardware. And let's not forget Apples and four chips are equally capable of running state-of-the-art LMS with incredible seeds and energy efficiency thanks to MLX. 31:45 In 2025, EJI isn't just a trend that will go mainstream. That was my prediction. And if I were to grade myself, I would give myself a Gold Star because in fact, GPT OSB is fantastic and that runs fully on my existing consumer hardware. 32:02 But not just that. Recently, like within the past month, Mistral release the new model for coding that is just as good as the 120B version of Open AIS OS Model O. If anything, we've just reduced the cost of getting even more of the state-of-the-art onto even cheaper devices. 32:20 So I would say Gold Star and then you should definitely wait for the full review. 32:25 Speaker 2 For the full review and then any also sort of maybe one sneak peek into 2026 predictions. 32:31 Speaker 1 This is really a great question. I think 2026, we're going to continue down this path, but I think one of the things that you're going to start to see more and more of is local Amazon really, really Edge devices. So I'm talking about Raspberry Pies and phones. There's already an app that I can install on my iPhone that lets me download some 4 billion parameter models and run them locally. 32:51 What I'm predicting for 2026 is that more apps are just going to do that behind the hood and reduce their own token costs pretty much by running things on Edge that are just as good as what they need, what they need. Before that they couldn't do it before on Edge. 33:05 Speaker 2 OK. Well, I guess we'll have to do this at the end of next year to get another grade, but I think that that's that's all for today's episodes. Remi, I really appreciate you kind of sharing your expertise and and taking the time and make sure to check out sort of remote reading on stop stack. We'll put the link and sort of in the in the description there. 33:21 And, and be sure to tune in for more episodes with some more incredible guests like yourself. And be sure to visit our website at framework dot VC and thanks for listening. 33:31 Speaker 1 Thank you so much.

Podcast Summary

Key Points:

  1. Remy, a software engineer leading generative AI engineering at Microsoft, discusses how AI has transformed search from keyword-based to conversational experiences.
  2. General-purpose foundation models are optimized for the average, so founders must build specialized products that solve specific user jobs exceptionally well.
  3. Context engineering, which involves selecting the most relevant data to feed models, is critical for delivering accurate and useful AI outputs.
  4. Startups should develop their own evaluation metrics and automated testing loops rather than relying solely on global benchmarks.
  5. The choice between human-in-the-loop and fully autonomous AI depends on the vertical, with latency-critical industries favoring automation and others benefiting from human oversight.
  6. Defensibility today comes from owning proprietary data, fine-tuning open source models, and delivering value to customers faster than competitors.
  7. Founders must communicate value in the customer's language, translating technical capabilities into tangible business outcomes for non-technical industries.
  8. Running models locally on consumer hardware is a growing trend that reduces costs and boosts productivity, with edge devices expected to expand further by 2026.

Summary:

In this episode of the Framework podcast, host Speaker 2 interviews Remy, a venture partner and engineering leader at Microsoft AI, about the evolving landscape of AI, search, and agentic applications. Remy shares his journey from software engineering at McGill to leading major AI features like Copilot Search and Bing Chat. He explains that while business fundamentals remain unchanged, the LLM era has raised the bar for user expectations, requiring founders to solve specific problems exceptionally well rather than relying on generalist models.

A central theme is context engineering, where domain expertise helps systems fetch the right data to meet user intent. Remy emphasizes that startups need their own evaluation loops and proprietary data to build defensibility, since open source models are widely accessible. He also discusses the balance between human-in-the-loop and autonomous systems, noting it depends on the vertical and latency constraints.

Founders must translate technical value into customer language, especially in non-technical industries. Finally, Remy highlights the productivity gains from running models locally and predicts that edge devices will increasingly host AI capabilities by 2026.

FAQs

Context engineering is the practice of determining which pieces of available data and grounding are most relevant to feed a model for a given query. It aims to surface the right context automatically so the answer meets user expectations without missing critical information.

Global leaderboards show general model performance, not performance on your specific vertical workflow. A vertical company should run its own benchmarks and automated evaluation loops as part of its build, dev, and test cycle.

It depends on the vertical and problem domain. Latency-critical environments like factories favor automation with supervisory human monitoring, while domains without strict latency constraints benefit from human-in-the-loop steering.

Founders need to understand capacity, cost, and latency, including GPU needs and utilization. They should optimize these levers and use automated systems to evaluate new technologies and techniques regularly.

Automation should cover the entire company operations, including go-to-market and pipeline generation. This lets founders quickly distribute the value of technology improvements to customers.

He uses LM Studio and Ollama to run machine learning models locally on his own machine. This avoids token costs and lets him build custom tools and complete side projects much faster.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.