How Intercom Cut $250K/Month by Ditching GPT for Qwen
53m 30s
Fergal Reed, Chief AI Officer at Intercom, explains how his team reversed their stance on fine-tuning over the past two years. Initially, they avoided fine-tuning due to rapid improvements in frontier models like GPT-4 and Claude Sonnet, focusing instead on prompt engineering. However, as their product Finn (an AI customer service agent) matured and open-weight models like Qwen became nearly as capable for core tasks, they shifted to heavy investment in post-training. This change was driven by three factors: saturation of intelligence needs for RAG tasks, significant cost savings (e.g., replacing GPT-4.1 for summarization saved hundreds of thousands monthly), and the need for finer control over behaviors like escalation. The team now uses a mix of frontier models (e.g., Claude Sonnet) for complex problems and fine-tuned Qwen models for simpler tasks, achieving a resolution rate increase from 30% to nearly 70%. Fine-tuning employs supervised methods and reinforcement learning (GRPO) with custom evaluation models, scaling from LoRA to distributed training on AWS. Internal benchmarks include real-world hallucination data and production A/B tests measuring soft and hard resolutions, prioritizing practical performance over external benchmarks. This strategic pivot, supported by team growth from 10 to 55 people, highlights how maturing AI applications can benefit from open-weight model customization.
(upbeat music) Welcome back to Chain of Thought Everyone. I am your host, Connor Bronston, head of technical ecosystem at Modular. And today we are talking about something every AI product team is wrestling with right now. When do you fine tune? When do you use frontier models? And how the heck do you make these decisions when everything changes quarter by quarter? If not, week by week at times with some of these recent model releases? My guest today is Fergal Reed. Fergal is the chief AI officer at Intercom where he has scaled their AI team from just 10 people to 55 in about two and a half years. And planning to double that team size again next year as they continue to heavily invest in AI, including their wither agent Finn, which we'll talk a bit about. But what makes this conversation particularly valuable is that Fergal and his team have completely reversed their position on a fundamental strategic question. Just two years ago Intercom's take was relatively simple. Look, frontier LLMs are improving so fast. We don't really need to spend that much time fine tuning. Just get really good at prompt engineering, providing context, and ride the wave of better models. Today, they've instead heavily invested in post-training and fine tuning open weight models with fine-tuned Quinn models running at scale and production, replacing GPT-4 for key tasks and saving significant money in the process. So what changed? How did Intercom and Fergal make these decisions? And what does that tell us about where the industry is heading? Fergal, welcome to Chain of Thought. It's great to see you. Thanks for having me Connor. Glad to be here. Yeah, I think this is going to be a really fun conversation because often in these talks, we don't focus on a use case and really talk through the trade-offs that are made along the way as much as I think we ought to. It's very common. And I think listeners will tell me this occasionally, like, hey, look, this is maybe a little too high level. Like you just talked about what's exciting, what's interesting, and those conversations can be useful, but I think it's important to drill down and really explore why decisions were made at a company like Intercom and I think it's going to be a fascinating test case. So let's just start with that elephant in the room. Two and a half years ago, your position was essentially, don't bother fine-tuning. The frontier models are improving too fast, focus elsewhere. Now you have, as we said, heavily invested in post-training open weight models. Walk me through what changed your mind? How did you get to where you are today? - Yeah, absolutely the corner. So things really changed out there in the external world. So again, two and a half years ago, we were just starting to build thin and we were building thin initially with GPT-4, GPT-3.5 at the time. It didn't follow instructions as well. You still listening to more. And GPT-4 was this threshold for us where we could give it instructions to try and constrain it to just answering from text that a business customer of ours controlled, which everyone calls Ragnar. And so that was really, we were very early to that. And there was really so much to do at the time by just getting better and better at prompting these large models. Like the difference between a bi-gly prompted GPT-4 and a very well prompted, a well engineered prompt, nice and day in terms of theatrical performance, in terms of theatrical accuracy. And so there was just this period of time for maybe a year, year and a half where models got better so fast. They got cheaper, they got more efficient. We had modeled like Sonnet came along, Cloud Sonnet 3, 3.5 came along really much better at like following instructions. And we could do a lot more by just getting very good at like prompt engineering, tuning, back testing, building more and more into the core prompts. And at the time there was other people in the space talking a lot about training their own models or maybe taking a model and post training, or mid training with this mix of day secure for their specific area. And we looked at that and we did a few experiments. I remember we ran an AB test with a fine tuned version of, and it was Text of Inche tree, one of the GPT, 3.5 kind of turbo models. And we ran experiments with fine tuning. We fine tuned models like in the GPT 3.5 kind of turbo series. And we did AB tests and production of those, but we just didn't see the kind of the improvement overall versus instead just getting better and better, like taking these models and just getting better and prompting them. And so that was really the case for a long time. We made this very deliberate decision. No, we're not going to invest ton of money and time into fine tuning your own models. We're just going to piggyback on this, this wave of, it felt like every quarter or two quarters, there was a new leading frontier model that was dramatically better as customer service and not the sort of things we wanted to do. And really in the last sort of six months, nine months, that's kind of changing. Now, I do think the frontier models are continuing to improve and they're getting better and better, like being a gentick and following instructions. But in our area, our domain of customer service, let's say for our core rag tasks, the models, it feels like they're saturating the intelligence level that we need for that core task. And instead, what's more important to us is, is really fine green control over when the models do and don't behave in certain ways. So we'll certainly care a lot about the intercoms, with Finis, when should you escalate? When should you say, hey, I went to involve a human in this? And that's something that we care a lot about, very fine control over. And so we spent a long time, you know, hand engineering are prompt, we're like certain times when you should escalate and when you shouldn't, given customers even the ability to provide guidance to that. But we've been able to get a level of control over and above that again, by actually post training and fine tuning, some custom models. And so I really think that like, when you're your application maturers, maybe when you're, you no longer need frontier levels of intelligence, it's a good idea to think about doing your own post training because you can reduce cost a lot and also you can get like, much more fine green control over the exact policy that the LM follows. So both of those reasons have been very effective for us. When did you start to question that original approach that you were taking? Was there a specific moment where you said, actually, maybe we need to change up how we're up hurting this or was it more gradual, maybe even correlated to simple team capacity growth? And the ability to say, oh, look, now we can do both. And it's definitely related to team capacity growth. Now the causality goes both way around there, which is that like we hire more people because we want to do more post training. And so we did up because we think there's more, there's more value to it, there's more upside-down. And I would say that like, you know, around the time of deep seek, and that was really when the open weight models started to become close to the closed weight models and started to become almost good enough to do some of our core tasks. And so for a long time, you know, fin underneath the hood, there's the hardest prompt of fin, which we use Cloud Sonnet for in production, which is like that kind of core, answer the end user's question. But then there's always been a set of ancillary prompt that are very important to the overall experience that we've had for a long time. And we would always try and run these prompts on the smallest, cheapest, lowest latency model that we could to get the best end user experience. And so for an example of this would be like, when you ask a question to fin, the first thing we've always done is we've always had an LLM that will try and summarize and canonicalize the query that the end user has asked. And when we found it's worked using an LLM to do that, before going to retrieval, because end users ask questions in ways that sometimes are, you know, they can be very colloquial, they can vary a lot from person to person, and they can be in a crash. And so putting a true in LLM to sort of like, hey, can you summarize this query and canonicalize this query? Before going to the retrieval model in the rank, it's like under retrieval models are powerful, but they're a lot less powerful than a modern LLM. So kind of doing that, that kind of, that the pre-summarization or pre-canonicalization phase has always improved the accuracy over overall retrieval system, which then in turn, you know, enhances the accuracy of the LLM generates at the end. And we very early on realized that we didn't need to run GPT-4 or cloud opus or something very heavy like that, to do that sort of canonicalization piece. So we were using models like GPT-4.1 to do that, which is, you know, a model that's pretty, pretty good, kind of cost performance trade off. And we were spending a lot of money on that summarization task. And I don't remember exactly, but it could have been like a quarter million dollars a month on just that summarization task. And so when you're at that, and then you see, you know, the Cran models are coming out, and the Cran tree models are very good, and use a 14 billion parameter Cran tree model, it's really good. And we were able to go and take that model and train that model, post train that model to do that summarization task, and able to replace that and save almost all of that, you know, several hundred thousand a month on inference just by doing that one thing. And then we get more fine green control over it. So it was probably a combination
of our product maturing, also the gap between the open white models and the closed white models for the specific tasks we were interested in closing. And obviously, there's the Quen Tree series of models, an extremely capable and powerful series of open white models. I think a lot of people are using the other production for those smaller or uncedured tasks. I love to understand more about how you trained these Quen models to be so effective here. So obviously, we've seen this explosion in open source AI, particularly I think DeepSeek has led the way a lot in driving the edge forward. And Alibaba, obviously, with the Quen models has done a fantastic job as well, kind of falling through there. And it's interesting, I'll say, to see the juxtaposition between open source models coming out of China and then elsewhere, and then how it's just propagating. So that's a whole longer topic we can dive into if we have time. But I'd love to understand specifically, okay, how did you take that open white model and actually tune it to what you needed for your summarization tasks and other tasks that sounds like? It's definitely involved. So we work primarily on Amazon, like AWS, and so we use large and EC2 instances and dash. I think our standard instance at the moment is one node of H by H200. So we got a lot of RAM. And we have set up, we have an AI intro team that has set up a lot of infrastructure to make it easy for a scientist to log in and to run a distributed training job on those large GPUs. And we went through a process here. We started that experiment in which like Laura and like parameter efficient fine tuning methods like that. And maybe on slot, I think we used a lot at the start. But over time, we sort of invested more and more in this. And now we use, I guess, more distributed training. And most of what we do these days is a combination of super full supervised fine tuning. And then for some of the harder things we're doing as well, we use reinforcement learning and we have models in production that we've trained. And we found that, you know, a GRPO has worked well for us for some of the things we've been able to do. We've invested a lot in doing things like building evaluation models. You know, we've invested a lot in building a resolution model. We care a lot about has this answer sufficiently resolved the end user query. We've invested in building a resolution model that does a good job at proxy the real signal we get from an end user, whether they, you know, whether it has been resolved or not. And then we use that as an impulse into our reinforcement learning setup. And so a reinforcement learning setup now tends to, you know, this is multi-dimensional setup. Like the resolution model, another thing is a whole bunch of things we care about in terms of like, are using like bullet points, the right amount, not too much. How long is your answer, how short is your answer? And we also use other open ways out of M's as judges as well for part of that too. So we've pretty sophisticated reinforcement learning setup. And now, you know, the actual, you know, setting up your, your objective function and stuff is relatively the smaller part of it, the harder part has been just the infrastructure of them of like, you know, probably if I look at our investment overall, we have a lot of heads in getting good at running these models at scale on AWS. And then getting good at that with us, putting the bulk of the investment, although, but both are, are important. But it sounds like the trade-off decision you've made is like, sure, we've added headcount to make sure that we can run this system effectively. But the feeling is that if you were, say, using, you know, GPT models still for these same tasks, not only with the model and for instance, be significantly more expensive, but you'd still would have to be doing this fine tuning to get what you're doing effectively. So I'm curious what has made Kwen in particular work so well for you. I mean, obviously there's the cost perspective. I think that's a huge piece. But it sounds like there are other reasons you're also focusing on the Kwen family. Yeah. So the Kwen models, I guess we realized pretty early on after the Kwen models came out firstly at the benchmark 12. But then whenever a model benchmarks well, you've always got to look at it yourself and try it in some of your own back tests or your own evaluations because you know, some people accidentally contaminate the benchmarks or some people just argue it's fairly common. Yes, maybe common, yeah. And then even even even even people who maybe don't, you know, it's so easy for a researcher or a scientist trying to optimize a benchmark to get good at teaching the model had to do the kind of task that's in the benchmark and then suffer generalization failure. So you know, we saw the Kwen tree models. Obviously, you know, a ton of token strutum. So really generous that they were released open weight. Because you know, large training expense to it and it's great them. And then when we backtested them, they worked really well on sort of our standard back tests. And suddenly we were very interested in this because we were like, these these seem like legitimately performant models. And then we, you know, we started to do production tests of them and we would always do like, you know, we always like to do a production test of that well prompted, but untrained model before we go and then start post training for everything just to get some sort of a sense of like, hey, what, what, you know, roughly how good is this is a close to what we're trying to do. And yeah, the Kwen tree models performed well on our back test for for our simpler tasks. And again, you know, Finn still runs, you know, frontier model, you know, sonish for I think we're going to process the movement of song at four or five in production at the moment for the hardest problems, but for those those easier ones and all the all that surrounding stuff we find is is a major part of the actual performance of the system. And you know, the actual, you know, if you look at fin, you know, fin is really just kind of this, this cluster of, you know, 10 or 15 different problems and machine learning systems. And, you know, we have improved the resolution rate of fin from launch was about 30%. Now it's, it's heading towards 70 and unlike the vast majority of that improvement has not been in the core hardest frontier out of them that has improved and it's been great, but it's been in that surrounding cloud of systems that rose, trivial model, the re rank model, we built a custom re ranker, we built that almost from scratch using modern and birthed our own cost function. That's worked really, really well and trained on a whole bunch of data. And so, you know, over time, we've getting a percentage point here, percentage point there across the 10 different parts of the system, a really odds up. So, you know, we, I guess we've been investing in this area for a while. And so, yeah, we were, we were just excited. I'm, as open white models get better and better and we get the ability to fine tune them and control them. We're excited to see if we can turn that into more resolutions. And I think I should be explicit here since I don't know that we stated this earlier. I may be alluded to it, but Fin is intercom's fantastic AI agent for all your customer service needs. That's what we're talking about here. And there's quite a bit of information on their website. You can find about it. Super interesting stuff. I want to bring up one particular thing you said, Fergal, which is this idea of, you know, when a measure becomes a target, it ceases to be a good measure necessarily, a good, gut heart's law. And I think good heart's law is really interesting here. In part because of that contamination issue, we alluded to, but I am curious if there are particular benchmarks that you're looking at when you are deciding which models to prioritize. So you mentioned that there's a few of them you look at as well, some internal benchmarking would love to understand a bit more about how you're evaluating these models and what external benchmarks you're using and if you're able to share it all, what internal ones. Yeah. And so I mean, in terms of external benchmarks, like everybody else, we look at the model cards that release. You got some sort of a sense there, but then you also have to worry about like contamination or people like overfishing to the benchmarks. And you know, in terms of internally, we have a suite of back tests that we've curated over time and it includes things like, you know, every time we've a customer that reports a hallucination, we capture now and then that goes into our back tests. So we have a pretty good suite of real in the wild hallucinations that we can use and we can basically quickly quantify, hey, in a customer support setting and a rank setting, is this model like with hallucination out? We have that and we have a whole bunch of other like quality things we care a lot about, like, you know, how often will the model give an answer to the question and then we, we use LLM judges like did they get the right answer or not? So we started, we started to build benchmarks and back tests internally like that, but then nothing beats testing and production. We test everything in production. We've done this. This has been our philosophy for years at this point and what we really care about is like resolutions and in particular, we tend to run an A/B test in production.
With a combination we look at like the soft resolutions and the hard resolutions and the soft resolution is when you know Finna's given an answer. It thinks the answer is correct and the end user hasn't said anything and Finna said You want to talk to a human if you didn't get what you wanted and the end user has like disappeared and not said anything We can't that as a soft resolution and that's just going to happen a lot people will get their answer It's customer servers. They'll they'll go away though. They won't bother thanking the bot but typically about 30 to 40% of the time When ever you give an answer to an end user they will do what's called like a hard resolution Which is where like they would be like yep, and that's a done is actually resolved my question I say something like tanks they say yes, I got what I want and nothing so you about 30 to 40% the rate of the soft resolutions and That's our north star. That's our like ultimate ground truth and are you using that for our LHF? Not directly because it requires a human in the loop and when we're doing our LHF You know, it's kind of it's offline and rather than in production where we don't have an open loop or L in production and be too risky as far as you know what you're actually getting at back out of it and we might get there eventually and we So it just our workflow is about like take the model make the model grade of what you want it to do and like that's the context in which we tend to be like Tuning these models so we're like we're evaluating them on back tests and then we're taking a release candidate and we're a be testing that in production So our our our L is like I guess offline RL, you know, yeah, and I think it's understandable given how broad our data set is and Some of the pieces are just curious. I think I think there probably aren't that many people doing full RL open loop in production I believe that's my take as well. Yeah, I think I'm a me cursor or someone said they were doing it which was Pretty interesting like like you know, it makes it makes them it makes sense if you've got a very dynamic fast moving signal that you know and Like you know that so this is where you get into like boundaries algorithms and things like that in the past, you know If you were running a news site and you were using LLM's to predict, you know, who's interested in what news articles are if you're running the You know a X Twitter if you're running their recommendation algorithm for that where you know It's a very dynamically changing signal. I guess you'd want to do that but for us in customer service It can change fast it can change month on month but not day on day and so like an offline Model training setup is fine for us and we have There's other parts of our stack so like our retrieval model and And you know, we we use in production signals like our retrieval model is trained all you know very heavily on Direct real user feedback where the real user is like that resolved my question and then retrieval model learns Yep, this sort of article is good at resolving that question It's a little bit more in comedians a little bit more complicated for training our out of them But yeah, and our gold standard is always being those sort of resolutions and then we will always a B test any model before putting in production and we check to see like that the soft resolutions have gone up But also that the hard resolutions don't go down basically and we need to avoid building a deflection machine And so like checking that like the hard resolution rate You know at least doesn't decrease and ideally increases like basically that the ratio of soft hard Resolutions should be roughly constant for every model change and we've done this for a long time and you know Fin is that significant scale so we can get very highly statistically powered a b tests It's a bit dated now, but just give you an idea like when we switched from GPT to Claude Sart, you know, that was that was a multi-million end user interaction a b test and you know, so we're at substantial volume that like if in Now is resolving more than a million conversations successfully a week Think proposively go about over that and so so we're at enough volume that we can just afford to a b test in production Which real end user signals and I would say that like That is the gold standard and nothing else comes close in that like every Backtest and every evaluation that we've ever built no matter how carefully we spend the law we spend the lot It's possible that the just grand treat in production deviates from it the actual The you know and we we we we we care a lot about like a tenth of a percentage point of resolution So like we we need we need very highly statistically powered testing setups in order to do this and We can we just find this the weirdest things happen in terms of like a cut one great example This is like latency like for a long time. We always believe that decreasing latency would and would be a better end user experience We still believe that but we now know that increased latency almost always leads to more resolutions And you might be like oh well, I can understand how it would lead to more deflections But increased latency tends to lead to more hard resolutions Interesting. Okay. I love to unpack that a bit because I I really like that you have these kind of two north stars I'm like we want more hard resolution But we also want to ensure we're not losing soft resolutions by doing that and you keep that ratio I mean it must mean that you have I mean frankly I think there's a lot of people listening maybe a little jealous of like how strong of a data set you get in production here So I would love to understand a bit more of of you know that deflection rate and how that's all coming together The point it does a latency thing is just it's very uninsured of you really need to measure in production Like you you'd never build that latency signal which confines everything you would never build that into your back test because Suddenly like a longer answer might seem like it's got a higher resolution rate But the reasons got a higher resolution rate is because there's more latency before you serve it to the end user More latency may increase the end user's perception of work that the bottom of the Fascinating and something you get used you got a lot of confanding and so over and over time you learn to tease out these confanders But you know, I really think there's there's two types of AI products You know you all models are wrong from a useful. There's two types of AI product in the world at the moment What is where it's like we have hand engineered our prompt with a data scientist to this specific business This sort of for the poor engineer model and Almost everybody's doing that because you get a lot of accuracy quickly with that and there's a second type of product Which is what we've invested in we've built in which is build this scientific system where you don't have to hand engineer at per customer And if you go down that second route while you don't have to hand the engineer per customer, but I think the bit everybody misses is That means that the system you now have this scientifically optimizable system that will get better and better over time Whereas if you branch your prompts and you hand prompt per customer You can't build this system that gets better and better over time you stand up But like technical debt you end up with like a slightly different prompt for every customer and I think as we all know at this point Prompts like leak if you make a change one part of our prompt it affects the performance of something somewhere else And so you really want your prompts to be as standardized as possibly on your systems be as standardized as possible With well isolated well encapsulated different pieces that can be tested it you know separately from each other I am but always checking the overall system accuracy and that's really what we've built and I really think that over time People are going people are going to realize that you don't want the product that has been built for you by someone just like Hand hacking the prompt like it seems great right on week one It's like oh and I asked for this feature and suddenly they turned it around yeah They did but you you know you no longer understand the product And I think everyone would realize that if someone was like yeah, I have to sass fender and they made a great product for me They branched the database just for me Ever to be like oh, how are they going to maintain that well? They're clearly not going to maintain that and so I think AI people have not made that connection yet But yeah, we believe that in the long run and the future belongs to people who are building AI products Who were doing this sort of like a standardized well-designed AI system that then has this level of like testing and rigor and that it improves over time to Scientific standardized testing That is where we are by and high senior bats anyway. Hopefully it'll work out for us. Yeah, thanks to Galileo for sponsoring this episode Their new 165 page comprehensive guide to mastering multi-agent systems is freely available on their website at calleo.ai and provides you the Lens you need to understand when multi-agent systems and value Versus single agent approaches how to design them efficiently and how to build reliable systems that work in production Download it for free at the link in the show description to discover how to continuously improve your AI agents Identifying avoid common coordination pitfalls master context engineering for agent collaboration measure performance with multi-agent metrics and much more I'd love to understand how this correlates to your team building strategy and kind of the stages with which you've Continued to develop this approach. So obviously you you moved from this prompt focus to you know a post-training setup What was the the MVP for your post-training where is you know where are you today and how is the team growing alongside of that because I think it'll be really informative for other AI leaders who are listening and thinking through
Okay, how do I need to scale up my team? How should I be thinking through these stages of enabling much more scientific approach? - The scientific piece is really culturally and it was in the team from the very start. In a previous role, I once worked for a company called Optimizedly, A/B Testing, SaaS Company. And so I've all sort of believed in A/B Testing experimentation. And I think we really took that into the AI group. And so all our scientists will go and will run tests in production very fast and very rapidly on our kind of responsible for checking and making sure it. And when we'll only interview, our interview loops always contain a lot about like randomized control trials, a scientific process. We really want people to chat. We want really want people to be coming to the team as scientists who are familiar with the core fundamentals of science. Because we used that day in, day out to make a better product for our end users. And so we've ended up with a pretty technical structure in the AI group. We skew very tack heavy. And our scientists tend to be, they're mostly applied scientists. And they'll be good at like writing prompts. They'll typically have a background in machine learning. Often they'd have a PhD master's in machine learning, three to five years of industrial work. And see a lot of people have come in, maybe they've worked in recommended systems, things like that in the past. You see a lot of that same sort of testing and mature development process. And then we also have a large cohort of engineers. And the engineers are typically back in the engineers. We kind of draw about half of them from intercom, half of them are new hires. And it tends to be our most experienced kind of back in the systems engineers, but who are then kind of like crossing into machine learning and AI as a discipline. So they've learned prompt engineer. And we've talked in prompt engineering. And then also pretty common for them to be able to run an AP test, interpret the results, without able to run a classifier, use scikit-learns, do some NLP stuff, or whatever we need to do. Again, or as the LLM gets cheaper and cheaper, they become our go-to tool more and more. But yeah. And it also seems like you're doing a lot of context engineering. So as I know you invested in REG quite early, talk me a bit more about how that factors into this pipeline you've developed. Yeah, I think we must have been one of the first production REG deployments, certainly. Like we launched in GPT-4, launch day with GPT-4. And it was really like before GPT-4, when you did, when you tried to do REG, it was very difficult to get the previous generation models to actually respect the REG instructions reliably. So I remember when we launched Fin, I had made this decision to go this REG direction. And we had a whole lot of competitors at the start who were just serving the NAKEN model. So we'd be like, hey, here's a chat, boss. And we're just going to ask GPT-3.5, terabyte answer. And sometimes it would get a right. And sometimes it wouldn't. And we were like, oh, this feels like a bad product. The business doesn't do any control over it. And I think we've seen people standardize on REG. But we were certainly one of the first kind of production REG deployments. And back in those early days, yeah, context engineering was a huge part of what you did. I mean, it still is. But I guess probably in the first year after Fin launch, so for a lot of 2023, we were really realizing how important it was to assemble the right context and pass it to the LL. And we probably spent a lot of time in early 2024 doing that as well. And we're now at the point where we have quite a complex sophisticated pipeline for getting the right context into Fin. This starts where it's chunking. We take all the contents that Fin might possibly use. The contents is chunked by LLM. And then we go and in production, we take the query the end user has. And as I mentioned, we kind of canonicalize that query. Then we have a custom retrieval model. It's a fine-tuned version of Snowflake, but fine-tuned on actual production resolution signals. Then we have a custom re-ranquer model. It's completely custom re-ranquer using modern but there's a building block. And those together are responsible for a loss of the accuracy of Fin. And when we built our custom re-ranquer, we outperformed the best co-herry rancor at the time, which we were previously using. And then we were really happy with Dash. And that was just because we have, I think, tons of high-quality data of actually the CASDIS resolved the query or not. And when we had to get our customer's permission in order to use Dash, and we did comms and stuff where we were like, hey, if you want to opt Dash of training here, you can. But I think it totally makes sense to me as a customer. And I'd like, yeah, of course, a retrieval model, a re-ranquer model at this very little downside participating in Dash. And that makes the system better. So yeah, all of those things together really feed into this, I guess, this whole system that's responsible for providing an RK sac called Sona 3.5 or a Caution of that 4. With the right context that we need, and then we're like, well, I engineered prompt. And a whole lot of other stuff that goes into that prompt, some of which is customer-specific, where the customer has used our products to provide further guidance to us. And that's kind of-- that's how it's in generated answers. And that's how we've been able to get such high-quality answers. And as you continue to build out your team over the next year, you're going to get even more human power to support building a system out. Can you talk a bit about the goals you have for taking a next step within? And then conversely, I'd be curious what systemic risks you're seeing that you're paying attention to and trying to avoid from the start. Yeah. The second one is a hard question. It's a fast-moving field. Like, in terms of our goals, there's several big initiatives we have. One is we're building something we call customer agent, which we've publicly announced before, which is essentially that we think that the future here is not that you'll have different agents from different vendors all in the one conversation facing your end user. Because that's just not going to be a bit of experience. They're going to fight with each other, will the sales agent from one vendor hand off to the customer service agent from the other one? Is it going to be substantial overlap in terms of what they do in terms of the systems they need to connect? We also don't see a good future of some orchestrator agents that sits outside them and referees them all. That just feels like a pain. So we really think, we really see a world where to get good experiences, you really want one agent or one set of agents from one vendor. That's like talking to your end user's one beautifully integrated system. And so that's really what we're trying to build. We call that customer agent. So that means that we need to build roles for Fin. Fin needs to be able to act when the conversation is, is a sales conversation. It needs to be able to act when it's a customer service conversation that it does today. And probably other things as well, like customer success. Probably needs to be able to act well in the e-commerce setting. And so, you know, that is one area of investment for us. Just the other investments we're continuing to follow true on. We were pretty mature Fin voice product built using OpenAI's real-time API. I think we were one of the arid customers of that API. Very good API for low latency high quality. We're now in like version, I think four of our, are kind of like tasks or procedures product, which is this product that like takes actions and external systems. It took us a while to crack that. It was always relatively easy to take actions. Fin from launch back in March 2023. It could like call APIs. But the really nailed the product experience that made it easy for one of our customers to set up actions. It took a really long time. I feel we've just gotten there, really in the last six to nine months. We're starting to see real adoption and product market fit. And people are now using Fin where it's like, you say something to Fin and it takes an action that opens the garage door in the real world or something like that. It's really quite cool. And people using it for very consequential real tasks. So we'll be following true of all those investments. We'll be following true of our investment of continuing to invest in better and better AI models. And again, we've gotten good at post-training. And it's worked better than we thought. Like you can take an open-weight model. And if you have enough data and you have the expertise and it's a bit of work, you can get real benefits on your actual task. And again, when we beach, Cohera is model after re-ranking tasks and a great company and their model, we had previously done a big evaluation to find what's the best re-ranker we could when we were using Cohera's re-ranked tree, my mismemory in the name. And then we trained our own re-ranker from scratch, using modern and worked on a lot of data. And then even out of sample, even cross-customer, it beat Cohera's re-ranker at this task. We were like, wow, and like it beat its substantially. We wrote a blog post on it. It beat its substantially. Much more so than we expected. And we were like, wow, there's a lot of value to fine-shooting first specific task or training from scratch for a specific task. And so we still believe that the future is probably vertically integrated AI products. And I think we're starting to see that. Like I've been using Cloud Code, which Opus 4.5 recently. It's just a great experience. And completely agreed that the ability for. or cloud code to set up its models. Like, I mean, I've used Opus 4.5 with cursorance. It's a good experience, but having it within the cloud code setup, it just seems to flow so much better together. And like, they are going to have post-trained or mid-trained Opus 4.5 in that harness. And so, and we see this, we basically see like, there's at least, there is absolutely a value to vertical integration where the model has been not just prompted, but being post-trained or mid-trained for the specific application. And like, we see that this parts of thin that run way lower latency and way lower cost, like, you know, 10x lower cost, then we'd have been able to get by just prompting models. And these other parts where we would never get the exact right trade-off. Like, you know, we have a, like, this is part, you know, you end up, you're trying to have a prompt that does a particular task. And then you realize it's doing the task wrong. So you put a special case in to stop it doing that wrong task. And that works fine for a while. Then you're in production for while you got, like, 50 special cases in your prompt and you're overloading your prompt. And you just realize, no, I could like, post-trained these 50 special cases in. And then I wouldn't need to put them in the prompt anymore. My prompt is leaner and lighter. And so there is absolutely a value to that vertical integration. Is the value big enough to, like, carry it a day and to be the thing that everybody cares about? I don't know. But I do think you're starting to see it. You see, like, people at the model there. And obviously, like, you know, thinking machines came out with this, like, API to kind of, and I think you're going to see, you're going to see lots of people provide more and more interfaces. You're going to see lots of frontier labs providing interfaces to, like, specialize and find you in a post-trained, moving my guess, yeah. I think this is also maybe an important point to highlight because one of the common questions in, I guess, the prior era of software is, you know, why can't Google just do this? Why can't Microsoft just do this? And now we're, again, asking why can't Google just do this? Why can't OpenAI just do this? What if their model is good? It's good enough. And the answer, I think, we're increasingly seeing is, A, there's the unique data set angle of, look, we have this unique data advantage that others don't. And obviously, that's been something that's been kind of talked about by the data in AI industry for 10 plus years. But we're really seeing that come to fruition now as models get better and better. But this idea of a vertical stack, a harness belt for your task is clearly starting to differentiate. And, I mean, you have to stay on top of it. You have to continue to develop it. You can't, you know, rest on your laurels. But it seems like there is a real case for a defensibility that is arising within and other, you know, production use cases that have gone far enough down the road. So, I'll say a couple of things. I'll speak quite candidly. I'll say a couple of things. I first say it up. Something like cloud code is a case for defensibility and verticalization coming from the model there, right? Which is that like, giving that you have the best model, no one else can build a harness. That's as good for your model as you can. So like, it's a way of the model there getting up the stack, right? And so like, that's quite strategically threatening if you're not a model there company, right? And I can't even keep it in line, cloud code for a while as an exemplar of this great product. Okay, it doesn't necessarily mean that the vertical integration works the other way around. And so a company like us, we're coming from like up the stack. And then we're trying to vertically integrate to get the same defensibility, but coming from the top that. And so like we have been doing stuff like, and, you know, I would say we have, we have some real legitimate durable vertical integration undefensibility in fin today. I would not. Yeah, in that like, you know, all fin like, you know, and I think, you know, we do a lot of tokens in fin. And when everyone was releasing their tokens recently, we can't say that if we're like over 10 trillion tokens of fin and production. So that's just in production. So we have a lot of tokens. And like, you know, if we can make 30% of those, 50% of those using a lighter model, using our own custom model cheaper faster, better. And I think it's great. It's real margin, it's real lower latency. And then we have now specific examples where you get better performance than we could possibly get by prompting. So I think we're starting to see it. It's certainly the direction we're bending and I'm bending and I'm bending that like, we'll be able to create durable value for us and for our customers, but I've practically integrated in this way. But, you know, way to get into that is like, it's a bit or less than right? Which is not like, okay, data. How much data do you really need to train the horizontal model? You know, often these things are quadratic. You get like, you know, I don't know, you gotta go like 10 times of more data to get like, you know, two times more performance or whatever it is. And so, you know, pppb.clin.curve pretty fast. - Can't throw a compute at it instead. - Yeah, and so, some of the control computer that are like, you know, you can kind of get enough data to get the diminishing returns and more data pretty fast. And, you know, we even see that when we do post training with Fin, you need like, you need a lot of data, but you don't need like, you know, you don't need billions of data points. You can get pretty far with like, tens of millions of data points basically. And so, you know, so I do think that the large model providers definitely do have a hands to play. And I think a really strong one. And so, yeah, I did think that this possible just like still bit less than the dynamics here. But we have a hand to play too. I think like vertically integrating, I think, you know, the open models, like intelligence will saturate for different tasks. Like for the rank task of, you know, here is a whole lot of data and I give you the right answer to that question. You wanna need a certain level of intelligence in order to do that, you know? Like a human customer support agent can do it off. Like a, I don't know what, 100 IQ human or 110 IQ human is gonna be like as good as 140 IQ human basically up doing that task. Maybe better, who knows? Um, down, it's gonna be the same for like the out of bounds. Like you get to a certain point where like, yeah, the out of them is intelligent enough to do that specific task. And what matters now is about like the expertise, how it's tuned, is it greater customer service? Has someone gone in tune to be absolutely greater customer service? And you can do that with prompting you in get so far, we find you get a little bit further with post training. And so, um, see, I think there's some defense ability there that's certainly something we're trying to do. We're also trying to go broad, we've got customer agent. We think there's, uh, there's gonna be a lot of defense ability in terms of building something that is great. At not just customer service, but a customer service, at sales, at customer success, I think we can build a unified horizontal agent that we think is legitimately the best in the world with those things. And that, that feels pretty defensible. I feel like you need a lot of expertise. And then you build a product around that too, that make it easy to use through a parking, that that feels like a durable offering to us, or at least, at least as best we can get. If someone comes out with a singular here, something who knows what I'm talking about, I mean, yeah, it's hard to plan for that. But, but, but, which short of that's that, that's our strategic direction. And then we've a lot of validation in the best so far. So we'll see. We'll see. It's a really exciting story. And I appreciate you taking me through it kind of from inception to where you are today. It's been a super interesting conversation. And I honestly wish we had another half hour an hour because I feel like we had three other topics that I'd kind of bullet pointed out. It's like, oh, we should dive into those two. And, you know, forgley, you've been so fast running on the post-training front and data pipelines that I know we're not going to get to everything. But I did want to ask a couple final questions here because, you know, we brought up Cloud Code. I think you and I are both big fans of it. And my understanding is you were the first user of Cloud Code at Intercom. And now it's integrated throughout your development pipeline. Even I believe running autonomously on some bug tickets. Can you tell me a bit about that journey and what made you push for adoption company? Why? Yeah. I don't remember exactly why I started to use Cloud Code. But I had saw that, like, Antropic had released this new thing in beta. I checked it out, NPM installed. And it really kind of, it took me out for a week. Like every evening after the kids are in bed, I was like coding. I don't normally do a lot of IC work. I could do a little bit, but not much. But I was like, just coding with Cloud Code. And I was like, holy cow, this thing is amazing. This is like a strategy. It's kind of joyful. It felt magical. And I really spent a lot of time that week, like, you know, awake at night after I started to decode late at night. Awake at night. Particularly about the future of the software industry and like how horizontally disruptive it is going to be over time. And, you know, I reached out to the Antropic team. And like, you know, I was like, slacking them. I was like, this is amazing. You built an amazing front of you. And I think they featured me on, I guess, a little quote for it at page. I think it's cool. They did something really cool with Cloud Code. And then, I think it's an amazing product. And then when I went internally in intercom and I did like an AI group, all hands on it, telling people about it, some people like laughed at me. They were kind of like, oh, come on. We did this already in cursor. And somebody was so we serious. It's like, it will laugh to me. So we were like, oh, you're getting excited about it. But it's just the same as cursor. And I think there's a lot of explaining because I know that like, and the other ideas. How about that?
done a lot of work to become more agentic since then. But like at that point in time, it's like, no, it's different that when you put it in like that the full also dangerously skip permissions mode, which is like the only way to use this. (laughs) I'm like, okay, I'm on a secure machine, hopefully. And but when you do that, and you know, it just goes in a way that other things didn't. And so, you know, I kind of have to go and take several people in the company and kind of like call us and install like an install cloud code into them, like get them to sit down. Here's the laptop, look at this, check it out. I did that digitally from one or two of the execs. And then that started the ball rolling, okay? Once I kind of align those people and show them the value of an intercom started and initiative to try and that's still build a team responsible for taking this. This is gonna change our space. And intercom is really good. CTO are the people who really got it like, hey, let's be ambitious. Let's try it. Let's see how much we can get out of this. And so we ended up with like a good team of people who's jobless to make intercom engineering faster overall. And then they started taking a role with and build agents that live as part of the intercom system, intercom software engineering system, doing things like, you know, automatically attempting to triage and identify the source of an issue when a customer submits an issue. And you know, it doesn't always work, but when it's a hit, it may be say saves time on like a critical issue, you know, where maybe minutes counts in terms of, you know, having a great customer experience. So yeah, so it's been cool. It's been cool to see. And honestly, it feels like we're still just at the start of the adeasing, we're gonna show up all over software. - Yeah, I think there's a lot of exciting stuff that's gonna happen in the coming years around coding in particular, because we have just such an amazing data set for everything that's happening with GitHub. And then obviously there's so much focus on it. There's even like companies like Coolside, which we've previously had in the podcast, highly recommend folks listening to check out that episode with ISO and Jason Warren are the two co-founders that are talking about how they think code is gonna drive HGI. There's a ton of great stuff happening on coding. And I agree, I think Cloud Code is really, like a magical experience at times. And so it's a wonderful note to end on, because I think it hopefully highlights the opportunity and the excitement and what's possible here. And Ferkel, thank you so much for spending the time with me today. It's been such a pleasure chatting with you and really appreciate it. - Reapply to your comments. Thanks for having me. - And for folks who want to find out more about thin, I believe thin.ai is the best place to go. There's a ton more there. Definitely recommend checking it out. It's super interesting stuff that Intercom is doing here. And if you enjoyed this episode, we would love to hear from you, whether it's in the comments on Spotify, let us know what we, should we have asked Ferkel questions we didn't? Should we have, to have an end topics, we didn't cover, or did you really love something we talked about? We'd love to hear from you. We'd love to know what your experiences like here too, because that customer feedback is also valuable to us. I'm not yet training a model off it, maybe I should be, but I still am training my own mental model. So whether it's on YouTube, bonds, Spotify, on LinkedIn, or damning me on Twitter, or just adding me in publicly even, I'll take it. We'd love to hear from you, listeners. And thank you so much for joining us today. And for a go one more time, I really appreciate it. Thanks, come on. That's good. (upbeat music)
Podcast Summary
Key Points:
Intercom initially avoided fine-tuning, relying on prompt engineering and frontier models (e.g., GPT-4, Claude Sonnet) due to rapid model improvements.
Over the past 6-9 months, the team reversed strategy, heavily investing in post-training and fine-tuning open-weight models (e.g., Qwen) for cost savings and finer control.
Key drivers
Fine-tuning uses supervised fine-tuning and reinforcement learning (GRPO) with custom evaluation models, scaling from LoRA to distributed training on AWS.
The AI team grew from 10 to 55 people (planned to double), enabling both frontier model use for complex tasks and fine-tuned models for simpler, ancillary tasks.
Internal benchmarks prioritize real-world data (e.g., hallucination reports) and production A/B tests (soft/hard resolutions) over external benchmarks.
Summary:
Fergal Reed, Chief AI Officer at Intercom, explains how his team reversed their stance on fine-tuning over the past two years. Initially, they avoided fine-tuning due to rapid improvements in frontier models like GPT-4 and Claude Sonnet, focusing instead on prompt engineering. However, as their product Finn (an AI customer service agent) matured and open-weight models like Qwen became nearly as capable for core tasks, they shifted to heavy investment in post-training.
1 for summarization saved hundreds of thousands monthly), and the need for finer control over behaviors like escalation. , Claude Sonnet) for complex problems and fine-tuned Qwen models for simpler tasks, achieving a resolution rate increase from 30% to nearly 70%. Fine-tuning employs supervised methods and reinforcement learning (GRPO) with custom evaluation models, scaling from LoRA to distributed training on AWS.
Internal benchmarks include real-world hallucination data and production A/B tests measuring soft and hard resolutions, prioritizing practical performance over external benchmarks. This strategic pivot, supported by team growth from 10 to 55 people, highlights how maturing AI applications can benefit from open-weight model customization.
FAQs
Initially, frontier models improved so fast that prompt engineering was more effective than fine-tuning. However, as models saturated the intelligence needed for core tasks, fine-tuning open-weight models like Qwen provided finer control over behavior and reduced costs.
They replaced GPT-4.1 with fine-tuned Qwen models for query summarization and canonicalization tasks, saving significant inference costs. They also built custom re-rankers and other surrounding systems.
They used a combination of supervised fine-tuning and reinforcement learning with GRPO, leveraging infrastructure on AWS EC2 instances. They also built evaluation models, like a resolution model, to proxy real user signals.
Qwen models benchmarked well, and when backtested internally, they performed effectively on standard tasks. Their generous token context and open-weight release made them attractive for fine-tuning.
They use a suite of internal backtests, including hallucination reports from customers, and LLM judges for quality. They also run A/B tests in production, focusing on soft and hard resolution rates.
Reinforcement learning, particularly GRPO, helps optimize multi-dimensional objectives like answer resolution, length, and formatting. They use a resolution model and other open-weight models as judges.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.