Go back

The 15-day pivot that saved a $1.3B company - Des Traynor [INTERCOM]

0m 0s

The 15-day pivot that saved a $1.3B company - Des Traynor [INTERCOM]

The discussion centers on the transformative impact of AI on customer service and software development. Des Traynor of Intercom emphasizes that AI, exemplified by their agent Fin, is rendering traditional help desk models obsolete, compelling companies to innovate or face decline. Intercom’s swift pivot after ChatGPT’s launch—developing Fin within 15 days—highlighted AI’s potential to handle customer queries faster and more efficiently than humans. However, building reliable AI requires a fundamental shift from conventional SaaS practices. Instead of deterministic feature development, it demands experimentation, rigorous “torture testing” to validate capabilities, and continuous monitoring post-launch. Traynor warns against overpromising AI solutions without thorough real-world validation, as premature releases often lead to customer disappointment. The success of Fin, which now automates most of Intercom’s support, underscores the importance of this iterative, reliability-focused approach in delivering AI that truly works.

Transcription

11502 Words, 61211 Characters

English
Des Traynor - Intercom Honestly every single thing about how you build software has to change to deliver a liable AI software. If anyone can build an AI customer support agent, everyone who's relying on revenue from a help desk has a very limited future world. The alternative is you die. So don't play around with this. I would prefer a long, slow death into irrelevance, because that's the alternative. Speaker 2 Part There will be less customer service job in the future than there was in the past. Speaker 3 Today on Billions, I'm sitting down with this trainer, founder of Intercom. In 2023, his company was stuck at 10% growth. Customer teams were shrinking and the old model was dying. So he did something radical. He launched an AI agent priced at $0.99 per result conversation, not per seat per outcome. The results growth doubled to 25 percent, $343,000,000 in annual recurring revenue and a complete reinvention of a $1.3 billion company in just 18 months. That's thanks a lot for being here. The $1.3 billion bet on AI : moving 15 days after ChatGPT launched Thank you. It's a pleasure to be here. Speaker 3 So we've been like a long term user of, of intercom with my fast company. I mean, I think everyone has been really looking at the, the whole intercom success story and loving it from afar and loving it as a, as a customer. But I'm quite curious, you know, like in November, I think 2022, ChatGPT was just launching and 15 days later you're already building fin. What exactly did you see that made you move like so fast? Speaker 1 The biggest thing we saw was our head of AI, Fergal pinged me a link to it and said hey, play with this. And we all kind of, I think initially we were still working at what this thing was. So the first few queries everyone asked were things that Google probably would have answered anyway. But then it was, I think the unlock for me was, I think it was Kieran, our CTO, our founding CTO. He asks Chachi PD how would you install intercom on a mobile app? Which is the sort of question our support team would get quite regularly. And it answered it pretty much perfectly instantly, like in less than 5 seconds it rendered the right answer, which is a better experience than any. And our support team are good. But no one's replying to that in five seconds, you know, no amount of macros and all that sort of typical human efficiency software stuff would have got to that outcome. And then of course it could answer it in any question in any language and it could do a 24/7. And, and I think the reality very clearly was like, hey, even though at the time when AI launched to concern with if we throw our minds back, all people talk to us, hallucinations and guardrails and all that sort of stuff. But even knowing all that, it was still pretty clear that this was going to change the nature of customer service irreparably at the time. I think we were probably a little bit like, I mean, versus the rest of the industry, we were quite aggressive. But like versus where we ended up, I think we were conservative. We were saying, hey, this could probably do all the easy support stuff, but which by the way, for most support or exists like 60% of the work, you know. And so we're like, OK, cool. On, on the fullness of time, this could actually grow to being a meaningful chunk of support. When we launched Finn in March 2023 was doing 23% or 24% of support. Today it's like just about 70% umm of like of all support for our own support team. We automate fully automated and to and to be a fan of about 85% of our total customer support like the numbers are staggering now, but even back. Speaker 2 Then if this was a tool that just did 1/4 of the work of your team, that was still going to clearly be a smash hit. So like, you know, I think everyone realized over the following year that AI was coming for them. I just think we were first on the chopping block and as a result we had to move the quickest. Speaker 3 At the time, what was your arr? Speaker 1 I don't recall if we have that public at the time, but I, I don't know if we were sharing it back then, but like, I mean, you know, like the, the metrics that are true is like our SAS business was in the hundreds of 1,000,000, but I had gone through all of the usual business drama that every SAS can do entry. We went through a post COVID spike followed by a 2022 decline in in revenue growth. Top line revenue never shrank, but growth certainly came all the way down. And I don't like, I think Finn obviously was this like stratospheric hyper growth. You know, one of the fastest products I've ever been. I've also D fastest product I've ever worked on in my life in terms of revenue growth. And intercom grew fast with intercom itself went from 1,000,000 to 50 million in like 3 years. But Finn clips anything like that. So, yeah, it's like a, I think the, the biggest risk for us was like building and releasing this thing. I'm not prioritizing our help desk product for the period and actually then releasing a kind of a cannibalistic threatening product versus our help desk because the more work fin does, the less people need help desk. Yeah. Speaker 3 Yeah. And how do you feel like about it? Like right now is, I mean is Intercom and Fin totally separated within the company? Like do you have a very different persons working on the two products or is the intercom team still working on Finn? Why building AI is not building SaaS Intercom is the company and Finn is is one of our products. And if you like help desk is our other product. There will be more products and we'll have more stuff to launch soon and which will be another sort of group of people. But yeah, they're relatively separate groups of people. We effectively have a team Fin. We have team help desk and then we've been under another team called Team customer Agent who will have a lot more to share this year. But yeah, they're they're relatively isolated. But we do make sure because FIN works anywhere. If you can be a Zendesk customer or Salesforce customer and still use Fin. But we do try to make sure that FIN on intercom is like the Platinum Premium best version of FIN because we can, because we can tune to help desk and the agent to work to it. We should have perfectly. Speaker 3 Yeah, 'cause I mean for us to be honest as customers, I think it was very like smooth to add like Finn, it felt just like a kind of an add on. So I was wondering if within the company, you know, it was already like the the team was working on intercom or you created a team from scratch and you decided, OK, like we're going to build something brand new and. Speaker 2 So I think the. Speaker 1 We'll get a little bit into the nature of AI here, but I think our AI team kind of incubated Finn and that was an important decision. I think the way you build AI software is quite, quite different to the way you build SAS. I don't know how deep we want to go on this, but if I was to give you the kind of the top line reasons it's up. We assume when we're building SAS, you know, pointy clicky, you know, usual boring features, file upload groups, teams in rights, all that stuff. We assume that that just works like no one's ever scratching your head go, oh, I don't know if we can work out how to build an upload file dialogue. It's it's like this high certainty. It's deterministic. AI is not like that. You don't know what's possible. And if something's possible, you don't know how reliable you can make it. And both of these things matter greatly, right? And by that I mean like you can waste a lot of time trying to make something that isn't actually quite possible, or you can make it work once really well, but then later on find out that you can't. When you, you know, we deploy FIN, we deploy it to 7000 customers. The intercom help desk process is like 500 million conversations a month. If there's an edge case or a wrinkle in fin, we find out immediately. Our customers find it immediately too. So like it is a lot closer to mission critical than the average, you know, Dolly style render me a duck on a skateboard. Like it doesn't really matter if that doesn't work. And also no one has a good definition of what that duck on a skateboard should look like. So whenever they see anything close to it, they're like, that's brilliant. We're customers very different. Like people have a real high expectation that it has to really work and does a clear correct answer. Like there's a, there's a specific answer to how do I reset my password on Riverside. You know, there's like there's a correct and there's a wrong answer. So when you have this criteria, the way in which you build AI software is quite different. You need to take a think about it as a set of experiments that you can determine to a high probability outcome. So you can say, yes, we can reliably, given the knowledge base on a set of guiding principles, answer questions. And then you deploy it and you test and you monitor and you iterate. And it's the iterations that brought us from like say 23% when we launched to 6768% where we are today in terms of raw reliability. But our old software development process would not have built that. And I see a lot of like, frankly, AI out there in the world that just doesn't work is the best way I could describe it. Like they say this will do such and such a thing and I bet you it does sometimes, but it doesn't do it when you try it. And I think load of those scenarios are just because they've been taking an old SAS style attitude towards building AI software. The "torture test" for engineering reliability I'm really like interested in digging more about that topic because I think what you've mentioned, we've all felt it like the to be fully transparent with you when when you launched FIN, I wasn't like a big believer in it at first, no offense, but it's just like I had used like so many chat bots in the past where eventually like you get disappointed, you get frustrated, etcetera, etcetera. And now I think it's in, in our company. It's similar to to what you say. I think it's around like 70% of the ticket are now like dealt with spin. And it's it's so much better in many ways than an actual human in the way it does it that I think it's really impressive. You mentioned, like many companies promise and we've seen this like, I think recently, like for example, click up, they did like amazing marketing with MBE events on, you know, like agent, etcetera. And it's it's the vision is really beautiful. But people whenever they spend like an hour on it and get frustrated and it doesn't work. It's do the customer experience is really bad. So how exactly can you maybe digging into how you engineer your team so you make sure that you actually get to a level satisfaction that is a reasonable effort and for a larger amount of customers? Speaker 1 Yeah, I mean, I, I love talking about this, so I'll rely on you to shut me up or close me off. There's a honestly, every single thing about how you build software has to change to deliver a liable AI software. The way you think about dates, the way you think about Rd. maps, the way you think about evaluation periods, the product you design, where design comes in, in their process, everything is different. So if I subscribe to you, the old world, the old world will be, we talk to customers, we get a lot of feedback. We diagnose that feedback into feature requests. We go to the team. The team does some sort of loose T-shirt based sizing like, hey, that's a small project, that's an extra large project. Everything is put into a road map. We tell our customers this is what's coming and we get to work and we build it all. And we have like, you know, plus or minus weeks. We have like reasonable certainty of what we're going to build. We deploy it and everyone's happy. And then we also don't really care what happens when it's live. You know, product managers care to the feature gets used. But like, I'm not there. I chew my fingernails worrying, oh, what if the files don't actually upload in the files of cloud dialogue? I assume somebody looked into that. So that's old world software in the new world, it's just you're uncertain about everything. So the idea that you start with like wireframes or user requirements done, it's like that's not true anymore. You have to start with like with genuinely what new capabilities do the models make possible? So an example will be like, hey, it looks like with the latest version of say like Opus 52 or whatever, we can now reliably, let's just say, determine whether or not something should be a refund based on a very small amount of guidance. OK, How do we find that we can do that reliably? Well, we come up with effectively A torture test, right? Here's 100 scenarios or 1000 scenarios where we know what it should do, have an eyeball that unlike my. And then we check out the agent's performance. And what we want to see is does the agent perform at least at human quality, ideally beyond human quality, right? And and so you, you keep iterating and tweaking your agent until in the lab like is in in internal testing, you're pretty sure it is like it is not performing any worse than a human and arguably hopefully better. I'm going to say any worse. I mean like inaccuracy, it's obviously it'll be faster and cheaper, right? So that's what you just. Developing the "white smoke" moment for product Just to interrupt you on that example so maybe people understand better. Like let's say you have full conversation of someone would like a refund on a specific thing. Instead of having the agent the the customer person replying, it's the agent and you mimic kind of that conversation and you see the output of it and you score it. Is that correct? Speaker 1 That will be, that will be part of it. It might also be like the policy decision like as in if you're allowing the agent to make a judgment call, right, to say, hey, Guillaume, this is his first purchase. He seems that he seems legit. He seems authentic. He, you know, versus this is the 24th time this person's asked for a refund. In fact, they've never paid any money. Like you might have a what I would call like a fuzzy policy that explained that you might say something to the AI, like if the customer seems legit and authentic and they haven't abused the system in the past, give the refund. And that might be all the instruction that you have given your humans historically and they're getting it now. The question is, can the agent faithfully reflect your intention? So the way you find that out or not is you put it through scenarios and you put it through hundreds of scenarios to replicate the kind of noise of the real world. And this all, but this is all happening before a normal, you know, classic engineer or anyone else has thought about this. This is entirely experimentation in the, in the AI lab where they're just kind of working, working at some point you kind of we get this like white smoke moment where they're like, we're pretty sure we can do this. OK, now we move into productization, right? So now we're like, OK, build all. Now we talk to our designers, now we talk to our product managers and we say, hey, we have a capability. We think we can do this. What do we need to build around it? And we build OK, there's going to be a section called policies. There's going to be a section called refund policy, or maybe we have an abstract class called policy and you can basically create different policies. You say do's and don'ts and you give it 2 examples and that's all the agent needs to. So you hit save and then we say OK, Grant now. So now we've built it all. Now is the first time we believe we're going to ship something now so we can sort of put a deadline. And that's an answer to plan marketing the way we might you, you used to have, but open to that point. We had no certainty, so there's no point in telling our customers it's coming soon. We don't know if it's coming soon. You can tell people it's coming soon like Apple did with Apple Intelligence, and then it never comes because you've over promised and now you can't deliver. So you have to be careful what you say to your customers. Speaker 2 Now once we launch it, we are not done again in old SAS world, you launched, you're happy on to the next project in the new SAS world we're going to run. We, we have been testing this on thousands of scenarios. Now we're going to run millions of scenarios, true. And we're going to see how it performs. Do the customers configure it correctly? Did it provide good guidance? And then if they do all of that, how does it perform in the real world? And how do we measure that? Like what's our instrumentation or telemetry? So we can actually see, hey, it's doing most refunds correctly. In fact, it's outperforming humans or, and it's saving a lot of time or wow, it's thrown off a lot of edge cases and we have to see if we can get tighter. And all of that is just these are muscles we had never had before. Speaker 1 And so it's both experimentation to determine what's launchable. Only when you are confident you can do something reliably should you then plan a launch and think about classic software rollout. And then once it's live, you have to monitor the hell out of it and then potentially build your your customers more and more tools so that they can use it better. So you might have to build reporting tools or like visibility tools, oversight tools so that they can see the refund policy playing out in real time. So it's it's, it's a very it's like, I can't stress enough how different it is. And I think a lot of software companies what happens is like you might, you know, I don't know about the tick up example you shared, but I'll just give you a kind of a hand WAVY example might be like, oh, we do expense tracking. I bet we could run an agent that just scans a load of receipts, works out what categories there are, works out after valid expenses. They're not issues to issues the reimbursement to the employee updates, QuickBooks, blah, blah, blah, right. You might speculate that product in your head. I mean like I'm sure what could possibly go wrong and you might even build it and you might give it some really obvious, extremely embarrassingly like prototypical examples of receipts where it's like business dinner for six people expense line 7 or whatever. I and of course the AI performs great. So then you go when you record your screencast, your demo, you tell your customers, then your customers throw in real world taxi driver receipts from New York or they throw in drunken receipts from like A7 cocktail dinner that they handed their customer. And you see how the AI really performs and oftentimes it is frankly terrible. And now you are, you've basically exposed your customers to this thing that you've hyped. Speaker 2 The shit. Speaker 1 Out of it marketing, but it doesn't work and like there's so many products. Speaker 2 Where it's just like, hmm, you should have spent more time interrogating water or not, you could do that before you told everyone you were going to do it. Speaker 3 I agree. And you mentioned actually the the change in models and I felt like for some product the change in models has dramatically improved what they were capable of doing. So how like how often do you test like different models? And secondly, when do you decide that a model is really, like, much better than another? That it's worth pushing it to everyone in production? Speaker 1 Yeah. OK. So there's a lot here on the model specifically like the, you know, when a new model comes out, it's always a a worrying sign when I see a company say and we now operate on the latest model that just went live two hours ago. Because what to me that says like either you would extremely early access, which is possible but rare, or. Speaker 2 You actually have no. Speaker 1 Evaluation criteria whatsoever such that you were literally willing to hot swap the engine of your AI without testing anything. And that like is, is like not a good sign for us to properly evaluate a model, which we do a lot and I'll explain why in a second. It's like it's a significant amount of work. We look at thousands of scenarios. We have this what we call like a torture test. We look at thousands of scenarios. We look at what current Finn will say is the answer. We look at what this candidate release would say is the answer. We look at the difference in quality between the two. We look at like, what is the ideal? Like if God's own support agent had a, he had written an answer like the most perfect human answer from the most perfect person. And we're looking just at the distances between these, Is this ahead of that and as a close to this? Or is it actually going in a different direction entirely? And that's what gives us confidence to release. I would say honestly, since we've been iterating offend, the majority of our improvements have actually not come from model updates. They've actually just come from better AI architecture, better ways to disambiguate, better prompting, better use of different models. The reason we care a lot about the model specifically is like if Finn, we have actually been post training a lot of our own models. So Finn is a system that involves like maybe 25 different subsystems, a lot of which work on our own family of models that we've been post training as well to get them like really, really good at specific aspects of customer experience. So we have our own Reranker as an example. We have a fewer of our own models. Speaker 2 For like different types of of parts of the of the Finn setup and every single architecture rewrite or like reconsideration is a is a. Speaker 1 Full. Speaker 2 End to end re evaluation of the FIN system. So even though it's a 25 part subsystem, if you change one of these things, you have to evaluate the whole thing again because there could be trickle down or trickle up effects that you might not have seen. So again, like it all comes back to this. If you're serious about doing AI property, you have to take it very seriously. And that's where a lot of this kind of extra discipline comes from. Defining what "good" looks like in AI And in everything you mentioned, like I think there is a notion that is quite hard maybe like to, to, to understand and tackle, which is what good looks like. Because typically, like we've seen a lot of AI wrapper, for example, who could do like a basic task, like write me a LinkedIn post and then you have something and someone was never written. Any LinkedIn post might say, yeah, this is great. You know, like it looks smarter than I am. It's a stuff that looks good on paper they post and it's actual crap or whatever in your like on your end. How do you know what good customer support looks like? Like, do you base yourself on data from previous conversation that have been rated by customers or do you also have like internally some people who actually understand it? Because the the question I have is that like AI is, is really like towards science and scientifics and engineers, but engineers like sometimes don't really know what good looks like. So how do you manage both parts? Basically I. Speaker 1 Mean we do it with scenarios and we have definitions of good. So like, what is the right answer to a question? So in the torture test, one of the columns is like literally this is the perfect answer or in our opinion, this is the perfect answer. Here is the current financer and we're trying to get closer to perfect. But there's a great topic here for which was more probably applicable to your listeners, which is like everyone needs a definition of what good looks like and in a lot of domains it's very hard to work that out. So in the case of like, say, write me a LinkedIn post, Well, there's no actual definition of a good LinkedIn post. And if there was, it would be you know, if there ever was it's been like, you know how you say hacks to pieces and become the template for the LinkedIn post, which means LinkedIn then punish it hit the algorithm, which means it stops working, right. That's why we stopped all these like I was hoping to share this morning and I thought about like, you know so like what is actually the definition of good? Well, good, you can take it as two things in in the say the scenario of a LinkedIn post, you can already say good means to the user. Good. Means it performed well when they posted it on the LinkedIn and I've you know if you started using this thing and every single time you wrote a LinkedIn post with it it was a banger and it got 1000 things you book shoot this thing's brilliant, right? That's not a very responsive or like accessible metric, right? If you're building the LinkedIn, you know, whatever, like. Speaker 2 The host take generator, it's going to be really hard for you to see how your pieces performed in the world. The probably the best proxy you might have to value might be human acceptance, right? So you could perhaps build a like send to LinkedIn thing and see how often people click it. Or you could perhaps like once a user clicks accept, maybe part of the process is they have to authenticate and follow you on LinkedIn so you can see their updates. So you can get a sense of what percentage of the stuff we generate actually makes it to the wild. And maybe that is your proxy. Speaker 1 And then maybe you do have like, maybe let's just say 102030 golden examples where you think your product did really well and you maybe have 50 examples of the humans who did really, really well. And you've had to like factor out audience and all the others nonsense that will be going into this. We have to go over the scenario where like, hey, given a few bullets, this is a great outcome. This is a terrible outcome. This is a median outcome. And then your AI team, assuming you're actually doing a good job building this product should be done trying to improve the AI to make like it more like the best and less like the worst. And if you have a diverse enough set of scenarios so that it can perform well for dentists, along with for like, you know, software soft boys who like like to write like long pieces or whatever, we kind of need to make sure that we're covering the full spectrum. So that's how you do you. You always need a definition of good. Otherwise, I dare say, what are your team doing? They're just kind of performatively running around in circles, right? You need to have an understanding of are you making the product better. Speaker 3 And I want to build up on something you, you mentioned earlier, like all your explanation are like super clear and well detailed. So thanks a lot for that. But earlier, you know, you mentioned like the, the old way of building a SAS product. That is why it's again, you know, like quite straightforward. If if you're a developer, you know that you're going to be able to have this upload button or whatever. But now you're becoming a lot more like an R&D company because you actually don't know what it's going to look like. You actually don't know if you're going to be able to make it happen or not as a business owner and as someone you know was like building as before. How exactly do you decide how much resources you're going to put, how long it's going to last? And when do you want to ship something like what's what, what was your framework back then when you made that switch to OK, we're going to have AI researcher. We're going to have like a team that works on it and maybe for a for a year without shipping anything. Like how did you frame it? The Blockbuster warning: Adapt or die Yeah, like the interesting thing is for us, for Finn, there was a bit of a scenario, you know, the there's a famous book on AI at the moment, it's called If anyone builds it, everyone dies. And they're talking about like, you know, AGI and super intelligence and all sort of stuff, right? Like our version of that is kind of if anyone builds it, everyone else dies, right? Like in if anyone can build an AI customer support agent, everyone who's relying on revenue from a help desk is how's how's it go? Has a very limited future world. Like you know, and you know, in software, what that means is you become uninvestable as a business, which means you need to get back to your like free cash flow dynamics, blah, blah, blah. And like ultimately, like you don't have a bright future. So for us, like it was kind of there was a real fear of like, look, if this thing is possible, we have to know and we have to be the first to know and we have to move extremely fast on it. And I, I speak in the sort of more absolute terms in practice. No, Zendesk haven't died. They're still around. You know, of course, like it's not, it's not as as severe, but you're better off thinking it is that severe because you'll make the right decisions earlier. So for us it was like Fergal, who's our chief AI officer was like, I think this will work. And when he says he thinks it works and you trust you're out of AI, you're like, all right, well, let's see. So he put most of his group on it, which was at the time a very small group. Today it's like 50 people because obviously AI has happened, but at the time was before. But very quickly, they're able to demonstrate a. Speaker 2 Class of questions where it would perform really well. Then it was a case of all right, how much does it need to get this to a point of certainty. And like it wasn't, we weren't sitting there trying to like like take an accountancy approach. It was like, give it everything it needs and let's see what happens. And what happened was we got a product that could do like 23% resolution, right? And we're like, that seems good enough to go. So we put that live. We made loads of mistakes in the early days. We should have gone harder earlier. We should have launched on all platforms earlier, etcetera, etcetera. But I, I think like that's the first thing. And to your point about R&D like I think? Speaker 1 Like for all of your listeners, especially those who have a SAS business and are thinking about how to adopt and react. It's not you know, the wrong approach is like what's the right size experiment or how much of my team can carve off. It's basically, do you believe an agent can can basically replace the majority of the work that your product exists to automate and contextualize and visualize. So if you're an expense tracking up, do you believe an agent can do all the worker humans and they log in to approve and file and all that. If you believe that that's possible or if you have a strong sense as possible, there's not, it's not even a useful question to ask how many people should put on it. It is the only future of your business. It is the only future. And what that means is it could take all of your resources. You might have to go and hire 100 more, but like the alternative is you die. So don't play around with this, if you know what I mean, right. And there's a lot of examples like I of late, I've been reflecting a lot on Netflix versus Blockbuster or a blockbuster base. You just didn't take online seriously enough. And they kept saying, well, it's marginal. Well, it's this. Well, it's that. Well, it surely no one's going to watch a movie on a laptop or surely every single it will. They won't do this, but I won't do that. Every single one of their like narrative lines crumbled over time. And then by the time they realized they should take this thing really seriously and start trying to build Blockbuster online, they were dead. It was too late. And I just think that's do like there's a thing. I think Marc Andreessen said it recently on the Cheeky Pine podcast, which is like humans. And as a result, businesses will tolerate any amount of chronic slow pain over a sharp acute pain, especially self-inflicted sharp acute pain. And what that looks like in business is? Speaker 2 Rather on making painful decisions such as ripping up Rd. maps and reorgs and changing hiring and also like disappointing your current customers and all those things that are hard to do rather than doing all that people. They never say this out loud, but what they're saying is I would prefer a long, slow death into irrelevance. Speaker 1 Because that's the alternative part. I just think the, you know, if you're, if you're listening to me say this and you're one of the billions listeners like and you feel it a good degree of anxiety about this. Let me just say that anxiety is your body's way of giving you the future anticipated pain today in the hopes that you behave differently because of it. So that anxiety is actually your body's defense mechanism saying we need to do something differently because there's a very, very big pain or a terminal pain coming at the end of this current road. And my advice to you is adopt and react. And if it helps them writing a book on this, it'll be out soon. But anyway. Speaker 4 Nice. Speaker 1 But genuine I did. That's the best way I could describe the changes to make. Don't try to be an accountant about it. Don't try to work out what's the minimum amount of resources we need to carve off to play the AI. Think that's entirely how you die. You have to be like, if this thing is possible, we are screwed. We have to go fall in it. There's no playing around here. Speaker 4 Do you feel like if your company at the time like would have been growing faster and you know, you didn't have this kind of like like we all we all had this, you know, 2122 etcetera. Like do you feel like you would have made the the same choice or? Speaker 1 It's nested right? Like on one hand, there's definitely a lot of companies that were like, what I would say asleep at the cash register, right? As in they've just had success for so long that they just kind of forgot how to adopt and react. So I've definitely seen other companies who basically like hadn't had the pain early enough, react slower and and then on top of that, so, so we were definitely like, it wasn't like we were like pointed towards a perfect outcome. It was definitely like we were already in the middle of a of a pivot. We've this our CEO own had returned years. Our founding CEO had kind of returned after a few years and we were already saying like we need to rebalance, restructure, do some layoffs, change the robot. So we were already. Speaker 2 In a sea of monarch change, which kind of helped. Speaker 1 Because this was just one more change into a pot of change, which was useful and then and credit to own because the decisions you have to make here are painful under under like they're, they're easier to not do them. You know what I mean? They're a lots, a lot easier to not do these things. So, but all that said, I think the only reason maybe we still would have had to do this is because it's like maybe the mistake we would have made if the business was doing great and we saw this opportunity was that we just would have under invested quite a bit. And I like, I like the folks I was just criticizing 2 minutes ago, right? Maybe that's what we'd have done with yourself. Oh, that's cool. Fergal has a fun little experiment. Let's, let's leave him work on that. But I, I think we still would have had to keep scratching on it though, genuinely because it wouldn't like, it would be very difficult to sit in in your office every day knowing that like 1/4 of the work can be automated at the click of a button. And like, it was truly the click of a button was to turn fan on, you know, So like knowing that you kind of realize, wow, like we have to do something. But maybe we would have said, hey, like, jeez, don't like, yeah, you know, don't take people off help desk. It's doing really well. Why? Why would you screw that up, right? Speaker 2 Like, you know, and I, I definitely, I can see when I look, you know, we're talking here and now in like January 2026, as we speak, every public SAS company is down somewhere between 20 and 60% depending. Like SAS is taking a hammering. I think a lot of that is because they don't look like they're successfully transferring to AI. They don't look like they're getting AI right. Most of the big dogs in our space can't seem to figure out AI at all. And, and I think like what's happening is the narrative, which was like, hey, we'll pay you back, you know, a free cash flow for the next 1020 years. Speaker 1 Is now becoming one of hmm will you guys even exist in 10 or 20 years and that's showing up in the stock prices So like I think now everyone is finding the religion again like shit we better take AI pretty seriously I but I do January. It's possible you're correct that we would I don't think we would not have done Finn. I think we definitely were going to deal with too many smart people telling us we should do it I think we might have undergone it a little bit we certainly might have chickened out of launching Finn that AI and giving an entirely separate name. Maybe it would have been like the intercom AI agent and maybe we we wouldn't have built it on Zendesk and Salesforce like we might have pulled back quite a bit. I think that would have been a mistake. So we ought to do is just there is some like, you know, people say never waste a good crisis or whatever. There's definitely some some sense of like while you're going through manic change, you may as well make all the right changes. It's almost like, you know, if you've got the floorboards up in your house, fix all of the problems. Don't just fix the one you looked at. And for us, it was like, you know, there was clearly one huge problem, which was like, there's this technology that's getting really good that looks like you can do your job, you know? Speaker 4 No, that's, that's awesome. And thank you for your honesty and transparency. And earlier, earlier on you mentioned, you know, that a lot of people were coming at you and saying you don't like AI can also hallucinate and whenever you know, like it's, it's your own company. You don't want like fin AI to just start making up things or calling people out name, etcetera, etcetera. Killing hallucinations with actor-critic logic How have you managed this risk of hallucinating? Because from using like fini. I mean, obviously I'm not in the support team, but it's like I, I look at conversation and I read conversation like a lot. It doesn't really hallucinate. So I mean, so you must have done stuff that other people don't understand and don't. Do you mind sharing? Speaker 1 Yeah. I mean, I'll share at a high level, like in practice, if I share, if I give you all the details, it'll help our competitors a lot, so. Speaker 3 But. Speaker 1 I'll give you directionally like so we need a GPT 4. One of the breakthroughs that unlocked fin was GPT 4 at the time and GT35. We we couldn't guardrail it enough. It was still too, too willing to invent information to fill out to fill out. It's a it's kind of answer in practice, what does it look like? The the the technique that we use is called actor critic, which basically it's like there's a red team that kind of like attacks everything answer and says what is the like, what is the basis or what is the citation or what is the ethnological grounds, if that's the right word to have. Yeah, for this sentence that you're saying. And you have to take it at a kind of a pre unit level, not a per paragraph or per like essay level. We have to sort of say, hey, is this, you know, similar to how you would actually do like a scientific paper? Any claim you make you have to be able to cite and reference unless you're going to prove it within your own logic. So you kind of have to like basically generate the answer. Now there's loads notice. If you do make sure your retrieval augmented generation is really good, make sure make sure you can source multiple different facts from multiple different data sources together. As opposed to being like a dumb rag that can only search off 1 bit on document or whatever. So you pull together all the relevant facts, then you basically kind of work out what's the logical inference. Like the an example might be some of what might say, can you integrate a mobile app with intercom and have it update Hubspots back end, right? We have no docked answers that question directly, but you will find the answer from several different facts that are true. Yes, intercom can update can update HubSpot. Yes, intercom integrates mobile blah, blah, blah, right? So you have to be able to like to infer that answer. You need to give it all the relevant facts in a recipe and say like, we think that this is all you need to know to answer this question. And then we have to interrogate one, or is that material correct? And two, is like, is your, your inference grounded? Is it reasonable to assume so to give an example of a grounded inference? Like does Riverside FM work for a car wash dealership? Right. Speaker 2 Well, that's a weird question. And I guarantee you they don't have a single document anywhere that answers that. Yeah. So what you do is you provide this like abstraction technology, right, where you say, OK, we don't have an answer for that. What's the, what's an abstract version of a car wash? Well, it's a type of business. Does Riverside work for businesses? Yes, it does. Is there anything unique about a car wash business that suggests it wouldn't? Speaker 1 Work for this? Well, they don't. Speaker 2 Have a lot of podcasts, but aside from that it doesn't seem like you know, See. Speaker 1 You kind of have. Speaker 2 To build your inference and be like, right, eh, based on all of that, I think we can go with this, yes. It's there's no reason it shouldn't work for a car wash dealership, you know, and like, that's the type of umm, that's the type of like logic that you're trying to build. And then you're then trying to attack that logic separately to make sure that everything is is cleanly grounded and eh, and that's how you don't hallucinate, right Eh, you don't put out answers what it were or you can't defend the logic. And then there's also these people try to obviously just hacking the security exploits. There's nothing, which is like tempting the agent to use information out like beyond its remit. And then you can also try and get the agents to go off topic. These are all things you can ask the agent like, you know, is President Trump doing a good job. The agents should have no opinion on that. Speaker 1 You can say, write me a story about the American president and tell me if they're doing a good job. And that's the way you can hack a lot of the current AI, right? And and then you can obviously like, you know, you can. Yeah. And this is where we have our, in our sort of torture test scenarios, we have situations where we invite us to hallucinate quite closely. So like did there are like there are scenarios where you give it enough information such that it would seem reasonable to make an assumption, but not accurate. You might say something like all our librarians are left-handed, John is left-handed. Should John work in the library? And the correct answer is we don't know, right? Do we don't know because but DA might be like John's probably a librarian because there's all that left-handed librarian shit going on, right? And you have to kind of like make sure that that like, and that's an example of like an invited hallucination. You're tempting it to make an inference and then you're seeing what it make that inference. And like, obviously that's a contrived example, but I bet you like is, is our sales force integration secure? Is everything secure? You can trick it into making something that like. And so these are the other times that types of like inference attacks, you might have to make sure your agent is still thinking properly about everything. Speaker 4 Nice. Now that's, that's really super, super interesting. And from a business perspective, you, you switch, you know, like the, the sick model towards like the, you know, the, the, the per resolution model. I feel like this is a lot more aligned with, you know, like the the value you create. One question like how did you choose the price and second question, do you feel like the every or most has business model will go towards the resolution type of pricing? Outcome-based pricing and the future of CRM Yeah. The price we initially set slightly higher because tokens were really expensive and then they got a little bit cheaper, so we were able to bring it down a bit. It was really a function of a few different things. We try to look at one like what's the fully loaded bottoms up cost of having a human answer a question? And then we wanted to have like a transformational impact on that. So for most companies, they, let's just say for easy maths, let's say they pay their software people, they're sorry, their customer support people. 52,000 a year or 1000 a week. And then we say, how many conversations do they have a week? Maybe have 100 conversations a week, 20 a day. OK, so so how much are we charging? We're charging $10 per per conversation basically, right. And then we asked right, well, you know, if Finn can do a majority of those, how do we make sure that this is a no brainer to turn on? And that's kind of like that. That was one type of calculus that got us there. There was another consideration, which was like how much do our seats cost? And let's say somebody hands back all their seats and turns on Finn. Is that good for business or bad for business? We have to make sure that the answer was it's good for business. So like the seat might cost $100, Finn might do more than $100 worth of answers. If, if it has access to say 100, it'll at least, you know, if that's has access to 200, it'll at least, you know, get 70% of that. So 140. So we're up, you know, so you have to kind of get all of your, you think about all these variables that we're kind of pushing us up, like what's our actual margin and how do we defend margin in the future? I think I genuinely think AI margins is going to be a big discussion point over the next two years and our industry, Umm, but anyway, so you're kind of like the pricing point and then obviously like you're not going to come up with a price of like $1.00 and 17 1/2 cent kind of like what's a marketable price? So anyway, that's how we ended up a $0.99 per answer or per per per resolution if you like. Do I think it'll go that way forever? And I. Speaker 2 Think like when there's a clean outcome. I think outcome based pricing is the clearest signal you can make as a business that your product works. Speaker 1 And I think anyone sales force recently moved away from outcome based pricing. You could speculate that there's a reason why, right? Like if your product definitely works, then you should charge for the value it delivers. And the most crystallized form of that value is, is it delivering outcomes. There are definitely categories of work where there's no clean definition of an outcome. We spoke earlier about like a LinkedIn post generator, right? It's really hard to say. Could you charge $0.99 for accepted LinkedIn thing? Is it clean enough to tell what acceptance is? Can they not just take a screenshot and rewrite it, blah blah. I don't know, right Dolly. Like generate me a picture of a duck on a skateboard. There's no right answer there. There's no wrong answer. If you generate the same duck 50 times in a row, is that $50.00 that 50 times outcomes or is that one you know? So I think on a lot of other areas it might make sense to look at like a usage based pricing, which is basically how much work did the user command the agent to do? And like that's just assume everything is a function of that. On the extreme end, you end up like what I call it marked up tokens, right? Like what cloud code is whatever. It's just like, hey, you're just spending tokens. Don't worry about like Amazon do this with EC2 units, etcetera, right? Like just like we've invented a middleman metric that you're just going to trust, you know, it's kind of like a a form of gasoline or whatever. It's just like, use it. You know, some stuff costs loads of it, some stuff costs a little and that's fine, fine. I think when, when you're not doing any of those and you're relying on just a static, it's an extra $10 per seat or, or anything in that, in that like vein, I think you're probably admitting that you're either your product doesn't work or people don't use it or, or both. And, and as a result, you're just hoping you can like kind of slap it on as a, as a price upgrade, but it's, that's actually not that useful. So I, I, I, I suspect we'll see a transition to outcome based pricing. I also think when there's a clean definition of an outcome and the business isn't charging based on it, it's often a bad smell. Like it's it implies to me that they don't have as much confidence in the product as as they project. Speaker 4 And I want to build up on what you said about Salesforce and like the CRM space because I mean within intercom obviously like you have all your conversation with prospects who are asking question or on the free trial. If you have like a software you have conversation with customers were either potentially like going to upset or down set. I mean, you have a lot of information. I think no one likes their CRM. Like I've never heard someone, you know, say like, hey, I'm, I'm so excited about going and spending like 2 hours on Salesforce. What's your view about that space and how do you think it's going to evolve with AII? Speaker 1 Think they'll be a substantial like most CRMS, well practically all of them bar maybe Atio and maybe in clarifies and their own like, but most CRMS aren't AI native CRMS. And but because of that, what they are is just key value pairs. Like, you know, Dez is an account and Dez has name equals Dez, surname equals trainer, e-mail equals whatever. And like, I think the shift we're going to see is towards like, you know, contextual products. So like when you actually come to like, you know, to the DAZ page, if you like, what you want to know is what's the relevant material I need? And you don't want to be spotted a list of key value pairs. You want kind of a summary or whatever, you know, like an example of a company that's seen this really well, I think because like Gong, so Gong when you go to review, so Gong is a tool that like records your sales calls and does analysis on them. But Gong just presents you this really powerful frame, which is like, here's an AI summary of everything that just happened. The New Deal. Here's all the transcripts if you want to check why I think that. But like the summary is really good. I imagine the future of CRMS will be a lot like that. It'll be a bit better context and and the current state, but it won't be like based on spitting out a lot of key value pairs. It'll actually just be at a higher level of knowledge. And then separately, it will be like heavily like MCP powered. It'll both be both an MCP server and a client. So it'll pull information from like your, your like, let's just say Des is a really active user. And we got that from Amplitude because Amplitude runs a server where it can returns a string of of interesting stuff about Des. But then it'll also be somewhere to can plug in and say, hey, tell me what's going on in the DES deal. And it will. Speaker 2 Give you an accurate state of play. And I could do this because it's not limited to just key value pairs, but it can actually draw inference from these things. So like hasn't used a product in ages is marked as hot prospect. Hmm, something smells bad there. We should let you know. So I think that's the future of kind of like where CRMS need to go from an intercom perspective. What we're, what we concern ourselves with right now is entirely being the customer. Speaker 1 Agent So right, right now we've spent the last while talking with Finn as a customer support agent. The future of Finn starting in a number of like maybe a small number of months or a few weeks, will be to actually flesh out the city of Finn as a customer agent. Finn to be able to handle inbound sales queries, Finn to be able to make people success, make customers successful. So it'll be Finn as being this kind of conversational intelligence that talks to your users, hears from your users, starts conversations, handles conversations, delivers outcomes. And obviously we will be CRM agnostic knowing that intercom has CDP, if not CRM like like functionality, but we will be agnostic to the back end. But we think that the actual value is, is like how do we make sure what Finn goes into A to start a conversation with a customer, It has all the right state of play as to who they are. So it says the right thing at the right time. How do we make sure it knows how to start a conversation? Well, it's good at that. And how do we make sure it handles the conversation to a conclusion? That's like, that's a lot of what we're working on right now. And it's, and we're going through the exact process we said at the start of like scenarios, expected behavior, ideal behavior, all of that. That's all coming right now. Speaker 4 Nice, super interesting. Looking forward to to check this out and when it comes to like the the global vision, because at first, you know, you were really like helping out customer service people and and eventually, you know, you start helping them even more because you can actually like answer questions. But with AI, at some point, you know people were saying like your job will never be replaced. It will be replaced with someone using AI, etcetera. The end of frontline customer service jobs But as we move. Speaker 3 Forward we still see. Speaker 4 That some jobs are being replaced. So what's kind of like your vision about it? Like do you feel like the the customer service job will be replaced, transformed or in some ways like how, how do you picture it? Speaker 1 Yeah, it's, it's the D question of our time. I think, you know, we're going through a similar thing with engineering. We just don't realize it right now in my opinion. But. Speaker 2 You know, agriculture people used to dig with trails, then they invented the shovel, then they dug with shovels, then they invented the combine harvester, then they invented the multi engine combine harvester, blah blah, blah. And and like all the while, like agriculturists still existed. But the way in which people are deployed has changed. When you lower like the the cost of entry typically get more businesses. So a lot of people who can't afford to do support will be able to afford to do support. It's like the the entire business topology changes. Speaker 1 But I don't want to hideaway from saying like there will be less customer service jobs as it relates to frontline customer service in the future than there was in the past. And the reason for that is that AI can do. Speaker 2 A shocking amount of the work. So I think if you're ACS person who really wants a career in CS, the good news for you is it's about to get exciting. There's going to be a lot of like new roles people your businesses will need people who can understand and train AI and use AI to like devastating effect to like to deliver perfect customer experiences to to execute ACX strategy via an agent. And that's that that will be the skill of the future. If the skill of the past was being a really good OPS person who could think about building a team by time zone, by language, by product area and could and like staff them all over the world and make sure you could provide 27 global multilingual support. And the way you did that was by like, you know, frankly, like mass hiring and, and you know, and all of the usual skills, you know, CS is also like an entry level job. So there's a lot of churn and a lot of like attrition or whatever you have to manage through all of that. If that was the old skill, the new skill is knowing you have to like use AI to strong effect. And that's a really popular skill, eh, that will, I think be in massive demand. So if you're in CS and you want to stay in CS, which is, by the way, I think maybe that's only maybe 1/3 of the category for most people, CS as a job on their way towards the sales role, product role, you know, whatever. But if you're in CS, you want to say NCSAI is the, is the new skill we're talking about conversation design, automation specialist, prompt engineer, all that sort of stuff. And that's the sort of skills that people will need in the future. And there's actually, in my opinion, a far more high impact career ahead for you. But again, I, I, you know, at intercom, we have this thing we, we never want to talk down to people on patronize and say, don't worry, all the jobs will still be here. Speaker 1 We see our competitive suita. We think it's just disingenuous. It's not authentic. There'll be less jobs, but if you want to stay in CS, there'll be better jobs. And both of those things can be true. And by the way, humanity has been introduced shift hundreds if not thousands of times with every piece of technology that's been created. And this is just the latest. Speaker 4 OK, yeah, I agree. I know we're almost running out of time. So last question, what advice would you give to a founder, you know right now who's watching kind of AI eat into their market and isn't sure like what they should do? Speaker 1 The first thing is you need to get somebody who you deeply trust, who's very good at AI to give you an approximation of how much of your software in the future will still need humans. And it's going to be a far smaller percentage. The wrong way to approach AI these days is to say, oh, look, we can build a little copilot over here, or to say, hey, where might we apply AI? And that's where you get all these like a little pet projects, like, you know, the sparkly emoji and we've a little auto AI summarized feature. Like that's the exact incumbent attitude to AI, which is to think about it like, oh, we'll build a little bit of AI over here to keep the guys happy. The actual right way to do is to say, assume nothing exists. If you're starting this product again today and you were good at AI, what would you build? And where are humans, if anywhere? Where are humans essential? And you would build a product that uses AI everywhere. It can be made reliable and performant in this new AI process that we spoke with earlier in this new style of soccer government. And you would then build, you know, boring old SAS UI wherever you need. Humans evolved and that's your new product direction. And I think it's worth the first step really is and you need somebody who knows AI to help you with this. But that's the first thing is what is division like? You know, if you were to do it all again, what would you actually do? And then you have to then charter a course from where you currently are to where you need to get to. And that will include robbing resources from your current business. It'll include watching some customers quit because they're not seeing the features they wanted. It'll include all sorts of messy stuff. Engineers who don't want to learn how to use claw. That'll include designers who think that they feel disempowered because of the new style. You have to kind of bulldoze through all of that, like inertia and resistance, but you have to find now do we have a North Star where we need to get to? We need to chart a course, and that might include, yeah, we'll have to keep servicing bits of our existing stuff. And yeah, we'll have to keep talking to our existing customers. But you need to charter that course and get after it as quickly as you can because the ship is sailing like where we're four years into AI. You know, it's high time that you should be like moving. And I think you maybe have 1-2 years left before it's too late. And I say that like depending on the category, if you're in, if you're in mobile apps, whatever, you, you, you're dead already. But like, you know, in enterprise B2B software, you probably have a bit more time. Like the market's still shaking out. Three years ago we talked Cursor was #1 two years ago, we thought it was Windsurf. Last year we heard it was cursor again or Devon. Now we think it's cloud code. It's a perpetual bottle. Like some of these things are still shaking out. And maybe you're in one of those categories. And so, yeah, you still, you might still have a bit of time, but but you do, you might have time, but we don't have this comforter luxury. So you need to move quick and move hard. Speaker 4 Awesome, this. Thanks a lot for being here. We can people follow you, especially since we're going to launch a new book. What's the best platform? Speaker 1 Yeah, I mean, you'll probably hear about true all things fans so like Finn dot AI and we're Finn on all the domains. And then I myself, I'm das trainer and it's just DSTRAYNOR. I'm just das trainer basically everywhere it doesn't account to be made, I'm das trainer on it. So X and LinkedIn is where I'm most active. Speaker 4 Awesome. Thanks a lot this. Speaker 1 Cool. Thanks a lot.

Podcast Summary

Key Points:

  1. AI is fundamentally changing customer service, making traditional help desk models obsolete and forcing companies to adapt or risk irrelevance.
  2. Intercom’s rapid development of its AI agent, Fin, after ChatGPT’s launch demonstrates the necessity of agile, experiment-driven approaches to building reliable AI software.
  3. Success in AI requires a shift from traditional SaaS development—focusing on deterministic features—to iterative testing, rigorous validation, and continuous monitoring to ensure real-world reliability.
  4. Companies must avoid overhyping AI capabilities before thorough testing, as premature launches can damage customer trust when products fail under real-world conditions.

Summary:

The discussion centers on the transformative impact of AI on customer service and software development. Des Traynor of Intercom emphasizes that AI, exemplified by their agent Fin, is rendering traditional help desk models obsolete, compelling companies to innovate or face decline. Intercom’s swift pivot after ChatGPT’s launch—developing Fin within 15 days—highlighted AI’s potential to handle customer queries faster and more efficiently than humans.

However, building reliable AI requires a fundamental shift from conventional SaaS practices. Instead of deterministic feature development, it demands experimentation, rigorous “torture testing” to validate capabilities, and continuous monitoring post-launch. Traynor warns against overpromising AI solutions without thorough real-world validation, as premature releases often lead to customer disappointment.

The success of Fin, which now automates most of Intercom’s support, underscores the importance of this iterative, reliability-focused approach in delivering AI that truly works.

FAQs

Fin doubled Intercom's growth from 10% to 25%, generating $343 million in annual recurring revenue and reinventing the $1.3 billion company within 18 months.

Intercom realized AI could answer customer support questions perfectly in seconds, in any language, and operate 24/7, which was a better experience than human teams could provide.

AI development requires experimentation to determine reliability before productization, unlike deterministic SaaS features. It involves testing capabilities, monitoring real-world performance, and iterating based on outcomes.

Intercom uses 'torture tests' with hundreds of scenarios to verify the AI performs at or above human quality in internal testing before considering it launch-ready.

They focus on rigorous testing, real-world monitoring, and iterative improvements to increase reliability, moving from 23% automation at launch to about 70% over time.

If AI fails with real-world data after marketing hype, it damages customer trust and exposes the product as ineffective, unlike controlled demos.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.