Go back

Stripe's Payments Foundation Model: How Data & Infra Create Compounding Advantage, w/ Emily Sands

84m 16s

Stripe's Payments Foundation Model: How Data & Infra Create Compounding Advantage, w/ Emily Sands

This podcast episode features Emily Sands, Head of Data and AI at Stripe, discussing the company's innovative "payments foundation model." Unlike standard language models, this AI is trained specifically on Stripe's dense, structured transaction data, treating payments as a unique modality. Its key strength lies in analyzing complex contextual sequences involving all parties in a transaction (e.g., buyer, card, merchant) to detect patterns invisible to humans or traditional models, dramatically improving fraud detection and other optimizations. Stripe's implementation strategy is notable: instead of rebuilding systems, they provide the model's output embeddings as inputs to existing machine learning systems, making everything work better efficiently. The conversation highlights how this model creates a formidable competitive advantage, as Stripe's scale generates the necessary data, leading to superior products that further entrench its market leadership. This raises broader implications about AI favoring large incumbents and the potential for similar domain-specific foundation models in other industries. The discussion also touches on Stripe's use of LLMs for data enrichment and ensuring reliability in their AI-powered products.

Transcription

14694 Words, 82821 Characters

English
This podcast is supported by Google. Hey folks, Stephen Johnson here, co-founder of Notebook LM. As an author, I've always been obsessed with how software could help organize ideas and make connections. So we built Notebook LM as an AI first tool for anyone trying to make sense of complex information. Upload your documents and Notebook LM instantly becomes your personal expert uncovering insights and helping you brainstorm. Try it at notebooklem.google.com. Hello and welcome back to the Cognitive Revolution. Today my guest is Emily Sands, head of data and AI at Stripe. The programmable financial infrastructure company that in 2024 processed $1.4 trillion in payments or roughly 1.3% of global GDP for everyone from solo entrepreneurs to the Fortune 100 and which continues to grow at a blistering pace. We begin by discussing the many fascinating details of Stripe's new foundation model for payments and how Stripe is using this model to deliver improved performance across their broad suite of products. While it might seem unassuming at first glance, I would argue that the payments foundation model has several important lessons to teach us. First, while payments are represented in text, the payments foundation model is not a language model in the familiar sense. On the contrary, payments are treated as a distinct modality and importantly, no payment is an island. To properly understand a single payment, require Stripe to assemble extensive context, including recent activity associated with multiple entities, the buyer, the card, the device used to make the purchase and the merchant. So much context quickly becomes overwhelming to humans, but this is exactly where neural networks can shine. And indeed, when Stripe first deployed this model to detect card testing, which is a process that fraudsters used to determine which stolen cards actually work, they saw a jump in their detection rate from 59% to 97%. Obviously, it hit massive wind, not just for Stripe, but for the entire e-commerce ecosystem that collectively bears the cost of fraud. Now, if you've listened to this show for a while, you know that one of my pet theories is that the surest path to superintelligence is to integrate today's reasoning models with models that are trained on other modalities that humans aren't well adapted to understand. I'd say it's safe to say that the payments foundation model is superhuman when it comes to understanding payments. And this conversation left me wondering how many other businesses are training foundation models on their own modalities, as well as how many other interesting modalities might still currently be hiding in plain text. I've been imagine that this proprietary modality strategy might work on any number of domains, including health, cybersecurity, logistics, energy, and insurance. But to be honest, I haven't found too many other examples of this strategy being used today. So if you happen to know of any other foundation models being trained on any interesting proprietary modalities, please do ping me and let me know as I would love to do more episodes exploring this theme. The next lesson, perhaps as important to Stripe success as the model itself, is the way they are using it. Rather than trying to design the foundation model to support all use cases directly, they are exposing payment foundation model representations, and thus allowing engineers to use them as additional inputs to the many classification and other ML systems that they've already developed. The richness of the foundation model signal makes everything else work better, but doesn't require a major rethinking of existing systems. Again, outside of social network companies, who I do believe make their user and content representations available in this way, I've not heard of other companies taking this approach, and it seems to me now that more of them should consider it. Finally, the most important lesson from a societal standpoint might be that AI strongly favors the incumbent platforms that have the data necessary to train such differentiated models. The flywheel that Stripe has created here, which translates their incredible scale to commercial advantage, is allowing them to reduce the cost of fraud for their customers even as fraud is rising across the broader ecosystem. This makes Stripe the obvious choice going forward, which in turn further strengthens their data advantage and product lead. It is genuinely hard for me to imagine how anyone, aside from a few of the world's largest technologies, could ever compete with Stripe, meaning that even as history begins to unfold at a dizzying pace in many respects, competition in many key markets may effectively come to an end. This isn't necessarily a problem. I've never supported punishing companies for their excellence, and I've never been convinced that we should break up American tech companies. But it does seem like something that policymakers will need to think long and hard about as they envision the AI future and hopefully begin to imagine a new social contract. There's a lot more in this episode besides these key strategic insights, including how Stripe is designing processes to iterate quickly enough to stay ahead of fraudsters, including by using LLM as judge to fill in missing data. How they ensure reliability in their LLM powered talk to your data product experiences, how developers can accelerate product development by treating Stripe as their payments database of record, what Emily and team are seeing in agentech commerce today, and how they think about scoping their AI ambitions and investments. All in all, as you might expect from Stripe, it's a high alpha episode, with practical lessons for rank and file AI engineers, and big picture implications for executive level AI strategists. So, without further ado, I hope you enjoyed this deep dive into how smart use of AI is transforming one of the world's most critical financial infrastructure companies. With Emily Sands, Head of Data, and AI at Stripe. Emily Sands, Head of Data, and AI at Stripe, welcome to the cognitive revolution. Thanks for having me. I'm excited for this conversation. Stripe obviously is a global recognized leader in payments and doing some really interesting things in AI with high standards everywhere, and obviously a lot of shared DNA with some of the big frontier AI developers. So, a lot to get into today. For folks who want to do a deeper dive into Stripe and the payments ecosystem, you did maybe six months ago now, a podcast with our sister pod complex systems, with patio 11. patio 11. And I would definitely recommend that for folks who want to do a deeper primer on the payments world, which is a fascinating and fizzing team, one with many rabbit holes to go down. We won't do nearly as much of that today, we'll kind of stay more focused on some of the cool new AI stuff that you guys are doing. But maybe just for like a super quick primer, how would you describe the role that Stripe plays in the economy and then we'll use that as jumping off point to get into the AI stuff? You said payments infrastructure, we started as payments infrastructure. Absolutely true. We now build broader programmable financial infrastructure. So in plain terms, we give any business, right? It could be a teenager who's selling a figment template or it could be, you know, any one of now more than half of the Fortune 100 that run on Stripe. The rails and intelligence to move money online and to grow faster. So last year, companies processed $1.4 trillion through Stripe and we'll talk about AI today. You know, every one of those charges becomes training data for the AI systems that we'll talk about. But that flywheel also means that we're no longer just the payments API. We optimize the entire payments lifecycle. Yes, the gory details are covered with patio 11, but it's the checkout UX, fraud prevention, bank routing, retries, even things like, you know, dispute paperwork so that businesses can really keep more of every hard $1 and scale up with very small teams. So we think of the tools we're building as structural growth tailwinds and we're already seeing it in the data. I mean, businesses on Stripe grew 7x faster than the S&P 500 last year. Wow. Okay. A lot of good nuggets there. I have been a customer actually for what it's worth since nothing, maybe even the earliest early days, but pretty early days, like at least 10 years, that I've been a Stripe customer with my company Weymark. So we've seen a lot of the evolution from the customer side. The biggest thing that has caught my attention in terms of what Stripe is doing with AI is the payments foundation model. And I'd love to just spend, you know, a good chunk of time really going into details on that, because one of the things that I have been fascinated with and kind of trying to see around the corner and better understand is to what degree are we going to get a form of super intelligence via AIs that become sort of natively capable of understanding potentially a huge range of different modalities. And, you know, people are familiar now with like image generation. Of course, we had like text images, we had images, images, image generation models. Now those have kind of come together in this really tightly coupled, you know, deeply integrated way with the nano banana and other recent innovations in that space. And so I have this theory that like one thing that people really in general underappreciate is the degree to which training on these other modalities of data is just going to create superhuman capability in these domains that are sort of familiar to us, but also in many ways like very alien. So maybe for starters, like what can you tell us about sort of the fundamentals of the payments foundation model? Like what does the data look like? Obviously this transaction data, but you know, give us more details. detail on that, what is transaction data when you really get into the weeds of it? - Yeah, and it's a good point. I mean, there's been a ton of coverage of sort of the large scale traditional LMs and a lot less coverage of domain specific foundation models of which the payments foundation model is one. For us, it's been really a step function change in the speed and quality with which we can deliver all of those optimization solutions I talked about in auth, in fraud, in disputes. At its core, it's a transformer model that turns every payment. So the tens of billions of transactions that run through Stripe into a compact vector. So it's like giving each transaction its own kind of latitude and longitude. And then once you have that map, you can use it for all sorts of downstream tasks, right? To figure out what's fraud, to figure out how to authenticate, to figure out what's a valid versus invalid dispute without having to train a new model from scratch every time. And I think, you know, what makes it work, you know, the reason you can build a domain specific foundation model in the payments context is Stripe scale. So we process about 50,000 new transactions every minute. And at that density, payments start to look in a lot of ways, not in always, but in a lot of ways like language. So there's kind of a syntax to a payment, right? There's the card bins and the merchant codes and the amounts. And then there's sort of an analog to semantics. So sort of how a device or card gets reused over time. And so in the same way that like language transformers are learning embeddings and words with similar meanings cluster together, the sort of premise of the payments foundation model is just what if every charge or sequence of charges, and we can talk about that too, had its own vector in a similar space. So the inputs are, you're right, just the raw payment signals as they come in, the card details, the merchant categories, the IPs, but also those sequences. So what a given card or device or merchant bin or customer has been doing in the last few minutes or the last K transactions. And it's actually that history that turns out to be a huge unlock. And then from those inputs, the model produces an output which is just a reusable embedding, right? It's a dense vector for each payment or short sequence. And then we can layer lightweight classifiers on top for real time detection. We also have a slower higher latency variant that generates explanations through a text decoder. And I think we'll get to a stage where that can be real time ish as well, but we're not there just yet. - Cool, okay, there's already a number of interesting things there. In terms of scale, the blog post that introduced the payments foundation model said tens of billions of transactions. And then it also indicated hundreds of subtle signals. Is there a couple of examples of the long tail of these signals that kind of illustrate just how much information the model is able to ultimately take in that might be hard for a person to rep 'cause we can classically handle seven items in working memory, right? So what are we missing with our feeble human working memories that the model is able to take in? And from there, I'm kind of interested in the overall scale of data sounds like it's getting into the trillions of tokens, which would be like not at the high end of text foundation models, but not too far off, maybe like one order of magnitude less. So I wanted to just kind of sanity check my estimates with you on that. - Yeah, your math is legit. I'll answer the second question first and that first question second. Yes, your math is legit and the data is very different from the freeform text that you'd use to train a model to write like Shakespeare, right? So payments data is highly structured and dense and so we actually build a custom tokenizer that compresses the numeric and categorical signals really efficiently. So yes, the data set is big, but it's also packed with kind of this purpose built information that's incredibly rich for the set of tasks that we care about in our context. You ask about like what's hard for a human to eyeball. I think the thing that's hardest for the human to eyeball is the looking across those dimensions, not within any one payment, but within any, combinatorial sequence. So if you think about it, what you need to look at in order to figure out if, say a fraud attack is happening or how to get a payment authenticated, has very little to do with that particular transaction and everything to do with where that transaction sits vis-a-vis the transactions that have come around it. So you're not looking at like a single screen, you're looking at like a clip of a movie, but there are a lot of different clips that include that screen that are relevant to look at. Like you wanna know what I was doing, you wanna know what the merchant was doing, you wanna know what my card was doing, you wanna know what my IP was doing. And so that's really where the model sings, making it really efficient, not just to look at the individual payment, like it's hard to do with the scale of 50,000 a minute, but you know a human code, I suppose, if you had it a few minutes, it's really about the sequences that make the problem intractable for humans, but also very hard for sort of traditional MLA purchase where you have to hand engineer features to capture what's happening in each of a range of different sequences. And so in our context, the foundation model pays off dramatically because it expands really three things. One is how much data we can learn from, like we can learn from literally all of Stripes history, not just a task-specific subset of history. It changes how richly we can learn, right? So these dense embeddings capture very subtle interactions that manual feature lists, like counterfeatures wouldn't capture. And then third, which is more about sort of how we work internally, but it changes how efficiently we can build, right? Once you have a shared embedding, then spinning up a new model becomes a weekend project, not a quarter project, and that means we can sort of open the aperture for the types of ML-powered solutions we can build. - Yeah, cool. I really like the idea of kind of multiple clips. So I take it that that just basically reflects the reality that there are obviously multi, there are multiple parties through any transaction and I'm kind of inferring that like the pattern of behavior of each of those different parties is really where the strong signal is. It's not the, if you looked at this particular transaction in isolation, you might not get much, but when you combine recent history for all of the parties to a single transaction, the combination of those recent histories is really what tells you what you need to know. Do I have to have right? Like there's nothing about me using my card in Boston that tells you it's fraudulent, but if I just use my card on my device at my home IP, which is by the way, at Palo Alto, not in Boston, and tend to be buying on things that are totally different from what you suddenly see someone doing in Boston, that's a red flag that that's actually like fraudulent use of my card. Or conversely, if you see someone using, you know, rotating across a small number of cards to buy thousands of accounts from a given AI provider, maybe the card is truly there, but you're almost certainly gonna see some sort of reseller refund abuse happening where they're trying to steal your compute. And so it's, and it becomes more complicated when you add more entities like the merchant where there can actually be internal collusion happening. And so you're exactly right, like it's not about how anyone entity acts in isolation. It's about like an entity is an individual or a card or a merchant, and then that's like the node, right? And then the edges are the transactions. And it's like how, how much sense do these edges make in relation to each other and in relation to the combination of edges that we've seen in the past? - Does that go out? - I can imagine that that web could extend easily farther or you could sort of imagine, including sort of the rendered judgment on previous transactions. So for example, if I am trying to buy something from you and the model is looking for the signal of fraud, you could also say, okay, well, all these transactions that you have recently done as a seller, maybe you just sort of have the determination of they were fraud or not fraud, or you could even look at, okay, well, who are all those buyers? I'm like, what's there? So how far out does this sort of path through the graph have to go to get you the, what is the sort of shape of the curve in terms of scale versus diminishing returns? - Yeah, totally. So for a lot of the, so I talked about the scale of the Stripe Network, right? It's like this $1.4 trillion, but it's not just a big network. It's also a very dense network. So for example, 92% of cards that a merchant sees for the first time, Stripe has seen before on another merchant. So, okay, well, in those cases, you don't have to do very many hops, although you do want to validate that nothing's changed about the card or how the cards being used in the time sense. But fraud and conversion are kind of tail events in some sense too, right? Like if you can get 1% more conversion or 1% or 2% less fraud, it goes a long way. So you get really far from the dense network, but you also want to be able to try to, for more novel traffic that you see. Hey, we'll continue our interview in a moment after a words of our sponsors. AI's impact on product development feels very peace-meal right now. AI coding assistance and agents, including a number of our past guests, provide incredible productivity boosts. But that's just one aspect of building products. What about all the coordination work, like planning, customer feedback, and project management? There's nothing that really brings it all together. Well, our sponsor of this episode, Linear, is doing just that. Linear started as an issue tracker for engineers but has evolved into a platform that manages your entire product development life cycle. And now, they're taking it to the next level with AI capabilities that provide massive leverage. Linear's AI handles the coordination busy work, routing bugs, generating updates, grooming backlogs. You can even deploy agents with Linear to write code, debug, and draft PRs. Plus, with MCP, Linear connects to your favorite AI tools, Claude, Kerser, ChatchyBT, and more. So what does it all mean? Small teams can operate with the resources of much larger ones and large teams can move as fast as startups. There's never been a more exciting time to build products and Linear just has to be the platform to do it on. Nearly every AI company you've heard of is using Linear, so why aren't you? To find out more and get six months of Linear Business for free, head to linear.app/tcr. That's linear.app/tcr for six months free of Linear Business. Build the future of multi-agent software with agency, AGMTCY. Now in open source Linux Foundation project, agency is building the Internet of Agents. A collaboration layer where AI agents can discover, connect, and work across any framework. All the pieces engineers need to deploy multi-agent systems now belong to everyone who builds on agency, including robust identity and access management that ensures every agent is authenticated and trusted before interacting. Agency also provides open standardized tools for agent discovery, seamless protocols for agent-agent communication, and modular components for scalable workflows. Collaborate with developers from Cisco, Dell Technologies, Google Cloud, Oracle, Red Hat, and 75 more supporting companies to build next-gen AI infrastructure together. Agency is dropping code, specs, and services no strings attached. Visit agency.org to contribute. That's agnTCY.org. So, architecturally, this sort of reminds me a little bit of some of the stuff that Meta has done with their joint embedding video models. I'm not sure if that is the right intuition for me to have, but it does seem like there's a clear difference here where you're not trying to predict the next token. It seems like it would be more of a dedicated. So it's not like an autoregressive type model. It seems like it would be more of a dedicated channels, like a mask type situation where you could imagine doing a training setup where it's like mask out whatever kind of randomly and have the model learn to fill in details. Yeah, exactly. So our V1 did use a sort of BERT-style, like, mask modeling setup that you're talking about. And then we paired that with a second stage, which is explicit similarity fine tuning. And so most of the heavy lifting there was like, "Okay, let's curate the right sequences." To learn from, let's build the right encoding, let's do post-training with that kind of similarity objective so that near neighbors in this payment space cluster together and the oddball separated, and you can start to a reason about those oddball clusters. And again, the big unlock here was modeling short histories, right? So what a Carter device or merchant bin is doing over some number of minutes or last-k transactions rather than the isolated payment. Now in the, we call it V1.5, but we're actually moving towards encoder decoder setups and compressed memory sequence. So like a few vectors together, that makes it actually easier to catch subtle abuse in real time, because you're not averaging across noise. You're just stilling the full story into this compact representations. And so yeah, the mental model is like V1, mask modeling, plus similarity training. V1.5 is compression first with a tight sequence embedding. And then we can put kind of lightweight task-specific heads on top, which are for the charge path use cases. If you think about the charge path, you got to get the job done in tens of milliseconds at most. And so those lightweight task-specific heads are important for latency and speed. Yeah. Can you say how big the model is? I mean, the 10 milliseconds doesn't allow it to be that big, I would assume. Although, it don't have to do a lot of steps in poor repass. The cuff, but the task-specific heads on top are small. And I think when you reason about it, all you have to be able to do actually is place the new charges and sequences as they come through in this dense embedding space as they come in, which is a much easier problem than obviously like the upfront training. Yeah. And it's also just one forward pass of the model, as opposed to having to generate a whole sequence. So you can, that one pass nature of it definitely helps with the latency as well. Yeah. That's really interesting. It also reminds me of one of the first language vision language models that I studied deeply was the blip family of models. And I remember that they had really amazing success with a frozen language model and then also a frozen vision model. And just trained a few million parameter connector between the two to sort of bridge from one latent space to the other. And of course we've gone way past that now in vision language, but this was like an early 2023 thing. And you were able to get like really quite good captions out of that setup, even though neither of the foundation models that were used had anticipated that use case. So it's not like you're lip in a while, but that's, yeah, that's an interesting, that's an interesting analog. So it, but it sounds like you've kind of created a similar situation where people internally at Stripe can say, "Okay, I have a new use case idea for this. I can train something really small." You mentioned it can be like a weekend project instead of a multi-month project. And if I understand correctly the ideas, because you've got the foundation work done, you can train a few million connector or classifier head, whatever you want to call it, but very quickly, that step becomes a rapid iteration step. Yeah, exactly. And actually, it can be even simpler, which is, so the embeddings themselves, where most people start actually is most most modeler start, is just taking the embeddings themselves, which are stored in shepherd, which are feature engineering platform, and literally just adding them as a feature to existing models. Like whatever those existing models are, and saying, "Is there an add-in signal from these embeddings?" And I would say, "You get some false negatives there. We're obviously just shoving the raw embedding into some number of soda models that have been iterated on over the last six years, isn't going to produce uplift." But in other cases, where it's sort of a lower priority model that's only in its V1 state, like you actually do get something straight out of the gate, and you can start to reason about, we've been talking a bunch about the payments embeddings, but how much signal do I get from like, understanding the payment better? How much signal do I get from understanding the customer better? How much signal do I get from understanding the merchant better for each of these use cases? And then that's also motivated where folks have kind of, for which applications folks have leaned in harder? Yeah, that's really interesting. So just to make sure I understand that correctly, you've got obviously started to run around for a number of years. There have been many types of problems that you've brought machine learning to over time. Typically with a more classical feature engineering type of approach. Yeah. And. And. And. And. And. And. And so now the foundation model embeddings can become just tacked on as additional features, rerun that training and immediately back test against your set, and then you're like, "Okay, cool. We just made this better kind of for free because we were able to get additional signal." That's really interesting. I mean, our system handout application wasn't that, right? It was car testing where we literally like, it was a whole new approach to car testing with the foundation model. I think car testing is like, you know, fraudsters trying hundreds of tiny authorizations, sort of iterating across stolen cards or literally just doing raw enumeration, just like trying a bunch of cards. And they bury those attempts inside floods of legitimate traffic, right? A big retailer can have hundreds of thousands of charges come through, and then there's like a couple hundred or maybe a thousand peppered fraudster charges of like 30 cents or 50 cents. And like classic models couldn't really pick up those kind of needles in the haste act. So our first application of the foundation model was just treat those sequences again, like frames in a movie, right? And suddenly these 200 nearly identical requests, like, you know, same low entropy user agent, rotating across proxies coming in about every 40 seconds, like they light up as an island in the embedding space and they get blocked. And the impact of that one was huge, like our detection rate of car testing at large merchants went from 59%, which is, you know, not bad, but not great to 97% from that change. But then as we started reasoning about where else could it be useful? Yes, just exposed the embeddings and let them be added as features to these traditional single task models was the next step. That was never intended to be the final state, but it's a way to get signal on where is their incremental value or incremental signal from these embeddings that requires very little lift. Yeah, fascinating. That's a very modular approach to AI deployment and I can't recall hearing any organization that has had a similarly modular structure. Maybe meta comes to mind as another one that might have a sort of user model that could then be bridged over to any other space or problem that you might want to apply them to. This is like a fairly uncommon setup, I would say, right? Are there others in the. They'll know exactly what happened or how it went or how it worked, but I know because I know the former leader, I know there was a cortex org at Twitter that was basically doing horizontal models. Again, don't know exactly how it worked or the architecture. We've been talking about it in the context of payments, but we've done the same thing over the last year in the merchant space where we have, it's called the merchant intelligence team, but it basically has this MI serve. So it can go out and find anything in the web about a merchant and generate embeddings and be used to answer questions. Those merchant embeddings are also features in downstream, for example, merchant risk models, but in it's a service where the model owner can ask merchant intelligence, the agent, to come up with a more custom embedding or a more custom embedding that maybe you want to know what payment methods the merchant offers or whether they have anything that's counterfeit. That's actually been another horizontal layer that's provided a ton of leverage for Stripe, right? Historically, you've got a lot of use cases. You want to know things about the merchant to understand supportability, whether they meet the requirements of the current networks and the issuers and the banks. You want to know whether or not they're fraudulent. You want to know where they've had an account takeover. You want to know if they're credit worthy. You want to know whether or not we should give them Stripe capital like a loan and you want to figure out whether you should be going to market with them, like all sorts of things. You want to know about a merchant and historically teams at Stripe were, you know, when LLM's hit the scene were like out building their own sort of custom versions of this, but what we realized is there's actually just like one service now that does that much more efficiently than everyone rolling their own. The better lesson strikes again. I got a lot of different directions I want to go, but where does ground truth come from on some of these questions and how long does that take? Because I sort of imagine, especially in a fraud detection situation, right? I mean, fraudsters, I always assume, are going to be some of the most clever people in the world, diabolically so, but nevertheless, like, you got to respect the smarts of some of these folks, right? So, I assume that they are very savvy to like real world events. You mentioned, I think in the conversation with Patrick, you know, somebody might have a flash sale and that sort of spike, like obviously you don't want to turn them off when they're having a flash sale because that's, you know, horrible experience and loss of business for the company running the flash sale. But at the same time, like that sort of potentially a really good target for a card tester to come in and try to do whatever it is they want to do. You're sort of, I imagine in a kind of eternal arms race between fraud and fraud detection. And then what I know, what little I know from my experience as a consumer and as a business owner is like to actually close the whole loop and get to the point where whether this was fraudulent or not has actually been set in stone. Like that's a long process, right? So, well, if a long process, if you even get to a definitive answer, but something like card testing, right? Like, actually, I said the first thing we did with the foundation model was to play it for card testing. Actually, the first thing we did for the foundation model is to play it internally for card testing, pass those labels to internal expert humans, have them go and validate the labels, then feed the validated labels into our traditional ML model for card testing. And suddenly our ML model, traditional ML don't deploy the foundation model. Our traditional ML model for card testing started doing way better because finally it had like a more comprehensive source of truth for the labels. So I was actually the first version, although hadn't had revealed that fun fact before. Yes, attackers iterate and so do their models. And so our job is just to iterate faster. And we are, and I'll talk about some of the ways we get around the late arriving labels or the missing labels altogether. But just to give you a sense of like how we're comparing to the fraudsters, like industry-wide e-commerce fraud is up, I think it's up like 15% year and year. But the dispute rate for the businesses that are running on stripe are down 17% year and year. And that's because we, in a bunch of different ways, and I can give a couple of my favorite recent examples, are just consistently shortening the loop between like new tactics shows up and like defenses go and adapt. And that sort of loop shortening is happening in production and in some cases in real time. So an example that our users are getting a ton of value from that we recently released is dynamic risk thresholds. So it's basically like, you know, radars out, they've got their threshold score, block stuff above the threshold. But then when an attack starts, radar learns and attack starts and it tightens the defenses, right? And that kind of throttles and that allows like, okay, you know, revenues flowing freely when you're not under attack, but then we're much more aggressively blocking when an attack arises because again, like an attack is never, almost never a single event. It's almost always like a true cluster. And in that case, like, you know, the model is learning the policy of how to act. Now it's not learning that policy online just yet, but it's learning the policy of how to act. And that's another powerful tool. And I think, you know, in payments, it's easy to think like, I put in my credit card and then just like an objective decision is made to block me or not. But that's not actually true. And so we've been leaning in harder on what we call soft blocks. So adaptive 3DS is an example here at like applies that 3DS authentication. Like if you're in the US, most of the time you don't get 3DS. But we can- You tell me what that is because I don't feel it. You might have defined it in the in the conflict systems, but if so, I could use a real question. And you're just like, you have like a second, like a sort of like a two factor auth that you would have. It's like a two factor auth-esque experience where you're verifying that to the credit card, the bank network or the credit card issue or that it is in fact you. And this is very, very common in Europe and it's very, very common, very, very uncommon outside of Europe. And by the way, when it does happen, it often creates unnecessary friction. And so part of what we do at Stripe is figure out when we need to authenticate and when we don't. But also with adaptive 3DS, we are pushing for authentication selectively in cases where we have a sense that the charge may not be good. And so instead of just having this binary decision of block, don't block, you have this other arm you can go down, which is, you know, hit them with 3DS. And what ends up happening is the good guys get through the 3DS because they're excited to buy the thing and their legit users and the bad guys do not. And so a lot of the AI companies are using this. You know, AI companies being hit with fraud is extra painful because their marginal costs are high, right, unlike for SAS companies who care a lot less. And so like early adopters of adaptive 3DS, we're like 11 labs and character AI. And they're able to just dramatically cut down fraudulent disputes without any effect on conversion because 3DS isn't super heavyweight for the end user. In fact, US checkout users, so it's a little different in Europe because a lot of those folks are already 3DS, but US checkout users saw 30% average drop in fraud. And they just like turn this on with a single click in the dashboard. And then it lets us basically like learn the policy of who is worth 3DSing to balance conversion and fraud to maximize their profits. Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Amthropic, Makers of Claude. Claude is the AI for minds that don't stop at good enough. It's the collaborator that actually understands your entire workflow and thinks with you, not for you. Whether you're debugging code in midnight or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. Regular listeners know that Claude plays a critical role in the production of this podcast, leaving me hours per week by writing the first draft of my intro essays. For every episode, I give Claude 50 previous intro essays plus the transcript of the current episode and ask it to draft a new intro essay following the pattern in my examples. Claude does a uniquely good job at writing in my style. No other model from any other company has come close. And while I do usually edit its output, I did recently read one essay exactly as Claude drafted it, and as I suspected, nobody really seemed to mind. When it comes to coding and agentic use cases, Claude frequently tops leaderboards and has consistently been the default model choice in both coding and email assistant products, including our past guests, Repplet and Shortwave. And meanwhile, of course, Claude continues to take the world by storm. Anthropic has delivered this elite level of performance while also pioneering safety techniques like constitutional alignment and investing heavily in mechanistic interpretability techniques like spars auto encoders, both internally and as an investor in our past guest, Goodfire. By any measure, they are one of the few live players shaping the international AI landscape today. Ready to tackle bigger problems? Sign up for Clawed today and get 50% off Clawed Pro, which includes access to Clawed Code when you use my link Clawed.ai/tcr. That's Clawed.ai/tcr right now for 50% off your first three months of Clawed Pro that includes access to all of the features mentioned in today's episode. Once more, that's Clawed.ai/tcr. In business, they say you can have better, cheaper, or faster, but you only get to pick two. But what if you could have all three at the same time? That's exactly what cohere, Thompson Reuters, and specialized bikes have since they upgraded to the next generation of the Clawed. Oracle Clawed Infrastructure. OCI is the blazing fast platform for your infrastructure, database, application development and AI needs, where you can run any workload in a high availability, consistently high-performance environment, and spend less than you would with other clouds. How is it faster? OCI's block storage gives you more operations per second. Cheaper? OCI costs up to 50% less for compute, 70% less for storage, and 80% less for networking. And better? In test after test, OCI customers report lower latency and higher bandwidth versus other clouds. This is the Cloud built for AI and all of your biggest workloads. Right now, with zero commitment, try OCI for free. Head to oracle.com/cognitive. That's oracle.com/cognitive. So one way I like to frame, excuse me, some of these conversations is just in terms of like practical lessons that people can apply in their own AI pursuit. So one takeaway there is add middle ground outcomes to your classifiers so that they're not binary, but try to find that sort of middle space if one exists that can, or something other than the model itself can step in to help resolve the most challenging cases. It's almost like Cloud, you know, now these days can sometimes end a conversation, right? If it has to-- A mental information could help you, and you can get it from your users in a low-cost way. Don't constrain yourself to being a modeler, like be a product thinker, and go figure out how to get that information, right? And the model is really good at deciding, and then that's not--forced. You don't require that additional information from everybody, but then let the model decide where it needs more information and where it doesn't. Yeah. On the adaptive threshold concept, this suggests a state, basically a sort of state world state that is maybe being fed into. I assume it's not like that the model itself is calculating that on the fly. This would be a more like global variable sort of thing that the model would receive, or-- No, so it's always so--actually like, hey, this merchant is starting to see clusters of scores creep up. Like isn't that interesting? And when we look at the subset of transactions that have those higher scores, maybe they're still below the block threshold, but they're looking elevated. Is there anything about those that looks like it's something collusive or coming from a small number of attackers or some rotating across IPs or coming from a geography that they haven't seen before? And then once we get signal that looks like there's a slice that's in attack, you can actually start to lower the threshold from that subset for what it takes to block. So you-- But that's all happening with sort of the same short histories that you described previously or there's like a longer--it just sounds like there's a longer history at some point coming in to inform that kind of decision, but maybe not, I guess the short history could be enough? There's a longer history with less of the individual transaction level. Like, it's basically detecting anomalies and slices of traffic, right? So like, this geo, this--these bins, this cart size, like something anomalous is happening here. That anomalous thing has kind of elevated rescores. Hey, like, it reads a bit like an attack. And actually what's interesting, I mean, rules are really good in a lot of ways, right? And so, you know, maybe another general lesson is like, rules are good, but they're also blunt. So figure out where you can blend rules with models. I mean, you'd asked earlier when disputes actually come in, like, disputes are super lagged. They can take days. They can take months, right? I'm the card holder. I have to like, see my bill. Like, notice I didn't buy the thing. All my bank, my bank has to go and like, file within that work. And so those labels for sure arrive late, but we don't wait. We use proxy signals. And those sort of weak labels show up way earlier, all the way to real-time issue or feedback. And it can be very--so real-time issue or feedback would be like, CVC mismatch, right? Like, the CVC code, you know, the little 3/4 digit credit card code doesn't match. Or like, the zip code doesn't match. It'd be easy to write like a blunt rule that said if the CVC doesn't match or if the zip code doesn't match, block it. But you'd be blocking a bunch of good revenue because like, who doesn't? Sometimes, fat finger, they're CVC or there's a code in a hurry or on their phone or whatever. And so we have these risk-based radar rules, which are like, okay, take the model score, combine it with the issuers real-time responses, and make a decision based on that intersection. So like, if it's looking marginally risky and the CVC is wrong for sure block, but if it's like a pretty known good user and they fat finger to thing, like, let them through. And I think that blend of rules and models is--these easy for modelers to put their nose up at rules and it's easy for rule makers to put their nose up at models that aren't fully explainable, but in plenty of contacts blending the shoe actually does far better. So you actually do let transactions go with a wrong three digits. What was that called, if you sued? Nathan, like, I know you're good. You've bought from this person before. Maybe you even use the same credit card. You're coming from a legit IP and like, I feel good about you in a lot of ways. And yeah, the issuer comes back and says, "Hey, there's a mismatch." And we say, "Hey, let it through." And then by the way, once we let it through, we also have to get the issuers. We're to let it through. And there we actually have data sharing with the issuers where we pass them our risk scores so that they can also understand why we passed it through and that motivates them to also pass it through when they see our signals. So it's kind of a two-step. Very interesting. Let's go back to how you are tightening the iteration loop. Again, I think this is something that basically everybody who's developing AI products could stand to get better at. So what have you guys found to be effective needle movers in shortening your cycle time? Woof. I mean, this isn't one for us where it's like, there's some magical, you know, reinforcement learning that we need to be implementing online for every single use case. I think it's actually for us being quite context dependent. The things that matter are having enough labels and having good labels and having those labels fast enough. And actually, like, you can get pretty creative about what the label is. We talked about some examples. We also talked about human-generated labels. But another thing that we've been leaning into is LLMs as a judge. So especially for context where there actually is no source of truth. So a simple example. We've been talking a bunch about fraudulent disputes. But there's a lot of suspicious payments that never result in a fraudulent dispute, right? Maybe the person starts a free trial and then they cancel or they ask for a refund or they, you know, just spin up a bot account but never even get to the checkout page. That type of friendly fraud is actually really costly to businesses. And it's like almost half a business. I think 47% of businesses say friendly fraud, which is a total misnomer because it's not friendly, hurts them more, hurts their business more than sort of stolen card credentials or what most people think of as fraud. And that's actually, you know, there's a lot of AI companies running on strike. That cost of friendly fraud is particularly true for these AI companies, right? Very different than SAS. And they have, they have inference costs, they have compute costs, they have like very high marginal costs. And so when someone is engaging in free trial abuse or reseller abuse or refund abuse, it's super expensive to their, their unit economics. Anyway, so built on the foundation model, we now have these suspicious payments that we identify. So these are fraudulent e things, but not in the traditional going to result in a fraudulent dispute sense. And when we pass those over, for example, to the AI companies, we want to be able to describe to them why their flag does suspicious. So it has an enumerated email or it is like, cycling through a small number of IP addresses or whatever. So those labels are, those sort of those explainers are generated by the foundation model. But then the question of course is like, how do you know if they're right? And so we have this like LLM as a judge that sits on top that looks at every transaction label combination and asks, given everything you know about this transaction and everything you know about the cluster to which it belongs, how do you feel about the quality of the label? And what ends up happening is that there's, you know, a large share of labels that are good enough, trustworthy enough that we pass them over to name AI company, DeJour, to decision on. And there's some small number that are like too noisy and we like, okay, like we gotta go work a little bit to make that label stronger. But I call out that example because these are like, there's no source of truth. Like, like nobody, I mean, you and I could manually go through, I guess, but at the transaction level we're not going to. And so it's been really helpful to have LLM's kind of, as a judge where there's no clear nor star. - Yeah, that's really, so that's fascinating, but I'm still kind of confused about one thing, which is, well, I'm probably confused about a lot of things. But the thing I'm focused on being confused about right now is, when I try to advise people on AI broadly or when I kind of try to give people the lay of the land, one of the things I tell people is, AI's are not very adversarially robust. They are really good these days at the happy path, if you dial in the performance and you control the inputs, you can, in many, many cases, you can get to superhuman performance on routine tasks. However, if you don't control the inputs and you're exposing your AI system to the world, you do have to be mindful about the fact that the systems are not adversarially robust. People can usually find some weakness, right? And that's even been true. We did an episode once on superhuman go playing AI's that were beaten by really simple attacks that no human would ever fall for, but which the AI, even though it was superhuman when playing Go in the normal way against like other high quality Go players, it was just totally blind to this certain class of attack that was found through this sort of adversarial optimization. And so it seems like you would be in this environment where you've got it on like hard mode kind of everywhere, right? Because anybody can come test the system from kind of any position. You can't really deny people the ability to like try a payment. So they can kind of gray box you, right? They can test from a bunch of different angles and try to see like what's gonna get through, what's not gonna get through. And presumably there's always some vulnerability that you're not aware of that they can systematically attack or try to find through these sort of attacks. And my guess would then be like the only way to really deal with that is to just constantly be identifying and iterating, but that that sounds still hard. Despite everything you've told me, it still sounds hard to be as responsive as you would need to be given especially that the actual ground truth is so lagging. So like how do we not, maybe we do, but how do we not like just bleed a ton of money in one incident after another as attackers figure out that there's like some gap. And then just you know jam as much as they can to explain it for a while until it's closed. Like how does that not end up being a huge problem? - Yeah, a couple, a couple thoughts. So one is we expose capabilities through products and APIs to our users not through raw weights. So for sure anybody can try a payment and test and see if they can do a workaround, but just to be clear, like we're not actually exposing the model for them to, for them to you know, have an attack surface against, right? So I think that the products in the API is actually better meet the user needs and they also kind of narrow the potential attack surface. So just clarification one. I think for, you know, it's interesting to think like what's the relevant alternative? Like what is fraudsters job? fraudsters job is like find loopholes and exploit the system. Like that's what they make their money on. And so the relevant alternative isn't, you know, perfectly airtight. The relevant alternative is like baseline approaches. And actually when you start to think about foundation models or LMS, they, you know, the payments foundation model for example, like it's actually a lot more nuanced. The type of information that it's using to decision versus like you could think of like an early, an early transaction fraud model that's using like last seven day counters and like the fraudster figures out that like as long as I'm eight days out, I'm safe. I'm just gonna do everything like on day eight and then hit them hard and they go seven days back. Right? So so to some extent, I think sort of traditional ML is easier to get around, whereas the the foundation model is more comprehensive. But the other thing that we certainly have long done and continue to do is a layered approach. So it is not, there's not just a single set of defenses. There's a set of model defenses. There's a set of rule based defenses. There's the soft blocks I mentioned. There's the user's own defense set, which can also vary like all the way to how they treat you at sign up or how they block bots at sign up. And so fortunately for us, unfortunately for the fraudsters, they're not fighting against like one model. They're fighting against a whole system that is I guess until I set it opaque to them. - Yeah. What it's almost just about kind of any other big use cases, you mentioned like off-fraud disputes. There's some interesting talk to your data, product experiences in Stripe. What stands out to you as the most interesting applications? Not even necessarily from a like what move the most money, but like what would be most interesting to the AI engineers and the audience in terms of just interesting implementation details or surprises, quirky stuff that you've learned along the way. - Yeah, I mean we've talked a lot about sort of transaction level understanding and the path there, if you think about like modality, like it's mostly payments plus text. You have these like structured payment signals with and then like language using contrastive learning and you align the two and you got the text decoder and whatever else. But payments plus text is like only the start, right? So the system's actually designed so that new modalities are just considered like tools that the router on top can invoke. And so like if you wanted to add another encoder, maybe for financial time series, which I'm very interested in, but I don't have anything yet that I could share or for images, right? It doesn't require like a whole kind of rewriting of the system. It's just like a modular expansion. And I think the multimodality, like I'm starting to see it really shine at the merchant level. So actually yesterday I was testing two lightweight agents, either of these are in production. So you know, just full disclosure. But like the team has them in shadow and one crawls merchant sites to assess fraud and it's relentless, like incredibly relentless. The other spots counterfeit products. And it just does so like literal orders of magnitude better than trained human reviewers that we have it striped doing the same thing. So it'll find like there's a print shop and there's thousands of items in the print shop and the agent will like patiently zero in on like the one like spider Gwen sticker was like an example. I was there yesterday with like no sign of official licensing like route route. But then it also knows like, oh, this other site, the Canada goose that's like marked as with tags is second hand. And so it's actually fair game. So I think that kind of like that kind of multi modal roadmap is interesting. Not for a multimodal end of itself, not for the technology in of itself, but for where it's like gonna unlock, gonna unlock real value. - Cool. I'm the talk to your data thing in particular, that's something that a lot of people have tried to do for themselves when they've tried to use a product to do it. It strikes me that the, where most people have kind of gotten stuck there, right? Is like, well, I was able to get GPT whatever or cloud whatever to be like pretty good, but it still made some mistakes and I didn't really feel like I could confidently give somebody who wasn't a proper data analyst, this tool and be confident that they would get good insights out of it. So you guys have that problem, maybe the biggest scale in the world. How did you think about like, what is the right threshold of accuracy for I talked to your data model? I assume you didn't achieve 100% accuracy on this sort of thing. But what was the threshold that you felt you had to get to and what was needed to keep dialing in until you actually got over that threshold to where you could deploy? - So one of the reasons that this sort of talk to your data was interesting to us in Stripes context is like a lot of what a business wants to know is captured in Stripes data. Like, who's selling what, for how much to whom, who's retaining and churning their subscriptions, and so. And so that's thing one. Then thing two is the data is actually very well structured, because it has to be, right? Like it's generated from the transactions that are flowing through Stripe that are incredibly robust and well-documented, and the schemas downstream of that make sense and are well-documented as well. And so a lot of this talk to your data stuff, it's like the garbageing garbage out problem, where like my tables aren't well labeled, my fields aren't well labeled, and maybe like the underlying data actually isn't deduped. And so you can't really tell if the issue was like that sort of text to SQL, or if it was actually like the underlying data was bad and/or the data structure was not understandable. So we kind of were able to like leapfrog that, which is great. But still, LM, so I think you're referring to our sigma system, LM's do make mistakes. And so our approach there is actually like, if we have reasonable confidence, we'll provide it. But we, I don't know if you've ever used it, we overlay on top a natural language explanation of what we're doing. So like we thought you wanted to know, whenever you asked like, how did Black Friday this year compare to Black Friday the last two years? And then we'll be like, OK, these are the dates we use for Black Friday. These are the timestamps we use, because by the way, most things on StripeHap and in UTC, and many people aren't reasoning about their business only in UTC. We looked at Black Friday over the last three years. Here's how we computed the percentage growth. Like you're like, that's boring. Doesn't ever compute the percentage growth the same way. But that actually allows someone who's not a data analyst to build comfort in the output versus saying like, either you're lowering it, just like taking it and running with it, or throwing their hands up and saying, I can't trust anything, because I don't know what's happening under the hood. You just wrote a SQL query for me, but I have no idea how to interpret it. So I think when I think about talk to your data, it's like, well, is your data interesting to talk to? If yes, make sure it is well structured, well documented. And if it's not, invest in that before you invest in the natural language interface on top. And then just make sure that the LLM is explaining what it's doing, which they're, of course, very good at doing now. And that allows you to open the aperture a bit in terms of less certain questions you're willing to answer, because you know anyone can read the natural language and make a call on whether or not that was the right approach. For folks who want to do a double click on the process of getting the data into shape, the episode with the CEO of Illumix was really good on that. And just for what it's for you, they have built basically canonical structures of enterprises across like a bunch of different categories, like, drug company, for example. They've kind of built out a vast representation of data that in their study, the opinion represents like the canonical drug company. And then when an actual drug company comes to them, they do this painstaking process of mapping all of their actual data with all of its idiosyncrasies onto the canonical version that they've kind of made work well. And that mapping becomes the kind of cleanup process that gets them the reliability that customers obviously ultimately want. Pretty interesting. That sustainability-- Yeah, please. Well, I'm just going to say there's also, there's also kind of an interesting feedback loop here with users, right? So if you're a usage-based billing company, the types of metrics you want to know to reason about your business or to share with your investors are generally very similar to the types of questions that all the other usage-based billing AI startups also want to know. And so that both means that we can really make great the subset of questions that matter in a given domain. But also, forget the natural language to SQL interface or talk to your data. We can just push those commonly asked questions over onto the dashboard. And even benchmark you-- we have a smart benchmarking now-- benchmark you on those metrics versus a peer group. And by the way, that smart benchmarking is one of the applications of the Merchant Intelligence Service, which is like, figure out which websites are like this website in terms of would be good comps because have similar user bases and are at a similar stage of their development. Yeah, cool. There's a good pattern there as well for sure. I've been thinking about that in the context of agents lately. And there's kind of the Choose Your Own Adventure agent where you give it a bunch of tools. Here's some MCPs, whatever have at it. And then there's the sometimes better described as a workflow, maybe with a couple forking decision points that people also call agents in many cases. And I'm starting to see the emerging pattern be like, have that Choose Your Own Adventure agent sort of at the top level of user interaction? But then in terms of the things that it's choosing, make those actually pretty detailed workflows in a lot of cases where you know that as long as it makes the right choice at a high level, that the process that's going to be kicked off is one that you've really deeply understood, dialed in for accuracy, confirmed for yourself is going to work reliably. So I think that's another-- you're kind of talking about a push model instead of a pull, but nevertheless, there's a sort of isomorphism, I think, between those structures. Explainability is obviously huge. One thing I was interested in asking is, are you doing any mechanistic interpretability? Are there like spars autoencoder type things now happening on the foundation model so that you can learn in a semantic way, like what new features the thing is learning? Yeah, not literally. So we're not in dissecting individual neurons in the way some research groups are. I think our focus is really on making the outputs self-explaining in the way that we and our users, where it's user-facing, can actually trust. And so mechanistic interpretability is really important when you're releasing the full open-ended model into the wild. In our case, we control both application and the environment. And so our priority is that really practical explainability that's tailored to the payment's application. So when the foundation model flags a transaction, it doesn't just say, high risk, it says, gibberish, email, and enumerated name pattern and device concentration. And that's actually the layer of explanation that lets the fraud analysts, or even another system, like a follow-on agent, act confidently. And then as we were talking about a little bit ago, in many of those cases, there's actually no ground truth for that explainer. And that comes up more and more as we expand to new domains. So we're actually detecting fraud further up your customer funnel all the way when someone's creating an account on you and well before they're entering credit card details. And so that's where things like the LLM as a judge framework are really helpful. So it'll look at that and the cluster of similar events and the tag definitions and then output, how confident is the LLM? Basically judging how confident is it that the cluster really matches the label. So no, we are not peering in, neuron by neuron, but we are focused on interpretability at the output level and that's really valuable for us. And then for example, like if you're in your dashboard and you're seeing a bunch of suspicious users, you want to know exactly why we flagged them as suspicious so you can decide how to action. And so in our setting, that's really what matters most. Gotcha. You mentioned usage-based billing. And this just led me to a very practical question around how you would recommend people build on top of Stripe today. 10 years ago or so, when I first became a Stripe customer, it was already a respected company, but not such a foundational part of the economy as it's become. So we were like, well, we don't want to switch off of this one day or we might have who knows, whatever will have kind of our own database of all the transactions and all that kind of stuff. And Stripe of course have their view of it, but will maintain our view. That was a lot of work then with usage-based billing, it sounds like an even more challenging project now, especially for your proverbial couple people that are doing a hackathon and want to kind of get something started. Do you recommend that people-- the alternative I have in mind, which I wonder, ultimately, if you recommend, is like, could I just leave all of that to Stripe and just basically make nothing but API calls, trust Stripe to be like real-time ground truth across the board and not even have a financial side to my database, but just purely do that real-time API calls. Totally. Totally. Don't even have it. Like Stripe's APIs run at six nines of uptime. So they are safe to use as your system of record for very critical flows. Our usage-based billing APIs processed 100,000 events per second. And they've got all the built-in monitoring and alerting and invoicing. The alternative is also pretty painful, right? Building-- if you're going to build your own mirror of Stripe's data, it's pretty complex. You've got to sync across all the events. You've got to build your own monitoring systems. You've got to keep everything reconciled. And especially if we're talking about a startup, that's just a lot of work that doesn't create any kind of differentiated value. You've got a lot of these companies that are taking off have like 5, 10, 20 employees. They shouldn't be spending an ounce of that limited capacity on this stuff. And then on the flip side, if you treat Stripe as your source of truth, you get real-time signals that you can actually act on in the Stripe ecosystem, right? Like the billing threshold has been exceeded or whatever, without having to have this whole parallel system. And we talked about Sigma assistant earlier. But with products like Sigma and Stripe Data Pipeline, you can still run all your analytics and all your reporting without doing the job of building your own warehouse. You might be wondering about the downsides. Like historically, the biggest-- downside of not mirroring of like just leaning into Stripe was, okay, but what about when I want to join Basically Stripe data to like my own business objects and now you actually can so you can actually extend Stripe's objects with Metadata so like a lot of users will Whatever they're like attach their own order ID or shipment ID to an invoice and for you know most companies that closes most of the gap now It's different if you're a large enterprise and you've got a whole bunch of other follow-on systems that are off-stripe and data sources that are off-stripe But for most startups it's both simpler and just a lot safer to let Stripe be the system of record Yeah, cool. I imagine some companies have gotten pretty big revenue terms over the last half of many months while still doing just that I don't know if you would want to highlight any by name or if it's too secret but when I see the Curves you know from folks like lovable bolt repel it recently obviously things like cursor one starts to wonder You know in the head counts that they have I bet a lot of them are probably doing exactly that and just kind of Trusting which six nines gives you pretty good reason to trust Yeah, and you know that's just for the system of record right so lovable It's a great example they hit a hundred million in ARR in their first eight months and their stack is Basically a case study in all-in on strike right so they incorporated the business before they monetize they incorporated the business with Stripe Atlas they from the very beginning used our optimized check out suite So like the front-end customer facing services is our optimized check out suite which allowed them to localize payments in over a hundred countries and get Something like a hundred fifty payment methods out of the box They leaned on billing for subscription so they didn't build their own billing system or you know have to contract with another third party They leaned on link which is like our one-click consumer check out for fast check out by the way nerds Love to buy from nerds so like our concentration And on link of AI buyers is very very high they leaned on radar for fraud prevention They leaned on stigma for analytics and so really like striped took care of the financial plumbing and so that lovable Just really focused with that small team on on product and growth which of course they nailed but there's like small They're smaller ones right retail AI they build um Have you used them they're like coagents yeah, yeah for customer support and so they launched last year I think they have like you know over 10 million an error in their first year and mentioned link concentration Link actually powers 38% of their payments so 38% of their payments run through our consumer network Where the individual has an identity and their payment methods are saved on file and it's literally a one-click check out for them They use us for smart retries so like you know when the transaction usually like a recurring bill fails like we retry it at the optimal time Which allows them to recover about 60% of their failed charges they use us for striped tax Which keeps them compliant in a hundred countries and so it's just a great example of how these AI companies very lean teams growing fast going global and really just like Scaling up to look like a much bigger company was straight behind them I've heard you talk a couple times you know we don't have too much more time So just to hit on a couple last topics I've heard you talk a couple times about the Time you spend getting new clothes for your kids. That's a Mostly something in all honesty my wife does in our home Lucky food yes, I love would flatter myself that I you know do my share in other ways, but she's definitely better suited to pick out You know what will make the kids look cute What are we what's interesting you know but folks who listen to this podcast will sort of know the basics right that like Proplexity has a shopping thing and whatever and you know we know what MCPs are Are there like any recent developments or you know is this really happening or is it still from what you've seen like kind of the Wouldn't be cool if one day this were real sort of phase of Agente commerce It's kind of both right like it's definitely still early There's still a ton we're sorting out about how's this actually gonna work and how quickly is it gonna take off but and we're seeing Meaningful traction you mentioned perplexity rates you can like discover and book hotels You know directly inside the app But it's it's not just like big guys like perplexity like hip-kamp is a little site that uses agents with virtual cards to book Cam sites off platform and from Montana It's like impossible to get into Yellowstone National Park I hate using their website although I love the park and Value that they're not spending a ton investing in intact, but hip-camp is like actually solving that and then you know I think it's easy when people think about a Gentick commerce to think about commerce buying kids clothes consumer, but on the developer side We're seeing the same trend right so developers now and cursor can buy for cell services right inside their editor That's a brand new channel right like this really embedded commerce directly in the workflow and and stripe powers Those transactions too, so you know, we're not totally new to this like it was it was last November when we Launched our agent toolkit, but we still have thousands of downloads each week and I think just looking at the pace of adoption and looking at who's testing I think agent commerce will be a major channel far sooner than than most people think with that embedded stuff like Versaille and cursor. Yeah, it seems like That's much more about like the connective tissue of The user has an intent and it's a question of how it's gonna get executed as opposed to any sort of like Autonomous you know decision making by the agent or any sort of like meaningful delegated discretion to the agent Have you seen anything that is? Really interesting in the like I'm gonna actually trust you to go figure out what to buy and execute on it at this stage Or is that still not really material? Yeah, so I can't name names, but this idea of like a business in a box Like I want to build this business and I actually don't know what third party tools and services I need I just want the business in a box and go spin up the business and that's not just the payment provider or the front end service or the You know bot protection or the you know HR system, but like give me my whole business in a box I think could be an interesting direction now Getting that right for the whole world of businesses that might be created is hard getting that right for You know a pretty focused AI wave that's coming online isn't isn't a crazy thing isn't a crazy thing to reason about so I agree with you that sort of the option set in consumer is broader and so there's more Job to be done for the agent to select from this very broad option set But like SaaS procurement is also very inefficient Maybe we underestimate how inefficient it is and that's not just in the selection of vendors That's also in like the pricing and negotiation with vendors and so I don't think it'll be tomorrow, but I think there will be a there there Cool keep watching out for that last question Just about kind of the the future Of platforms the future of scale the future of kind of market power I Think back often to the week Danthropic deck from like two years ago where the Claim was made by Anthropic we believe that The companies that train the best models in like 2025 2026 May have such an advantage that nobody will be able to catch them from there Why because presumably like the models will help train its successor with all these data filtering and synthetic and Constitutionally I and whatever and like once you've got clawed for Contributing to the training of cod five like anybody who doesn't have clawed for it is still like you know Sourcing every through through scale AI or whatever is just at a massive disadvantage It seems like that basically applies to stripe as well like is there any Is there any hope for anybody to ever compete with stripe given you know the 1.3% of global GDP flowing through the system and the massive data advantage that already exists or Are we now sort of in a future where? We just need to rely on the calcium others to continue to be you know good actors like it seems like this position is like almost unassailable I think financial services is a big broad space and there are a lot of services that one can provide in that space in the context of data The 1.3 trillion a year is a lot right and volumes grow growing like 38% year over year That's like a massive growing data set the real advantage. I think isn't the raw size though to your comment earlier It's more the compounding loop and in our context that loop is the more data we process the better our better models get the better our models get the More value we delivered to businesses Incentives are super aligned the more value we deliver to businesses the more the businesses grow Which means the more transactions they run through stripe and kind of that loop compounds year after year and you know Obviously, that's why we talked earlier about why it's hard to make horizontal bets But that's why we can make horizontal bets not just because we have skill But it's because we're in a position to harness that scale to create even better products Which then is is the feedback loop and we're pushing this further right? You may have heard it sessions last year We announced a big push for modularity and so now products like Radar or fraud prevention product or our billing product or that optimized check out suite for your users are available Multi-processor so they don't just work on stripe transactions They work on transactions or on billing plans or on checkouts that are happening outside of stripe two and that actually gives us window into an even bigger data network and Kind of further reinforces that loop So you know, I think there's a lot to be done in the financial infrastructure structure space and I think there will be plenty of players playing important roles there, but I think we are quite differentiated in the intelligence that we can serve to users and it's just really fun to see how that intelligence in turn helps them grow more profitably. On the other end of that, do you ever think about trying to compete at the foundation model level? This is something that obviously not many companies are really able to do, but given the depth of ML experience and the unique data set that does exist and just the reputation of the company, I sort of expect that if there was a special fundraising round to raise $10 billion to go train a Stripe 1 to try to compete with Cloud 5 and GPT, whatever, the money would be there. I guess, do you ever think about going that hard or how do you think about calibrating just how ambitious to be with the AI investments? Stripe has always leaned into new technology ways. Back when we were founded, it was the platform and marketplaces' wave. I got us a lot of the way here. Today, it's the AI wave and our mission is to build the economic infrastructure for AI. That shows up in today for BigBets being the best partner for AI companies, so just helping them monetize effectively and scale globally and manage billing and manage tax and manage fraud. Two-thirds of the Forbes AI 50 already run on Stripe and we're very focused on co-building, whether it's usage-based building or whatever the next wave is, co-building with them and being the best partner. The second is enabling agent to commerce, so we only talked about it briefly, but agents are going to be buying on your behalf and we want that to work really well for the whole ecosystem. Yes, for the consumer, yes, for the seller and yes, for the platform or commerce facilitator. The third place we're really focused in the world of economic infrastructure for AI is making Stripe native inside the AI-enabled tools that developers already use, whether that's for sale or replete or cursor or mstral's lichat. Payments should show up right where the work is happening. That's thing three and then fourth is what we talked about today, which is deploying our foundation model across the network to improve fraud detection, yes, to boost authorization rates, yes, but also expanding the intelligence layer that we provide to every user. So those are the four big investments. I'm not going to say that there could never be a fifth, but today we're really hyper focused on the economic infrastructure for AI, not being an AI model shop directly. Tasha, cool. This has been excellent. I really appreciate the time and the depth. Anything we didn't touch on that you would want to leave people with or just any concluding thoughts? No, super fun. Thanks so much for having me. Emily Sands, head of data and AI at Stripe. Thank you for being part of the cognitive revolution. Thanks so much. If you're finding value in the show, we'd appreciate it if you'd take a moment to share with friends, post online, write a review on Apple podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries. Either via our website, cognitiverevolution.ai or by DMing me on your favorite social network. The cognitive revolution is part of the Turpentine Network, a network of podcasts where experts talk technology, business, economics, geopolitics, culture, and more, which is now a part of A16Z. We're produced by AI podcasting. If you're looking for podcast production help, for everything from the moment you stock recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And finally, I encourage you to take a moment to check out our new and improved show notes, which were created automatically by notions AI meeting notes. AI meeting notes captures every detail and breaks down complex concepts so no idea gets lost. And because AI meeting notes lives right in notion, everything you capture, whether that's meetings, podcasts, interviews, or conversations, lives exactly where you plan, build, and get things done. No switching, no slowdown. Check out notions AI meeting notes if you want perfect notes that write themselves. And head to the link in our show notes to try notions AI meeting notes free for 30 days.

Podcast Summary

Key Points:

  1. Stripe has developed a proprietary "payments foundation model," a domain-specific AI model trained on its vast transaction data to understand payments as a distinct modality, not as language.
  2. The model excels by analyzing contextual sequences across multiple entities (buyer, card, device, merchant) simultaneously, achieving superhuman performance, such as increasing fraud detection rates from 59% to 97% for card testing.
  3. Stripe's strategy involves exposing the model's representations (embeddings) for engineers to integrate into existing systems, enhancing them without major overhauls, a method less common outside major tech platforms.
  4. This approach creates a powerful competitive flywheel
  5. The episode suggests this "proprietary modality" strategy could be applied to other domains like health and logistics and raises important societal questions about AI consolidating power among incumbent data-rich platforms.

Summary:

" Unlike standard language models, this AI is trained specifically on Stripe's dense, structured transaction data, treating payments as a unique modality. , buyer, card, merchant) to detect patterns invisible to humans or traditional models, dramatically improving fraud detection and other optimizations. Stripe's implementation strategy is notable: instead of rebuilding systems, they provide the model's output embeddings as inputs to existing machine learning systems, making everything work better efficiently.

The conversation highlights how this model creates a formidable competitive advantage, as Stripe's scale generates the necessary data, leading to superior products that further entrench its market leadership. This raises broader implications about AI favoring large incumbents and the potential for similar domain-specific foundation models in other industries. The discussion also touches on Stripe's use of LLMs for data enrichment and ensuring reliability in their AI-powered products.

FAQs

Notebook LM is an AI-first tool designed to help users organize ideas and make connections from complex information. Users can upload documents, and it acts as a personal expert to uncover insights and assist with brainstorming.

Stripe's payments foundation model is a transformer-based AI that converts payment transactions into compact vectors, similar to embeddings. It analyzes sequences of transactions across multiple entities (like buyers, cards, and merchants) to detect patterns, enabling improved fraud detection and other optimizations.

By analyzing extensive contextual data across transactions, the model significantly enhances fraud detection. For example, it increased Stripe's detection rate for card testing fraud from 59% to 97%, reducing costs for the e-commerce ecosystem.

Unlike standard language models, it treats payments as a distinct modality, focusing on structured transaction data and sequences. It leverages contextual information from multiple entities to understand payments in a way that is superhuman compared to human analysis.

Stripe exposes the model's representations as additional inputs to existing classification and ML systems. This enriches signals without requiring major overhauls, allowing engineers to enhance performance across products like fraud prevention and authentication.

AI favors platforms with large-scale data, enabling them to train differentiated models. Stripe's data advantage creates a flywheel, reducing fraud costs for customers and strengthening its market position, making competition challenging for others.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.