Go back

Ep 72 - China is building a market for data. Why isn’t America?

65m 14s

Ep 72 - China is building a market for data. Why isn’t America?

The podcast discusses the rising use of Chinese open-source AI models by Western companies due to their significantly lower cost compared to US models. Host Andrew Polk and tech policy researcher Kendra Schaefer explain that for many low-stakes, high-volume tasks—like processing policy documents—US models can be 2 to 20 times more expensive, making projects unprofitable. Chinese models offer a "good enough" quality at a fraction of the price, enabling innovation and cost-effective prototypes. This cost dynamic is a key factor driving adoption. The conversation then turns to potential US government restrictions on Chinese open-source models, such as federal procurement bans or rules targeting cloud providers that host these models. Schaefer argues that unless the US develops a cheap, competitive domestic alternative, such bans would be counterproductive. They would be difficult to enforce and could harm US global competitiveness by forcing American firms to use expensive models while the rest of the world benefits from cheaper Chinese technology. The discussion highlights the tension between national security concerns and the economic realities of AI adoption, emphasizing that blocking innovation without a viable substitute is a risky strategy.

Transcription

12365 Words, 68177 Characters

English
[MUSIC] Hi everybody and welcome to the latest Tribune China podcast, the proud member of the Seneca Podcast Network. I'm your host, Tribune Co-Founder Andrew Polk, and I am joined today once again by a pod favorite or a pod fan favorite, Tribune's head of tech policy research, Kendra Schaefer. Kendra, how you doing? I'm good, I'm good, how are you? Oh yeah, I can't complain, excited for this discussion, always good to get back in a rhythm with the pod after being offered a couple weeks. So I got to talk to Dene last week, get to talk to you this week. So I'm excited about it. Thanks for coming on. Of course. I am going to talk to Kendra today about some of the research she's been doing, kind of on an ongoing basis for a while now, specifically around how Chinese regulators and Chinese policymakers think about data and how to sort of use data in the economy, how to govern data, all of that stuff. The framework is data as a factor of production. We'll get into what exactly that means. So we're going to do a deep dive on that. It'll be wonky, but super unique research that Kendra has been doing that I'm excited to get into. Before we do that though, we are going to talk a little bit about some of the latest developments in the kind of China tech space around AI, specifically around what's happening with open source models and more Western firms opting to use open source models for cost purposes and potential restrictions coming both from the Chinese and US side on those models. So we'll touch on that briefly before we get into Kendra's research. But before we do that, of course, we have to start with the customary vibe check. So Kendra, how's your vibe today? My vibe is actually really mellow. Nothing catastrophic has happened in the China space in the last 48 hours. I'm pretty excited. I'm going to Taiwan. I think I mentioned the last time I was on the pod. I had an Asia trip coming up and now it is imminent. I'm going in a couple of weeks to Taipei with the Brookings Institution delegation. So I am pumped. Yeah, that's exciting. Good to get back over to the Asia time zone. I know it's been a minute since you've been over there. It's always nice to get back on the ground and hear what people are saying. I know that'll be a great trip. Very cool that Brookings is having you along for that. So excited for you. My vibe, similarly mellow. I feel like we're sort of in the dog days of summer. Yeah. You know, it's like you said, nothing catastrophic has happened. Our clients in a good way seem like they're not having any fires. They need to put out. And so we'd not have people blowing up our email inboxes first thing in the morning. Oh my gosh, we need to figure this out, figure that out. So I'm just kind of leaning into the casual summer vibe. So we'll bring that mellow vibe to the podcast today. I don't know if I can promise that based on what we're going to talk about. Well, I was going to say, Kendra Mellow is sort of calm before the storm by definition. So it actually makes me more nervous. And you're like, oh, mellow. I'm like, oh, something's coming. But now we will channel your energy into the discussion today. So that'd be great. Of course, before we get into the content that we also have to do the quick housekeeping up top. So a quick reminder, we're not just a podcast here. Trivium China is a strategic advisory firm that helps businesses and investors navigate the China policy landscape. That of course includes domestic policy in China, much of which we'll talk about around tech and data factors today. But it also includes policy towards China out of Western capitals like D.C. London, Brussels and others. So if you need a help on that front, on any of those fronts, please reach out to us at [email protected]. We'd love to have a conversation about how we can support your business or your fund. Or if you just have comments on the pod content, reach out to us. We always love to hear feedback from our listeners. I mean, we prefer positive feedback. But we also will take constructive criticism. We'll make fun of you. Yeah, well, in the office. Yeah, behind your back. And then we'll respond. No, we don't do that. We never do that. Secondly, if you're interested in receiving more tripping content, check out our website, triviumchine.com, where we have a bunch of subscription products, both free and paid. They're all sort of focused around Chinese policy intelligence. So we've got a bunch of different options. Kendra's team produces a daily tech policy update. We've got updates on policy impacting markets and impacting sort of the business landscape. So check out the site. You'll definitely find the China policy intel option you need. And then finally, please do tell your friends and colleagues about Trivium, both about the business and about the podcast. It really helps us grow the company. And I say it every week, but we truly, truly, truly appreciate the word of mouth recommendations. They mean a lot to us. And a word of mouth recommendation is so much more powerful than someone finding us randomly through a quote in the newspaper or whatever. So we appreciate folks for spreading the word about Trivium. While you're at it, leave us a rating on your favorite podcast platform. That also helps us grow the visibility of the podcast. So with that out of the way, let's get into it. You ready, Kendra? I'm writing. Let's go. Well, so like I said, I think I want to start just with some of the recent developments in the tech space. The number one game or narrative I've kind of been looking at in this space for the past few weeks is companies increasingly thinking about or questioning the cost of, you know, AI investment of building AI processes into their, you know, internal systems partly because, you know, everyone thought, oh, well, we'll be able to replace humans more cheaply with automated systems in AI. But it turns out that it actually is quite expensive. And a lot of companies are finding out that their investments are actually having lower ROI than investing in humans. I think Alex Carp, the CEO of Palantir, you know, had an interview. I believe it was on TV where he talked about kind of the weak ROI and how it's making companies rethink how they are approaching the issue. And he's, I think, suggestion was that AI companies rethink their enterprise model. I don't know if that'll happen. But that, I think, is also related to this idea and increasing reporting that a bunch of Western tech companies and startups in particular, partly because of this cost issue, are really basing much of their tech build out on the open source AI models because they're either close to the cutting edge or they're good enough and miles cheaper that it makes sense from a cost perspective for them to rely on the Chinese models. So I just wanted to throw that over to you Kendra. What do you think's happening here? How do you see the state of play in terms of these cost inferentials and the dynamic of more and more Western companies taking a look at potentially employing Chinese models to a greater and greater degree? Well, this is an issue, as you know, that's near and dear to my heart because I not only run our tech practice, like our tech analysis practice interview, I also sit over our IT department. And of course, we are working with models internally. Have we talked about, you know, model cost on the pod before? Remind me. I don't think so, actually. Yeah, let's get into it. I mean, this is another one where it's wonky and this is pretty inside baseball, but I think what we are doing is actually quite illustrative of this bigger issue. So yeah, let's talk about it. I mean, I think what we're doing is the issue and it is sort of half the issue. So I think many of our listeners probably will have already used an LLM programmatically. They will have tried to interact with an LLM. They are coders themselves or our vibe coding apps and stuff like that. But there's also a large sub segment of listeners. I think you probably haven't done that and don't really understand what the cost issue is. We have had a sort of intimate experience with understanding where Chinese models are kind of winning the day and where they aren't. So I want to not make that such a squishy conversation but give a very specific example. So for illustration sake, so we use LLMs for processing massive amounts of policy documents. So just for illustration sake, and this isn't exactly what we're doing. But let's just say we need to take a million policy documents and flag, you know, it would take a human countless hours to read all of those and figure out whether or not they're related to a specific sector. Autos, semiconductors, whatever it is. Or if they have a subsidy amount in them and what that subsidy amount is, right? But we can take that giant pile of documents and we can pass it through an LLM and ask it to do that analysis and then maybe sell that output to a client or use that output in our researcher, whatever it is, or create a data product with that output. Processing a bunch of policy documents is a low stakes, low security use case. It doesn't matter if the model is Chinese or just parsing, boring, open source documents. There's no client data going across that channel. There's nothing, you know, even remotely sensitive that is sort of passing across to those queries. And we've tried these processes internally with both US models and with Chinese models. And the bottom line is that the US models are two to ten times more expensive. And I think for one project that we ran some R&D on, it was like 20 times more expensive. That cost differential decides whether or not our product is profitable. Can we even build this? Should we even do this, right? That's a huge difference. It's the difference between it costs us $100,000 a year to run this service or it costs us a million dollars a year to run the service and clients won't pay for it. So it's really kind of that cost is a real maker-break thing. There was one tech CEO, I think that was quoted and we quoted him in the daily a couple of weeks ago. I think it was the CEO. of Lindy, which is like an office productivity platform who's who announced on their blog that they're using Chinese models for some of their features. And he just said, I don't need God to write my emails. I don't need God to write my emails, which is true for so many use cases, right? And so that's not a US China thing. It's just a cost thing. There really isn't a US alternative where the model's pretty good. It's good enough to handle those kinds of things. And then in addition to that, the cost of it is cheap. So there's a thousand reasons that a company would choose cost over quality, R&D. You're just like testing a theory. You're making a prototype. You don't want to use the best equipment. You just want your proof of concept so that you can get to a place where maybe you switch to a US model after that when you want a better quality or you're kind of looking for a top dollar. - Yeah, that actually raises a point that I just want to throw in quickly, which is probably, I mean, is an obvious point to everyone like you who's using these LLLMs and to a lot of companies who are trying to figure this out, but maybe not to some people, which is there's no perfect solution, typically with this kind of thing, you're constantly toggling or adjusting the dials between speed, cost and quality, right? Quality of output. And so at various times, you're optimizing for different ones. Obviously every company wants the highest quality, the fastest speed, at the lowest cost, but sometimes you have to trade off and some of those things and the Chinese models give you a different tradeoff at times. I guess one other question for you, if you can talk about a little bit is, for what we do in terms of kind of looking at Chinese policy documents and other things in that area, are the Chinese models better with working with Chinese language material or is that not right? - Oh, 1000%. I mean, but our use case is so niche, it almost doesn't matter. I mean, how many, you know, maybe our listeners care, definitely the Chinese models are better at Chinese policy documents than the foreign models, than the foreign models. But I think for most people, that's probably not really that big of a consideration. But it's like, I do think that for most companies, unless you are a coding firm, unless you are a bleeding edge tech firm, there is a lot that companies can do with LLMs. I mean, and we've only started to scratch the surface of adoption, right? Corporate adoption. It really hasn't filtered out. And we work with lots of companies who don't use, I at all yet, right? So it's just, there's this huge space where you're gonna have companies who want to use all kinds of models for all kinds of purposes. It's not like we use four different models in our work. And we just use the right tool for the right job. But if the only tool available is the top of the line, most expensive tool off the top shelf, that is very problematic for our economics. - Yeah, we can talk more about this on later pods. I'm sure folks would be interested in how these models inform our work. And, you know, with some of the behind the scenes stuff is, I mean, I think it's interesting. I think people think it's interesting. But today, I don't want to spend too much time because I want to get to the day to factor piece. But before we do that, the additional piece of this is since there's been sort of more reporting about how US companies in particular are using more and more Chinese models, the US government, of course, has taken interest in this issue. And over the past week or so, there's been rumors particularly flying around in X and things like that, in a policy space where people are saying the White House, in particular, the US government is considering trying to restrict access to Chinese open source models. And there was a suggestion that an executive order to some effect on this might be coming out at the White House. And I've done that anyway. I just wanted to get your thoughts on what you think about that as an issue, whether or not the US government should do that, I can guess what your answer is to that. But also how that would kind of work and just provide some context to us about that latest reporting. Well, yeah, I'm sure you can guess how I feel about it. Basically, unless there is a really also not just one good US alternative that is a low cost and good enough alternative, but a robust ecosystem of competitive US alternatives, it is a real bad idea to restrict access to the models that allow innovation to happen in small businesses, in the laboratory, all of those kinds of things. There are other reasons besides pause to choose an open source model, right? That includes being able to download it and install it on your own machine at home or more likely in your own private corporate data center, which you can't really do with US models, right? So the US just simply doesn't have a great alternative and I think you said something to me earlier which really rang true, which was like, if the US decides to try to ban access to Chinese models, and I'll talk about how I think they might be able to do that in a second, but if they go that route, I mean, it's basically the same route as saying, hey, we can't manufacture a good NED either. China's got cheaper, better NEDs now, but we're just not going to allow them into the market. Did you see the, I think the CEO of Ford a couple of days ago, you know, it was like one of the New York Times headline essentially said, look, we support the US in blocking Chinese cars from coming into the market for now, but you absolutely aren't going to be able to keep them out forever and we have to be able in the long term to compete on a playing field with Chinese manufacturers. And it's the same thing here. It's like, okay, well, you can ring fence the United States for a little while and let everybody else use. Cheaper open weight models, but the economics get real wonky. The longer you hold that line, if we don't have a good alternative and we simply cannot be competitive. So I think that has to be addressed. If they want to do a ban, all right. But man, we better have a good alternative in a plan for how we're going to offer cheap processing to domestic companies or I think it's stupid. - Yeah, well, and I hate the reaction being the knee jerk reaction to being, we want to win XYZ part of the tech race. And so we are just going to keep China out, right? Of our market. Like, it just strikes me as such sort of simplistic thinking. Like, we want to win. So we'll just kind of block them. And like the way you win is be the most competitive. - Right. - Right. It's also what I mean. You don't tie your punnets shoes together that only gets you so far. You know, you might win a couple rounds doing that, but I just don't think over the long term, that's not a sustainable strategy. We can't just keep saying, well, okay, then you just can't sell. Oh, well, you mean, you don't have a better when you can't sell the air. I mean, it just isn't. - Well, and that doesn't even account for, you know, what does that do for the global landscape? Like, you know, you do reduce your competitiveness globally. And now, oh great, well, all US companies run on really expensive US models while the rest of the world works on, you know, just as good or nearly as good, very cheap Chinese models. Like, that's not a positive outcome. One quick thing before we finally pivot is you also, I said we weren't gonna give them this too much, but you're unclear exactly whether or not the US government can keep open source models out of, you know, how do you even enact a ban like that? - So I think from what I understand, there's a couple of options under discussion. The first one and the most obvious one, although this has already been done to some extent, I think, is federal procurement bans, basically, right? And which is what they did with TikTok, is the very first step the federal government took was that you can't put this on a government device, which is just that's very low hanging fruit, but they could also say any government supplier can't put it on, you know, can't use it either or you can't be a government supplier. So there's those kind of that could extend in that way, or you cannot use this tool on a government contract, basically, so they could go that route. I think the main concern is that the commerce department is gonna use the ICTS, like, sort of supply chain restrictions toolkit that they've thought, basically, the USG has a rule that essentially says if a tech product or service comes from a foreign adversary and could be used to spy on Americans or sort of threatened US national security in some way that then commerce can kind of ban it from the US market or force changes to how that is used. The problem is that this rule regulates transactions. So it's kind of awkward to try to characterize downloading open source models as a transaction. So the question is which touch point would they go for? They could maybe go to cloud companies and say no US cloud provider can host these models, which is mostly how people are using that. It's a large, not everything, but it's a large chunk of how US companies are using those models they're going through Amazon. So you could do it that way. They could try to go to, like, hugging face, which is where models are listed where a lot of these open weight models are listed and try to ban them from listing it in some fashion, which would make it difficult to download. People wouldn't know where to go to get it or it would be, I'm sure, in two minutes, somebody would put up another mask. Just like, host it in Sweden and in it. So this is difficult to enforce. - Our colleagues didn't think my joke was funny, but obviously it's just going to be on the dark web, which is where I'm most proficient. - Sure, you hang out. - Yeah, exactly. - Yeah, yeah, exactly. (laughing) - I mean, okay, and so there's that, and they could also, I think, use the, what is it, the emergency powers, IE EPA, or I think it could kind of declare an emergency and go for it that way. So there are things that they could essentially do. I very much hope that policymakers are weighing what it would mean for US firms to not have access to that kind of technology. And what I would love to see is if the US government focuses on how to incentivize the development and release of a cheap open source US model. All this goes away. I don't care if I'm using a Chinese model, we're perfectly honest. I'll deal with like a slightly crappier, if I don't have to deal with any US government problems, I don't care if I'm using a Chinese model or US model, I care if it's cheap and good enough. That's all I care about, right, as a developer. So why don't we just focus on figuring out some policy incentives to make sure we have one of those? I don't understand why that's not the primary topic of discussion, or maybe I'm just not in those rooms and maybe it is, but anyway, yeah, that's my thinking on that. - Yeah, well, we'll leave that in the to-do list, figuring out a policy agenda to advance US open source or US developed open source models. I'm sure some one somewhere is having that conversation, but that's all super helpful, very interesting stuff. Well, of course, stay on top of all of that as it develops because it'll be an important part of not only what we do, but very important for our clients as well. I want to pivot now to your research. We're going to get into your work on data factors or data as a factor of production. And this is kind of evolving thinking in the Chinese side around how the government treats data, how everything from taxing data to data ownership, all that stuff. So you've been doing this research for a long time. How long have you been doing this for like six years now? - Yeah, I started in 2020, it's been six years. I have been cornering people at parties about this and torturing them for six entire years. - Well, that sounds like a fun party. Remind me not to go into your parties. But so the topic overall is what, how China thinks about data, is that not something that sort of we already know the answer to? I mean, it seems like it should be relatively straightforward, but maybe I'm wrong. - Yeah, no, you're right. I mean, I think that's a perfect place to start because if you ask anybody in DC, what's the big US-China data issue? Or how does China think about data? You'll probably get something to the effect of, China's primary goal is to steal sensitive data from American citizens or the United States, and the US has to prevent that from happening, right? That is the vast majority of the DC conversation on US-China data. - Yeah, that's a little mind numbing for sure. I have had that conversations many times in Washington, but what's the conversation more if you talk to people about this outside of the DC, bubble hour people thinking about this that aren't so focused on national security and policy, and that kind of thing? - I mean, I think the other group of people that we talk to about this is foreign companies that operate in China. They're not obviously as worried about data exfiltration, but they'll kind of tell you the biggest issue is cross-border data flow, right? China's got one of the strictest cross-border data regimes in the world and for the last five years, I think multi-nationals kind of been tearing their hair out, trying to get their own information out of China, and so that's basically what corporates are talking about. So DC's talking about China's trying to steal our data, corporates are talking about, how do we get our data out of China, and how do we comply with Chinese data laws without screwing up our R&D processes and stuff like that? But as far as I'm concerned, both of those views, or both those conversations really only look at a teeny, teeny tiny corner of the conversation that is happening inside of China about data, right? In China, the government has been having a very broad conversation. They've essentially developed a sort of part theory, part national strategy about what role data plays in the economy, how to activate the economic power of data, right? How to use data to boost GDP and make gains, and how to kind of bolster technological competitiveness by increasing the supply of data. So we saw this start kind of six years ago, and then we've just been watching that theory evolve over time, and it's now driving this huge wave of Chinese tech policy. And I think that way, just sort of flying under the radar of it in the US, you don't often hear people talk about how the Chinese government thinks about data. Why do you think it is so under the radar? I mean, if this is like the fundamental thrust behind the conversation in China, why isn't it on more people's agenda here? Well, that's a good question. I mean, I think two reasons. One, all of the data policies we're going to talk about today, individually, if you look at them by themselves, they're just deeply unsexy. It seems very uninteresting. They're really interesting in aggregate, but they're very uninteresting by themselves. And so unless you can see what they mean in aggregate, looking at one particular piece of it, isn't that fun? And then two, I think the way that China's looking at this is so different. I mean, deeply different from how the US talks about data that it kind of doesn't even register. It doesn't pattern match to anything in the US policy conversations. We don't see it. Well, that, I mean, I think is exactly why you and I wanted to have this conversation, right, is to start highlighting this. But why do you, in particular, think it's so important at this moment that we, yes, the US policy community start to see it for what it is now? I mean, I think they answered pretty easy, right? Data supply is now a core input to AI development. The AI competition that everyone's obsessed with is in part a data competition. So what we're going to talk about today is a very heady idea, right? How do Chinese state views data? What is the long-term strategy? What's the big idea underneath these evil policies and what that means for the US? But that's also now very intimate-- like five years ago when we started looking into this, that was a very squishy concept. But now it has this immediate economic impact because of how important it is, because of how critical and central is data is to artificial intelligence. Yeah, good point. All right, well, let's get into some of the details here where you want to start in terms of diving in. OK, awesome. So this is me cornering you at a party now. Oh, no. [LAUGHTER] Look at the time. [LAUGHTER] All right, so this is kind of going to sound like a bait and switch. But I want to start this conversation with a concept that doesn't seem to have anything to do with data at all, because getting into how China sees data sort of hinges on understanding the sort of Econ 101 concept, which is what is a factor of production. And I think a lot of our listeners probably remember this from school. But I don't know, you're an economist. Do you want to give us the 30-second refresher? Remind everyone what is a factor of production? Yeah, I mean, I think I can do it in less than 30 seconds. I mean, traditionally, the factors of production are land, labor, and capital, right? So think about the agricultural economy. You'd need land, labor, of course, humans, the people who do the work, and then capital being both money and equipment. So equipment, of course, matters in agriculture, but also in manufacturing. So basically, the fundamental inputs that you need to produce economic activity is what we think of as factors of production. Right. So a factor of production is the input necessary for businesses or whoever to create economic value. And if they don't have those things, they cannot create output. And there are typically, I think, in traditional economics, there's four. You said land, labor, capital. And then China calls the fourth one technology. I think the US calls it entrepreneurship, but basically, like, IP, know how, you know, like-- Yeah, well, I would call it sort of productivity. Doesn't matter. I won't be potentially on that. But it's really how those things interplay. Like, basically, productivity is how well humans use capital and land. Right. That's a little bit-- Anyway, yeah. Right. So if you're going to do business, you need some work to operate. You need people to do the work. You need money to fund it. You need to know how to put it all together. So that idea of those are the inputs to the creation of economic value. That idea has essentially been stable for about, you know, a century, right? It's the sort of periodic table of economics and nobody messes with it. Yes. And I feel like there's a butt coming here in the China context. But in 2020, China did actually mess with that idea. So this is a kind of interesting part. So in 2020, the State Council released this high-level macroeconomic policy. And buried in that policy was something quite remarkable, right? The policy basically designated data as the fifth factor of production. So now, according to the sort of canon of socialist economic theory that China runs on-- and remember, that's like the foundational theory that the entire state apparatus uses to make policy, right? We've decided that this is the sort of economic theory. And based on the theory, we're going to make some rules and we're going to make some policies and incentives. There are five factors of production-- land, labor, capital, technology, or whatever-- and data. And what's the point of adding data? I think it's somewhat obvious based on what we have talked about so far, like pretty obvious input. What do you think the point is of China to elevate data to that level in the canon, so to speak? Well, I think by doing that, what the state is formally saying is in a digitized economy, companies need data to produce economic value, right? As you said, in the agricultural economy, let's say 300 years ago, if you wanted to create value, you need a plot of land, and you need to do to farm that land. So you need land and labor. But in the digital economy, I did. But now, in the digital economy and the modern age, you need data as an input. Or your company needs data as an input, in the same way that they need financing. And so that sounds abstract, but it actually has these enormous practical implications. Because think about what that means. It means the state is taking responsibility. If the state names something a factor of production, they're basically saying the state is responsible for making sure that companies can get this thing. Yeah, that does make sense. And I mean, in a way with Agriculture being such an important part of kind of how Chinese policymakers think of the economy They would never actually drop land as a factor of production But you can see for most modern economies land is sort of less and less and important one So it's almost like you could add data and take away land like for our business We don't need land, but we do need data. Yeah, but that's just a quick point But more like what do you mean like in terms of people or companies getting data? What do you mean by getting it? So okay, so it's the state's job to create a market environment where businesses can access the inputs they need to grow and contribute to GDP right so if companies need labor That's fine. It's on the state then to kind of build an education system that produces the right workers or to Right employment laws that like balance the needs of employers and employees so that talent can Flow smoothly between firms and hiring and firing can happen well balancing everybody's needs So it's kind of on the state to create the background The environment in which labor can get to companies where they can acquire it and use it well And then if companies need land to build a factory it's kind of the same thing right it's on the state to run Zoning to run deeds and titles to write property and ownership laws Those are things that we take completely for granted. It's like Invisible infrastructure of the market we never even think of it But those systems are basically what keeps factors of production Moving throughout the economy and keeps them flowing into companies Enterprises. Yeah, that makes sense. So you're saying basically that this same logic at least in the Chinese context now applies to the data The state is taking a role in making sure there's an ecosystem that sort of curates and feeds data Into companies broadly speaking is that right to have it right or is it different that yeah? Yeah, exactly exactly the state saying look there's already a capital market There's already markets for land and natural resources. There's a labor market and now It's on us to build a data market the systems the regulations the standards that basically govern how Data gets bought and sold and traded so that it can sort of circulate through the economy and so that businesses can get our hands on it And in order to describe that idea the state has Basically formulated or coined this term data factors meaning data when we view it as a factor of production data as an economic input Okay, yeah, that makes sense. I guess the question then for me is when you talk about quote-unquote building a data market You know strikes me the data gets bought and sold all the time without the intervention of the state right and their data brokers their entire Industries already existing around this both in China and elsewhere. So what does the state need to build anything? I mean like what specifically doesn't need to build I mean actually that's such a great question because I think there actually is a big open question about whether or not The state needs to do anything or needs to take an intervention as to approach to this at all But I think if you ask Beijing what the issue was they'd say that For every other factor of production humans have been trading it for in some cases hundreds of years Right we've been trading a land for hundreds of years and so the rules of the road are kind of ancient I mean we solved the fundamental plumbing problems that make those markets run to the point We don't even see them anymore, but none of that plumbing is there for data. Okay, so that all sounds squishy We've been very squishy. Let me get very concrete. Let's do a concrete example So imagine that you wake up today and you decide I want to buy an acre of forest land in Washington state So what is the first thing you do after you've decided to do this? Well either Google or ask an LLM or a chat GPT. Where do I buy land or why she's doing? I mean now I like I guess you like sign on to some third-party site like a Zillow for land Great you would know exactly what to do you want to buy real estate? You open a real estate website. There's a real estate market at your fingertips You would open one of a dozen well-known sites all of which are kind of pulling from the centralized property listing systems That have been there forever and you just browse what's available consumers know where to shop no bigs Now imagine you want to go by access to regularly updated shipping container movement data now. What do you do? I don't know. I mean truly that's where I'd start but there's not like container data dot com like container data dot com is not like a common marketplace where all data sales I do website idea. Oh, there we go What we're doing right now. So there's like their real estate. They're well-worn pathways for discoverable real estate and not so much for Other kinds of data right you've like you poke around online you'd Google it But there's no a business camp wake up and say I need this very specific kind of data And I know where to acquire it in most cases. Does the supply of data you want even exists? Who has it? Right and so the reason data brokers exist is because you go hire these people to find the data for you because there is no place That you can just simply go find it yourself in most cases right so that's one problem discoverability How do I discover the supply? Where is it? How do I get it does it even exist? Problem number two, okay, you're back on Zillow. You're buying your acre of forest. How do you figure out what you should expect to pay for that data? Compare well see what's out there right look at what's on the market and compare them to like I guess decide the parameters of what you want and compare them to other Comparable acres of land houses, etc. Whatever you're trying to buy there. Yes exactly you look at comps Are you look at house with the same you know the rally real estate? Lee a house with the same number of bedrooms and bathrooms that you're looking for in the same street and you'll say oh With the same square footage and you'll say I usually sells at this particular price you found your million dollar parcel Right whatever and then the value whether or not that value is correct Basically gets confirmed through an appraisal in the process of buying your property And it's the same with the labor market if you want to hire a senior engineer with 10 years of experience You check indeed or a zip recruiter or glass door you see what everyone else is paying for the same Set up skills and of course capital markets have decades of you sort of established valuation methodologies, right? So you can find the price for similar items easily Whether you're buying or selling now if you're buying or selling that shipping container data What should you expect to pay for that? How would you know that you're paying fair market value if somebody does put you a cost and If you were selling data, how do you even know what it's worth or what you should be charging for it at all? Yeah, I mean, I guess no real answer. I don't really know but I mean fundamentally I guess it's worth whatever someone's willing to pay for it There's no real standard metric for valuation this type of data is valued at this amount of money in general Right, it's very hard to do that. A) there's so many different types of data But B) we just haven't been selling it that long and it's hard to compare one data transaction to another data transaction right now Right and so I mean, I think this is very interesting But the inability to put a very clear standardized value on data actually creates the sort of cascading set of downstream problems and here's here's my favorite one Let's say your small tech startup. You don't really own that much physically You don't have any equipment you know real estate you know Tractors or anything but you're sitting on a genuinely valuable data set or you've collected or made You know some data that is worth a lot you think it's worth a lot That data is your most valuable asset now you go ask a bank for a loan and of course they want like collateral or something Right, right. They want collateral. You don't have physical assets physical assets work in collateral As in part because they've got a third value the bank knows it can resell You know your equipment for a million dollars a few default But if it takes your data, which is your only asset as collateral What are they going to recoup on that? We're really even going to put it how would they offer it to they can't price it They don't know what it's worth and so that creates the situation where data rich companies that don't have a lot of assets Which is to say like a lot of tech startups are it become at a sort of structural advantage when they're looking for for financing They can't use this valuable thing that they have Yeah, I guess I had not thought about it from that aspect in terms of Becoming a structural challenge for capital allocation I mean, I think maybe the US and the West broadly maybe a little bit better of that Through venture capital, but that's like basically your gambling's wrong But you're taking big bets on something you have no idea about and China has obviously a venture capital ecosystem But there's a long-term problem that small companies innovative companies can't get capital So this makes sense that it would feed into this issue of you know sort of lending issues capital allocation issues Yeah exactly. I'm going to give one more example just to kind of give a little bit more meat on the bones so Let's say you bought your land you have purchased it and now you go to closing and it's time to take ownership of that land Right there's a mechanism for doing that that is very well warned the deed gets transferred into your name And that transaction and whose name is on the deed gets registered with some kind of county reporters office So that forever after if anybody needs to verify who owns that land right now they can check the registry There's nothing like that for data. We don't really even conceive of data as something you would need to register in that way Right that you would need to kind of confirm that you have the right to buy and sell and the right to own and the right to use That there would need to be some kind of allocation Beijing does think that that is probably necessary so You can kind of see these four issues pulling back a little bit like all of these things are related to trade these kind of invisible pieces of it discoverability, valuation, can you figure out how much it's worth? collateralization, can you turn something into an asset that can be used as collateral and registering or confirming ownership or rights to ownership over some kind of property? Those are four of the many unglamorous, invisible, plumbing problems that have basically been solved for every other factor of production and just don't exist at all for data. So you're saying that basically establishing those four things for data is the underlying project that the Chinese state or policy apparatus is trying to achieve here. Do I have that right? Yeah, that's the whole project. I mean, not just those four, there's probably about 20 different unglamorous plumbing problems like that that the state has identified and gone, okay, we're going to have to launch a sort of policy initiative to do that. Yeah, I mean, when Chinese policy makers say data is a factor of production, what they're really committing to is just what we said, define the fundamental roles and processes and systems surrounding transactions so the market can grow. And the theory of the case is if we make data easy to find, if we make pricing standard and predictable, if we let companies illegally sort of establish and protect their rights to data so that they can trade it, then more companies will want to sell data. More companies will buy data. That means more companies will acquire and use data, empowering the data, economy and share data and trade data. And so supply goes up and circulation goes up. That's generally that's the fundamental data theory. Okay. Yeah. Makes sense. All right. So thanks for laying that out. I think that kind of sets the sort of theoretical and sort of contextual piece of this. But let's kind of go a layer down. What can you talk about like an actual policy here? Something sort of more concrete that solves one of these problems. The Chinese policy apparatus is putting forth. Yeah. I'll actually I'll give you three. I'll talk a little bit about how the state is actually trying to solve those three problems like this couple of the problems we just talked about. So first, the registration and ownership problem, right? How do you confirm you have the right to sort of use a specific data set and a specific way? What we're seeing now is that the NDRC, right? China's big sort of macroeconomic agency is piloting what they're calling a data property registration system. So you can think of that. Well, the way they've described it is a land registry or like it, like a patent office, a securities depository, but for data where data owners and users can register their claims, log rights to use and then trace the history of ownership of specific types of data. So in other words, I could basically say I made this data set. I'm putting it on this registry. I think they're talking about the underlayer, maybe being built on blockchain or something like that. But I thought this registered that I'm the owner of this data set. And then let's say I'm transferring, it's not really actually with data about transferring ownership. It's I'm going to allow you to use my data set for the following purposes and the right to use the data in that way is then logged in this registry. And the use the end user can then take that data and use it without worrying that there's going to be some kind of, you know, there's like a clear transaction that they can point to and a clear rights document that they can point to that is sort of part of a sort of central depository. So that's the general idea with that. And they're already kind of trialing that at the local level. Shen Zhen in particular is actually running a trial that's supposed to go national in a couple of months. And last year we actually saw the NDRC's national data administration put out this call for research proposals on how to construct basically asking researchers for ideas on how the base construction of that system should be run nationally. So we see a lot of movement, right, early movement on constructing a system like that, meaning that companies in China in five years, three years, that acquire data that sell data, that use data that leverage data in any way, and we'll probably have to transact with this registry. So this is like the county recorder office registration system, but for data sets, you're saying. Yeah, exactly. That's exactly right. So the second issue, right, discovery ability in the where do I even shop problem? This one's pretty simple. We've been watching this for many years. There's been this sort of wave after wave of state backed data trading platforms established. They call them data exchanges. Usually it's a local government that stands one up. It's basically a platform where you can browse available data sets. Most of the companies listing data sets on their state owned companies, indicating that the private market is not really that interesting. It's acting on these state exchanges. So I don't know that they're the best idea, you know, but the state has been essentially doing that. There's one in Shanghai. There's one in Beijing. There's one in Xinjiang. There's one in the way on still where essentially just a centralized marketplace where people can go and kind of shop for the data that they need. Or at least that's the fundamental idea. Okay. I mean, got that. But I guess a follow up question would be what are they doing that sort of resource allocation issue that we talked about before or how to get a bank loan based on your data sets? How are they looking to solve that issue? Oh, well, this one's actually my favorite because it's really concrete. State banks are running pilots that let companies use their data as loan collateral. And so we've studied quite closely the structure of those pilots because I think they're pretty interesting. It's a three party structure. So you have the bank that's making the loan. You have a data heavy and asset light company that wants alone. And then the third party is back usually one of those state back data trading institutions, so like a data exchange that independently certifies the value of the company's data assets like in a praser, their data praser. And so then the bank makes the loan sets the loan rates based on the value of the company's data assets. So there's like one example, I think from August 2024, when the Chongqing branch of Fwasia bank partnered with this data trading platform locally and offered a 1.3 million women be so not a big loan to a company in Chongqing that was doing smart city development. Right. And so then the trading institutions certified the data's value, the bank price, the loans interest rate off the certified value. And that's how the money was issued. So these aren't big numbers. 1.3 million women be as not like a massive loan or anything like that. But it's interesting just to watch them kind of see it is this work. Who we proceed here. Yeah. I mean, that's actually that that whole system depends on the bank or someone else, some third party, whatever it is, being able to credibly say what the data is actually worth, right? Yeah. Yeah. Exactly. So there's another piece, right? Another piece of unglamorous policy plumbing. So the Ministry of Finance has basically been supporting research into standardized data valuation methods. And we saw a couple of years ago in like 2023. There's this body called the China appraisal society, which is like an industry association tied to the Ministry of Finance. They usually just do the physical asset appraisals. And so now they've been publishing guidance on conducting data asset appraisals. Right. And so they asked people to look at basically creating a sort of framework for determining for how an appraisal should be able to instead of price on data. It's very interesting stuff. Okay. Let me step back for a sec. So that all makes sense in terms of domestic flow of data, right? Kind of trying to boost the infrastructure behind the pricing of data, how data can be used as collateral, where and how you can sell it, where you're in how you can exercise the rights to data. But I mean, as we talked about before, foreign companies who we work with, non Chinese companies are primarily interested in cross-border data, right? Getting their data in particular out of China. And China's regulatory regime on that front is incredibly strict. So we've worked with these companies trying to get their data out of China for months and months and months. So it strikes me as actually quite normal for China. But talk to us about that dichotomy where yeah, we want stuff flowing freely internally, but we don't want it to go across the border. What's going on with that? Well, so I'm glad you brought that up, right? Because actually, I think this is the single biggest miscalculation in how DC reads China's data security regime. I mean, the DC read is China's data security rules are digital protectionism. And that's it. Right. China wants to build a wall to hoard data inside of China's borders while they steal data from everybody else's that's kind of the standard rate, that's kind of the standard framework. But from Beijing's perspective, the data security regime isn't a wall around the market. It's actually the guardrails that make the market possible. Like, it's not unusual for markets to have guardrails, even really, really have a hand to guardrails for cross-border trade, right? Capital markets have a zillion guardrails for cross-border trade. Labor markets have a zillion guardrails and not necessarily cross-border, but there's some, right? And so the logic runs, if the state clearly establishes what kind of trading is not allowed and where the safety risks are, and a lot of those risks are bigger in cross-border trade. And if it clearly defines which categories of data cannot be traded and starts there, then everything outside of those lines can sort of flow more freely and with more confidence, right? I urge listeners, anybody who cares enough to, after you finish this episode, go read the actual text of China's data security law. Go read it. I think Digit China has a really good English translation. And I promise you, it will read differently than you remember if you've read it before, right? There's all this language in there about how data security is the fundamental building block of data trade and that security has to be strong before data trading can occur and that how all these data security roles are about, you know, enabling the safe trade of data and the state's job is to enable the safe-traded data. And I think we just kind of gloss over that because we don't, again, it's not really on our radar that this is the plan, right? I do actually want to say one other thing though. So that's the plan, but China's data security regime is still overcalibrated. I think they do want trade, but they have significantly overshot on the, let's secure this before we allow trade to the point where the current regime is not serving its own goals. Right? I think the one you undesired to enable safe data flows, but the state is kind of its own more stand-in with this like knee-jerk oversecuritization. And so what we're watching right now is the state kind of actively hunt for a balance point. How do we balance development and security? We heard that a thousand times, right? And we've watched the pendulum swing really hard towards security. We've watched it swing back a couple of times. It's just, it's a live negotiation. Yeah, I mean, that's not shocking, right? Like the security versus development debate, to the extent that it's even a debate or finding that balance is an ongoing endeavor among Chinese policymakers like in a range of areas, right? Data, technology, supply chains, you name it. They're always trying to strike that balance. So that's not shocking to me. And it's also not shocking to me that they've leaned a little bit further into the security side than the development side, which also is like normal for governments everywhere, but also in particular for China. But I think you've done a really good job here of laying out kind of the main rationale that China is using to put forth this data governance regime, some of the specifics around the very concrete plumbing and flowing issues or flow issues that Beijing was trying to solve. But flip that around, what do we, as people who are in the policy community in the US, to make of that? What should Western policymakers or policy thinkers take away from this discussion? Yeah, I mean, I definitely don't think that the United States needs to adopt the idea that data is a factor of production and rush into Beijing's footsteps and do exactly as they have been doing. That's definitely not the point. I think the biggest takeaway is that like when you lay China's approach to data policy next to America's approach to data policy, on our side there's this kind of massive gaping hole where a proactive pro-growth US style data strategy ought to be. Right? China has a pro-growth strategy so we need a pro-growth strategy. Every major US ally has already done this. We are the outlayer, right? The UK, Japan, the EU, Canada, Australia, all of them have looked at this issue. How can we use data to foster growth? What are the problems we need to solve? What are the pathways we need to take? What are the incentives we need to put in place? And we simply have not done that. And I think it's because, you know, as I mentioned earlier, when the US talks about data, it's almost exclusively as a security issue. And when security is all we talk about, then security is all we do. I mean, just look at the last five years. We've done a ton on security. We have secured telecom equipment, smart car software, port cranes, cellular modules. There was the TikTok fiasco. We went after WeChat. We're doing ICVs, you know, preventing Chinese cars from coming into the US because they collect data on this. So all of those actions was fundamentally about preventing the exfiltration of sensitive American data. And that's just the entire American policy portfolio right now. Yeah. Okay. So I understand that. I guess the question then to me, actually, I was thinking of this as you were talking through the Chinese side, is it that the US has just decided like we don't need a growth strategy for data per se, or is it that like does the government need to be involved to the extent that China is involving itself here, meaning like our US policy makers just saying like the market will figure this out or which I think he would be, you know, that might be also an appropriate way to go. I don't know. What is that conversation happening in the States? Is it just like a growth strategy is nice to have? Should it be left to the market and where do you land on all of that? I think every time I have heard policy makers talk about this kind of sort of pro-growth strategy in the US, it has been talked about like, like those are the Montessori kids. Like that is a quen by-ah, get out the guitars and sing together. Let's all talk about data sharing. Let's all talk about like almost it's taken on this. Like hard left kind of, I don't know, let's all hold hands and share data kind of initiative. Right. It's just got this very strange overlay in the US that I haven't really seen it take on anywhere else. I'm exaggerating there are certain initiatives that have, you know, made some progress. But there's been a lot of that. I mean, there was some government data sharing initiatives where the US decided to try to push more government agencies as another thing China is doing to release more of their data in a format that researchers could use to the general public. And that was treated as this like, you know, there's some open data laws about what, you know, researchers supposed to do, get government departments to share more data with each other so that they could be more, you know, efficient and improved bureaucratic efficiency, all this kind of stuff. But they don't have any staying power. They die. They go to the back burner. They get treated as not important. I think because my personal take on that is that in order to see the value of initiatives like this, you're looking at a 20 year investment. You're looking at a 20 year investment in research. You're looking at a 20 year investment in changing the way that the bureaucracy functions. You're looking at a 20 year investment before you see any returns. And on a four year or an eight year administration, we're not good at that. We're not good at making, I mean, that's one of the US's weak points, unfortunately. We're just not great at making investments that we hope will, you know, prioritizing investments that we're going to reap the dividends in two decades. We're great at, you know, let's reap the dividends next year. But we're just not really good at those kind of long term goals. And so I think that's why, you know, security strategy, you can implement that within the span of a single administration, you can ban TikTok in two years. Or I guess not. I guess you can't. You can try to ban TikTok in two years. That's a three administration. That's a three. That was a bad idea. You can institute semiconductor export controls or whatever. You know, you can put out an executive order in a minute. You like generating more efficiency and growth economically from data is like a little bit of a squishy idea. And that's a little bit of it's too long term. And actually, I just want to say, this is real money. This isn't just a sort of wishy washy. Oh, gross. But like, there are actually numbers there, right? The OECD kind of concluded back in 2019 that data access and sharing, if you increase the supply of data in the economy, that it can generate benefits worth 1.5% of GDP if you're just talking about public sector data. In other words, if you just make governments release more data, then you can really generate a bunch of significant economic benefit out of that because companies will jump on that data and they'll make new businesses out of it. There's more data available. Let's make an app that like uses that data to do something. You know, and then that creates jobs and then that creates productivity. And then if you also account for private sector data, if you basically get companies moving their data around between market actors more than they do instead of sitting on it, importing it or being afraid to share it or can't be bothered to sell it or whatever it is, then you know, the range gets a lot bigger. You can get a bump of like between 1% and 4% of GDP. So it's like really leaving, actually leaving potential gains on the table in a way that's pretty detrimental, I think. Yeah. Can you talk actually just a little bit more about the channels through which you see and again Chinese policymakers or others, non-Chinese policymakers see like what avenues are there for data to be a growth driver, generally speaking, I think that'd be interesting for the service as well. Yeah, I'll give a couple more examples. So I just kind of mentioned one of them, which is job creation, right? I mean, I think some of the studies that are coming out now are basically showing, as I just said, data is available to startups, to innovators, to entrepreneurs. They come up with cool ideas for creating businesses with the data. If the data is not available, then they don't do that, right? And that's especially interesting because you have a lot of situations where the government or a large company or a collective of companies is the only body capable of putting that data together. We have, I thought one example, I cite a lot, which is so in 2017, Deloitte did a cool study. They looked at what happened when transport for London released real time transit data through APIs. And I think it was free. If I recall correctly, I don't remember exactly, but I don't think they charged for it. Like transport for London, put this out. 600 apps got built off the back of that data. 500 jobs were produced. And the economic savings for the city were like 130 million pounds. That was one data set. This is one data set on the market, right? And if you aggregate that across the entire economy, what you could do with that is like pretty cool. There's some early research indicating that you will get a small productivity boost, and firms invest in collecting and using their own data. So if you basically encourage a company to go acquire data and then transform that data and make, I mean, we're seeing that in our company right now. We're using data more than we did before. And there's a lot more, like we're doing bigger things faster, right? We can see it kind of in the way that we're working at the moment. So the productivity is a way that you can kind of get growth out of that. And then third, it's like the government itself kind of gets better. The bureaucracy gets more responsive. People get better public services, right? And China's a really good example here too. Nobody really liked how China responded to COVID, but they responded really fast. And that epidemic control was all totally data-driven, built on 20 years of investment in data sets for public health, for transportation, that they just leveraged the minute this disease kind of appeared. They took all these existing data sets and they pulled them and started drying insights on disease spread. And that's kind of how they did the entire epidemic control measures. And they did that in just a couple of weeks because they made that investment already, right? And finally now, it's of course talking about this bit, but it's AI. The biggest human AI is like AI researchers and small AI startups, like specialized AI startups and niche industries really need a steady supply of this high quality data, especially data that's hard to get. So that would be things like, imagine what you could do if you had an entire data set of all of the mechanical equipment failures and smart factories across manufacturers. Not just one manufacturer's data, but every manufacturer's data. Would you improve uptime productivity, production speed of machinery? What insights could you gain from that? So tons of things like that across an almost every sector. And so healthcare, another great example, hard to get good healthcare data because of various privacy restrictions, et cetera, but you get tons of benefit from that. You can cure diseases with that kind of stuff. And so China's made that producing that supply. This is where we come back to factorist of production. If data is a factor of production, then making sure that supply exists so that these things can happen. It's a state shop now. It's a state's priority. They've taken on that responsibility. They've decided to move that ball forward. Right? So anyway, that's the game. China is obviously pursuing that. You would say that US policymakers are just kind of leaving that on the table in terms of not having a national strategy for data development and supply. There was a couple of mentions of data in the Trump administration's America's AI Action Plan. I read those I got real excited about. Some of those are really good. They're actually really good ideas. And they have not at all been prioritized as much as all of the securitization stuff in that plan. The funding has not gone to those initiatives yet. Tick tock, tick tock. It's that kind of stuff. It's like somebody will recognize that, yes, mostly those initiatives were about funding, consortiums, that pool sort of high quality data and compute for leading edge researchers. It was solving that access to research data problem for AI specials and stuff. So it's not that somebody hasn't written it down. It's not that somebody hasn't said, hey, we ought to do this. It's that when you look at where policy maker time and energy and attention is going, that's not what anyone's talking about. When you walk into a room where they're talking about AI and DC, nobody's sitting around saying, how can we really squeeze economic value out of data? How proactive, positive, long-term roads can we lay down so that we really get benefit from data? That's not the conversation that's happening. It's not that it's not recognized. It's just not prioritized. Yeah. Well, and again, I just sort of anticipate listeners saying, well, that's not the state's job. I guess my thought would be, of course, the US is never going to take the same state-heavy intervention approach that China is. But that doesn't mean there's no role for the government to help build this ecosystem. I mean, of course, as you talked about the government, whether it's the city government, county government, national government has taken a role in governing and overseeing transactions and putting guardrails around all the other factors of production. But we just don't seem, I mean, we haven't cut up in terms of kind of treating data fundamentally as so structurally important to the economy. I mean, the old obviously cliche is data is the new oil. We're certainly not acting like it, right? Yeah, exactly. Exactly. Well, this has been super, super interesting. Obviously, a ton of work that you've done on this. And just in case it's not clear, the work that Kendra has done on this in case not clear listeners, specifically with an eye towards informing US policy. So everything we do at Tribune is kind of trying to understand China. But this was like an effort to understand what China is doing in order to kind of make strategic recommendations on how the US might want to be thinking about these issues. And so that's one of the reasons that we kind of leaned so heavily in the last part of the conversation on what the US is not doing here. I think this is great. I hope that this work gets some uptake from policymakers and people in that space. We will keep sounding the drum or pounding the drum, sounding the alarm. I don't know. Sounding the gong? Yeah. And yeah, well, I'm sure there will be a lot more opportunities to talk about these kinds of things. It's always good to kind of take a step back and do kind of a wonky or higher level, wonky higher level. Those who may be at odds. Anyway, I'm rambling now. But this was amazing. We'll just leave it in that. Thank you, Kendra, for the time and for walking us through that. I found it super helpful and fascinating. I'm sure our listeners did as well. Awesome. Well, always going to be here. All right. Well, thanks so much. And thanks everybody for listening. We'll see you next time. Bye, everybody.

Podcast Summary

Key Points:

  1. Western companies are increasingly using Chinese open-source AI models due to their much lower cost, which makes AI projects profitable for low-stakes tasks like document processing, compared to expensive US models.
  2. US models can be 2 to 20 times more expensive than Chinese alternatives, leading firms to choose "good enough" models for cost-sensitive applications.
  3. Chinese models are superior for processing Chinese-language content, but the key advantage for most firms is cost.
  4. The US government is considering restricting access to Chinese open-source models via executive order, potentially using federal procurement bans or commerce department supply chain rules.
  5. Restrictions would be difficult to enforce (e.g., targeting cloud providers or model repositories) and could harm US competitiveness if no cheap domestic alternative exists.
  6. Blocking Chinese models risks making US firms less competitive globally, as the rest of the world adopts cheaper Chinese technology.

Summary:

The podcast discusses the rising use of Chinese open-source AI models by Western companies due to their significantly lower cost compared to US models. Host Andrew Polk and tech policy researcher Kendra Schaefer explain that for many low-stakes, high-volume tasks—like processing policy documents—US models can be 2 to 20 times more expensive, making projects unprofitable. Chinese models offer a "good enough" quality at a fraction of the price, enabling innovation and cost-effective prototypes. This cost dynamic is a key factor driving adoption.

The conversation then turns to potential US government restrictions on Chinese open-source models, such as federal procurement bans or rules targeting cloud providers that host these models. Schaefer argues that unless the US develops a cheap, competitive domestic alternative, such bans would be counterproductive. They would be difficult to enforce and could harm US global competitiveness by forcing American firms to use expensive models while the rest of the world benefits from cheaper Chinese technology. The discussion highlights the tension between national security concerns and the economic realities of AI adoption, emphasizing that blocking innovation without a viable substitute is a risky strategy.

FAQs

The episode focuses on Chinese regulators' approach to data as a factor of production, along with recent developments in AI open source models and cost dynamics.

Chinese models are significantly cheaper and good enough for many use cases, such as processing policy documents, making them cost-effective compared to US models.

US models can be two to ten times more expensive, and in some projects, up to 20 times more expensive.

Possible restrictions include federal procurement bans, limiting government suppliers, or using ICTS rules to ban hosting on US cloud providers or listing on platforms like Hugging Face.

Both describe their vibe as mellow, with no major crises in the China tech space recently.

Kendra is going to Taipei, Taiwan with a Brookings Institution delegation in a couple of weeks.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.