Go back

The Hidden Risk in Crypto No One Is Talking About

33m 44s

The Hidden Risk in Crypto No One Is Talking About

The conversation explores the gap between blockchain's promise of transparency and the practical challenges of data reliability. Using the Layer Zero hack as an example, Victor explains that on-chain data can be manipulated at the RPC level, leading to false information that protocols trust. While blockchain itself is immutable, the infrastructure around it—such as indexing and data providers—introduces vulnerabilities. Indexing is essential for translating raw blockchain data into user-friendly formats, but errors can cause financial losses, such as missed trades or incorrect airdrop distributions. DeFi's fragility mirrors traditional finance's supply chain risks, but its openness allows for both security research and exploitation. As AI agents begin to analyze blockchain data, they excel at generating high-level reports but are not yet trustworthy for executing trades due to latency and accuracy issues. The biggest data-level attack vector remains the security of centralized data providers, which are often less rigorous than the blockchains they serve. Ultimately, while blockchain offers a decentralized record of truth, the ecosystem's reliance on intermediaries means that data quality and interpretation remain critical points of failure.

Transcription

4969 Words, 28032 Characters

English
This is the crypto hipster podcast. This is not a traditional interview show. These are perspective-driven conversations with founders, builders, and independent creators shaping what comes next. We go beyond headlines, beyond hype, and beyond price to explore ownership, freedom, and opportunity in the digital economy where builders talk freedom, not price. Today, we're not talking about price, we're narratives. We're talking about something that's far more dangerous. What happens when the data that we trust in crypto lies? I have Victor, from Army Labs today. So Victor, let's ask you first, when does on-chain data stop being truth and start being interpretation? This is first of very relevant question because this is right following the heel of the recent hack in layer zero with ABE, with draw frozen all that. So this is crazy because, for example, with on-chain data, first of all, what source is RPC? In fact, a lot of the data comes from RPC. What happened in the recent event was that that hacker was able to manipulate the RPC between the false data. Well, I mean, well, de-dossing the other providers. So in one sense, they end up having bad data. And because of that, the hack was able to happen. This is just one of the potential failure scenarios in the entire crypto or web3 stack. Let's slow that down for a second. The blockchain didn't fail. The data people trusted did. Data quality or data accuracy, just potential failure scenario by what they really truly considered. But I think after that hack, this actually has been at the forefront of the industry. It's a problem that people actually do not know what's the best solution to that. So people say that crypto is transparent. Is it understandable and transparent? It is transparent, yes. And it is understandable, yes. But when there are hundreds of billions of transactions to be able to normalize and to actually find out what is actually true, what's actually not, that's a non-trivial, you know, infrastructure, you know, to get to that point. So in a way, it is transparent. But to actually verify that accurate transaction, it does take, you know, resource. So that's the draw, that's the, you know, yeah, that's the hardest, difficult, difficult side to that. You mentioned Abe. From what I see and from the comments in the market, there's a lot of people didn't understand what was going on. So when I say is crypto transparent and understandable, we have, we have proof that it's not understandable, right? I think let's put this way. It's, I think you might need something close to a PhD to really truly understand the Empire Stack because it is transparent, but it's a very complicated system. All the source code is out there. For obvious cases, all the failure scenario is there, which is good, right? It's much better than like, you know, some, you know, centralized, you know, a couple of people making decision in the room. But just a matter of fact is because this has to be globally resilient. So therefore, there's so much complexity built in. And that's why, you know, normally people have professional teams to audit. It is transparent. It is there, but it does take a lot of, you know, mental understanding to really grasp on what's going on. So this is just, you know, in a way, the trade-off, yet to have something public, but complex, verifiable or have something that's in the black box, you have tried to, I don't know, authority. So this is just the trade-off, I think there is. They're saying that that in order to understand the blockchain, you need a PhD. Maybe not a PhD, but I'd say somebody who tell you they understand exactly all the failure scenarios in DeFi, they understand every single piece of infrastructure that goes on crypto completely. I think that's a little bit overstated because, you know, similar to the modern cloud infrastructure, right? You know, it's fairly complex, but I think it's a sign of maturity because the complex system is mostly dealing with all those corner cases. So I'd argue, yeah, purchase action and validation, whether you're funds are actually saved, that is easily verifiable. And it's an open platform that people have been building towards that optimization. But you're saying, hey, do we understand, can we see into the future that the ABA or layer zero hack happened that we could have proven in that in a way? Yes. But, you know, that's kind of one of those, you know, surface that people have not really fully, you know, thought through. And so that's the part is I can't need a security in PhD under the truly understand this. Yeah. Yeah, most people don't have that. And in love not having that, they have to, they have to trust on sharing data. Right. So where does the on-chain data mislead people the most? So maybe let's, if you don't mind, like, maybe I will kind of come back to this question because I think I had something I wanted to add earlier. Like it's also related to this. So in a way, it's not the on-chain data misleads people. It's just at this very moment, right? When you look at on-chain data, people go to block explorer, people look at hacks, a decimal, transaction. Those are machine friendly, but those are not human friendly. This is not just individual person. That's why every institution is being well-crypto. Every institution is buying crypto. But they're looking at Bitcoin because that's easy to understand. They're looking at stablecoin because that's easy to understand. DeFi, it's like about five percent adoption. Why? Because it's complex because, you know, there is a risk associated with complexity. We're building towards there, but it's not, you know, there yet. And part of that is just, you know, the complex is the data. Right? Now, when you look at on-chain data, okay, there's balance. You have to understand the details of every single smart contract in order to know two piece together, how the routing works and all that. And then you would need the secondary service, you know, to build data model on top, right? whether that's supplying live data to application, that's what we do, or like analytical, you know, type of dashboard. But that would also involve secondary human resources, the piece together, you know, a human digestible data model. So that chain is, you know, kind of that application level of filtering, forming data model, representing to an end user, that historically has been very expensive. Yeah. So it's not, it's not, you know, it's not like it's false. It is true. Everything's there is true. There are differences, right? Different RPC provider, maybe returning slightly different truth, but those can be normalized. But, you know, in order to normalize generate the accurate data model, it does take a lot of resource to do that. You're saying everything's true. Why did AVE in particular, you might want to use another example, or if you want, but why did it fool so many people? So the, it's, it's food, the protocol, right? That's because of the protocol had, you know, three or four RPC source that they trust. And those are pieces of centralized provider. When you manipulate two and deny the rest on the application level, when they're doing that validation normalization, it just attacked surface that you did not, you know, sit there. But that's more like a failure in like the system in a way, like application levels, not blockchain levels, more further stream application. So this brings you to another kind of point, right? You know, blockchain, when the first course conceived is very much optimized for writing. It's not really optimized for reading. People talk about, oh, how can we make sure one transaction return there is not being built, but spending is not being happened, you know, like kind of transacted the second time. Same money isn't spent twice, but people really haven't thought about, hey, once a robust application ecosystem started to form on top, how you get data reliably and very identifying that data, all the blockchain, that's a whole other piece of complex infrastructure. Where the hack, where the issue happened, yeah. - So I want to get into that, I want to get into the system risk. Says, you know, it's one more piece of architecture and talking about the architecture, you know, indexing. Indexing, where does indexing sit in the stack? And why is that important? And how does that look? Hope to lower the system risk. Okay, so indexing is essentially the most critical piece, right? Because the applications we interact, if you're a DeFi user, if you're not dealing with a crypto, you're lending position, you're trading, just anything that happens on the blockchain, that gets through an user via indexing. It's simply, you know, tracing what's happening on chain and present to user in a way that they understand. And then there are different indexing methodologies. Some of them give you larger amount historical data, generating a snapshot, you know, some of them are really need to stay on top, like within one second, especially you're doing trading. So those are theoretically different architecture, but they're all called, you know, indexing. So indexing is just getting data off of the blockchain. And then if that data, if that first thing, if that data layers wrong, right? - Yes. - What downstream is gonna break? - So from our real experience with the clients, people may have raw money, like raw amount distributed if they state or they participate in some on-chain, you know, airdrop, some trades can disappear, they try to place a trade and that doesn't show up. And then essentially you miss that window of opportunity. So real money is at stake, essentially. Yeah. - Everybody should, people can lose a lot of money. - Yes. - Absolutely. - This isn't theoretical. This isn't maybe someday. We're talking, real money based on data that might or might not be right. - It doesn't show up because one transaction or one block is misindexed that trade will not really go through. - So you're saying everything is based on their interpretation of the data? - It's not based on the interpretation. So think about this way, right? You know, indexing is what data taken from blockchain, blockchain is distributed ledger of truth, right? So indexing is this short way, hey, you know, I'm a trusted provider at that data backed user. This is the app that displays the user. The user takes them, believes it. And then you make a decision in trading, right? Based on that data is a transaction back to the blockchain. Blockchain is, you know, accurate. But if that data you got from indexer and you make a decision, but that data is wrong. Like the blockchain is not going to accept your transaction. So in that case, you know, you would be losing essentially that opportunity, that trade opportunity. And for cases such as, you know, token drop, sorry, air drop token distribution, that type of thing, you know, user will say, hey, you know, you give me the wrong amount or I put down this amount. So user will end up disputing with application application have to spend more money and double verify what's that transaction. So all that is additional cost essentially. I mean, your money is still there on a blockchain. It's just like when it's time sensitive, or when the data is wrong, you have to spend more money to validate that, or you, yeah, so. - So I know Bitcoin is supposed to be anti-fragile, right? And there's been recent concern that DeFi itself is very fragile instead. So how fragile is the system if something upstream such as data or indexing or anything else shifts, how fragile is DeFi and how sustainable is that? - I think it's less fragile than the street of hormones right now, because that can be for whatever geopolitical issue that can be interrupted. So I'm bringing this example up, you know, it's actually in a way it's similar to blockchain, right? So yeah, again, blockchain is decentralized, all your money is safe on blockchain, but there is a bunch of chains or supply chain, right? To get to the final application to work. Any of that supply chain can break. The good thing is there have been, you know, a few like major providers in any of those, you know, supply chain. So but any of them can break. So you can say, if you're asking the honest opinion, it's as fragile or as accurate as resilient as any of those supply chain vendors or infrastructure providers are. Hence that's why, you know, for DeGens, you know, people just, you know, Blany trust or they think this will not happen to them. But as blockchain or crypto gets into more institution, there's going to be more other requirements, there's going to be more over sites. Yeah, so just kind of to secure that, you know, supply chain infrastructure supply chain risk to make sure that doesn't happen. But as with the all financial systems, right? You know, global financial crisis, even bankers on Wall Street due to greed or whatever, they can create a systematic risk. That similar risk absolutely exists in DeFi. I mean, this is just a nature of a global financial system. But the good thing is, you know, it's not closed, it's open, right? You know, for better or for worse. The better is North Korea, better is like some security research that can submit a security proposal, versus North Korean hackers coming while also attack. So it's a free fall play field. Yeah. Yeah. You just said institutional risk and I used to work in corporate America and I have no trust as institutions are providing you accurate information. So, you know, so you trust the blockchain that's provided information by institutions, you know, how do you ensure that institutions are giving you the accurate data or are they lying to you? You know? Yeah, well, I mean, this is unfortunately is the existing state because there have been projects, right? You know, such as the graph trying to, you know, set up a decentralized way of getting data of the blockchain with different decentralized way to verify that. And there are many projects that are out there. But, you know, the so far, the economics of that have shown it's not really very sustainable to operate that. So eventually, you're more or less settled down into a scenario like a who are the four big cloud providers, Google Amazon, Microsoft, right? You know, three cloud providers. So that's very much is what's happening in the blockchain infrastructure space. But obviously, if you want to go truly decentralized without 100%, you know, vendor risk or, or, you know, kind of tied to any registered institution, you can do that, right? But, you know, it's similar to how, you know, you can call the dark web or the internet that's not really be indexed. It's just going to have a very niche use case. But for blockchain to be adopted by billions of people, there has to be more than institutional scale. And what has shown work is really just a centralized provider auditing data out of the blockchain. The record of truth, obviously, that blocking itself is decentralized. - So let's talk about the recordage group because, you know, we're now taking that record of truth and we're layering AI agents on top of it, right? And there are these agents are consuming the data. And, you know, a lot of times, AI is not an accurate representation of the world. So what do you think could go wrong? And are we headed toward like automated decisions based on flawed signals? - This is a very good question. And the short version to this is it's a work in progress. So, you know, at Warmeer, what's coming next is actually we're building like an end to end AI infrastructure just for blockchain data needs. So in this, kind of from our actual work we've seen, are, you know, having AI generating smart contracts, that's probably a bad idea because, you know, a lot of attack services, right? Even people's wipe-coded apps get hacked. Yeah, it can be good at finding exploits. On the data retrieval side, on the core data infrastructure, getting data out accurately, that you might still want to have a human to double check or AI to validate. But what AI really is good at is, once we have this pipeline built out, right? you know, get in an emphasize. people understanding what's actually going on blockchain is difficult is because it's very much machine language. So what AI is really good is that it takes all their raw transaction. Generate various reports reports that would have normally taken, you know, cost $100,000 just to have on report AI and generate at 10 seconds. It gives you an instant snapshot of what's going on the blockchain. But if you really need to go back and validate sure, you know, you can't spend resource on that. So that's what the AI is really good at generating data, producing, you know, the final product and yeah, and eventually AI will be consuming that report and making decision. And then there's a transaction transaction process, which AI right now is not really good because that needs to be very fast and then needs to be, you know, accurate, because trade decisions. millisecond, in milliseconds, in sub-second depends on that that you cannot wait for it to go to AI because they can actually screw it up. That you have to really have a very clean pipeline or very clean circuit that goes to the end trade application. Because that's actually tied to real time financial decisions. So the agents actually then since they're not good at transaction, they actually are the amplifying the truth or the amplifying noise. So right now I would not trust the agents to do trades. I mean, they might be good at, you know, doing macro back testing or analytics or the patterns, right. But the actual placement of the trade, I think most serious products there are it's actually the traditional software engineering way. So in the way you're at AI is very good at giving the macro trend, but the micro trends actually I think might might still be good for human to do that. But you know, that is changing because there's more and more tooling for AI's to validate the transaction truth. So hopefully later in the future, you know, AI can be 99.9% correct correct. And it's much better than I say two years ago. So you're thinking that are we are we creating the systems that act faster than than we can verify right now. We're definitely creating a system that's producing more data actually real data, not not, you know, convoluted or made up data, then what we can consume. And then on the transaction side theoretically yes, because an agent, of course, can place more trades right in its automated than a human ever will. But you know, from my limited understanding is it's on such a critical path, you know, the real adoption that's still waiting for AI agent to be a lot more accurate because you're handling real money. If there is a mistake right who's going to be responsible for that. But if a post mortem analysis or pre-trade analysis, AI is very good at giving that report. You know, AI could be flagging, hey, now this pattern happened in the market, you know, do a trade or something. You know, those are rest of that is a software problem. For the low low latency stuff, probably AI is not the best right now, but you know, AI can alert, hey, either places trade right now. You mentioned earlier, you mentioned Lazarus group. And there that that's just one of the one of the many hackers that attack systems right. Right. If you could, if you can name one data level attack vector, not necessarily a hack, but like attack. That's the potential. Yeah. What would what would that be? So the biggest risk right I'm seeing is very similar to what happened with the layer zero hack rate, no, because you know, way there are data providers out there, RPC providers out there. But they themselves are in a way that the where the host servers, how they manage security is actually much less rigorous than the traditional, you know, enterprise or blockchains right. So. And you know, that's just at a very fundamental like hey, all the blockchain data source come from say five six. Okay, maybe there are 10 provider are there. Yeah. So the same thing could actually happen again right for different. Daps out there. So that that that's a very one very critical risk. And the rest is mostly just application level. How did you decide to treat the data that return. Do they have secondary, you know, security or validation normalization of the data, the retreat. Those are more questions that's up to those individual applications. So that's why you know, DeFi has rescribed in we can trust a top two or three portal costs. The ones that are less rigorous, you know, who knows, maybe they're too small. That's why they have really been targeted. Yeah. But but hopefully with all these attacks, right standardization will form people get smarter hackers getting smarter, but a provider is also getting getting smarter. Yeah, just similar to now you don't really be barely see any windows virus, you know, anymore, like 20 years ago, right. So it's a it's a similar catamount scheme essentially. I don't I like the word you used. I don't like it as a word, right, comes to technology and money and use hope. So hopefully, you know, what actually, you know, that major attack vector that you said, I want, you know, nobody's talking about it, right. But it's actually real. You just shared it with me. You know, why why is nobody talking about that potential failure mode and what isn't that where are you. Because I think it's the first time that happened. I really think it's the first time it happened. Maybe it will happen second time, but I think the second time happens people will be, you know, be more aware of that. And that problem is actually a fixable problem. Because in the application level, as long as they add actual layer of verification, things like that probably will be patched will not happen again. So that's why I see, you know, like someone as you're writing software, there is a bug. You may not realize, but user uses it and they discover that and that gets patched. But the problem here is, you know, there is real money here. It's just an evolutionary process. So I don't see it as a long term systematic risk. Yeah, but you know, short term, we make cost a couple more hacks. But I think that you can be patch up. So going back to crypto data, the kind of the major theme is like that's what's so difficult about crypto data like data is there right real truth is there, but it's just so much blockchain, so much data on the blockchain. To take out the actual piece you need and validated. You know, that's a lot of work and the hack like that happened is because, you know, people just didn't see the need earlier to spend more money making more robust to double down on that data integrity, you know, side of things. Yeah, so that's what happened just just the fact that, you know, infrastructure on blockchain data is very expensive. It's very difficult to build. So I think that going to your project to to to warm me, right. How are you improving this landscape. And what's the hardest constraint you've had to build so far why or to face so far while building. Okay, so the funny thing was, you know, talking about back to the RV hack, right, you know, the reason we kind of for saw that happening is because we dream our time working with our clients, we actually encountered a very similar issue. We do realize, hey, you know, the source data from our PC can be wrong. So we build a redundancy layer to validate to check to make sure data passed on the client are correct. And as a background, we power some of the topics changes that trades, you know, billions of dollars per day, you know, and essentially every 10 seconds we go down or we slept, you know, a hundred thousand dollars are at risk. So we have that guarantee we need to make to our clients and make sure. You know, the data has to go through so part of that is, you know, iteration process over the years, like the short of two things we need to make sure data is fast and data is accurate and we need to turn that into, you know, digestible data to the application, meanwhile, make sure it's fast and accurate. So it's an iterative. So which means that you constantly improve it with every iteration. Where were you? Where were you wrong? And now you predicted, are they? You must have been wrong early on. And and when you were wrong early on, how did you, how did you face that problem? What is the problem was, and how did that early not be right on help you to shape and build your interim process for the future? So, take a, for example, in the beginning, like if it going a little bit more like technical, right, you know, what's called the raw blockchain that would provide us, or RPC providers, right, there are big names out there. Some of them are even, you know, valued in the tens of billions dollars. So you would trust, hey, because they're big, they would give you the accurate data. And we kind of initially made the mistake of, you know, relying on a single provider, and we passed the wrong data to the end user. And then, you know, they kind of give us feedback, hey, you know, this is the financial loss we experienced because of, you know, it trace all the downstream, hey, who's the provider that's there. So that's what I mean by like, hey, we realize this is the problem, we realize, we checked all the existing providers in the industry, that does initial, you know, that runs the blockchain node, right. None of them have actually accurate data. So you do need to, you know, do another level of consensus, filter normalization. And, you know, that's kind of part of the iteration process that was talking about. You know, we, because we deal with such large, you know, number of transactions, and a lot of real money and stakes. You're able to, you can say like adapt evolved during a very, you know, stable infrastructure in this way. But, you know, many smaller, like, daps right, who would be our clients, who are, who have not come to us yet, you know, you know, they probably will have a very similar issue. Those people who are trying to do the house, but just because the scale on their side is not big enough, that problem really hasn't been like, you know, forefront, you know, on their mind. Yeah, some data may be missing, but, you know, I don't know, it's like $3,000. So people wouldn't care. But it's a different whole different scale when, you know, one of your clients traded billions, you know, per day, right. That's a whole different scale. What I just heard you say is that you relied on one provider. They were wrong and had a massive cost all the way down the stream. Correct. Yeah. Yes. Exactly. Absolutely. Yeah. And a better method is a better approach is to take data from several providers so you can make sure there's no outlier. Correct. Yeah. That's the, that's the most foundational basis, right. So you're talking about like, hey, the first step of getting data out of blockchain is me to run a node, right, blockchain node, or use a provider. And the funny thing is when you run your own node, it's even worse data quality because of all kinds of issues that can happen. And then the other thing is, okay, once you have all the data right in your own indexer or your own database, you need to the next stage when you're transforming that data, you need to make sure nothing goes wrong either. So that's more like a separate layer, the next let layer. So you're asking me what indexing earlier is that in the traditional term, they call it EPL, like in extract, you know, transform load. So you know, raw data into finished product, finished data. So, yeah. Very critical piece of infrastructure. Yeah. Excellent. So most people think crypto is biggest risk is volatility. It's not where you're describing is something much deeper. And that's uncertainty in the data itself. Correct. Yeah. Director, I appreciate you going there with me today. Thank you very much, Riton. Thank you for your time.

Podcast Summary

Key Points:

  1. On-chain data is not automatically truthful; it requires interpretation and can be manipulated, as seen in the Layer Zero hack where RPC data was falsified.
  2. Blockchain is transparent but not easily understandable due to complexity; verifying accurate transactions demands significant resources and expertise.
  3. Indexing is critical infrastructure that retrieves and presents blockchain data to users; errors in indexing can lead to lost trading opportunities or financial disputes.
  4. DeFi's fragility depends on the reliability of its data supply chain, including centralized providers like RPC services; institutional adoption requires oversight to mitigate risks.
  5. AI is useful for generating macro-level reports from blockchain data but is not yet reliable for real-time trade execution due to accuracy concerns; it amplifies trends but may introduce noise.

Summary:

The conversation explores the gap between blockchain's promise of transparency and the practical challenges of data reliability. Using the Layer Zero hack as an example, Victor explains that on-chain data can be manipulated at the RPC level, leading to false information that protocols trust. While blockchain itself is immutable, the infrastructure around it—such as indexing and data providers—introduces vulnerabilities.

Indexing is essential for translating raw blockchain data into user-friendly formats, but errors can cause financial losses, such as missed trades or incorrect airdrop distributions. DeFi's fragility mirrors traditional finance's supply chain risks, but its openness allows for both security research and exploitation. As AI agents begin to analyze blockchain data, they excel at generating high-level reports but are not yet trustworthy for executing trades due to latency and accuracy issues.

The biggest data-level attack vector remains the security of centralized data providers, which are often less rigorous than the blockchains they serve. Ultimately, while blockchain offers a decentralized record of truth, the ecosystem's reliance on intermediaries means that data quality and interpretation remain critical points of failure.

FAQs

On-chain data is true at the blockchain level, but it becomes interpretation when it must be normalized and modeled for human use. This process involves resource-intensive steps, and different RPC providers may return slightly different data, requiring secondary services to build accurate, digestible data models.

The hacker manipulated the RPC sources that Aave trusted, feeding false data by denying other providers. This caused the protocol to act on bad data, leading to the hack, even though the blockchain itself did not fail.

Crypto is transparent because all source code and transactions are public, but it is not easily understandable due to its complexity. Understanding all failure scenarios in DeFi may require deep expertise, like a PhD, making it difficult for average users.

Indexing sits between the blockchain and user applications, tracing on-chain data and presenting it in a human-readable form. It is critical because applications rely on it for functions like trading and lending; if indexing data is wrong, users can miss trades or dispute token distributions.

DeFi is as fragile as its supply chain of infrastructure providers, like RPC and indexing services. While the blockchain itself is decentralized and secure, upstream failures can break applications, causing real financial losses, similar to systemic risks in traditional finance.

Currently, most institutions rely on centralized providers to audit blockchain data, as decentralized alternatives like The Graph have proven economically unsustainable. For mass adoption, centralized providers are the norm, but the blockchain remains the decentralized record of truth.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.