In this podcast interview, Stuart Reed, CEO of Notable, discusses his journey from quantitative finance to building infrastructure for real-time web data monitoring. He reflects on his quant experience, noting that using widely available data with complex models like deep learning often fails to generate alpha; instead, success comes from combining simple models with well-justified alternative data. Reed highlights critical factors for data evaluation: coverage across securities, point-in-time accuracy to avoid temporal biases, and overall trustworthiness. He cautions against over-reliance on vendor backtests, advocating for detailed metadata and transparency. At Notable, his team crawls millions of web pages daily, employing techniques like cross-referencing, web archive verification, and detecting AI-generated content to ensure data integrity. This process helps financial institutions trust web-derived data for decision-making, focusing on robust, thesis-driven approaches over sheer data volume.
Welcome to Selling Signals, the podcast focused on how businesses actually monetize and sell data. Each episode we interview an industry insider to hear their experiences and lessons learned. The series is powered by Valsis, the company that transforms your data into investment ready intelligence products. If you enjoy the episode, please subscribe from wherever you get your podcasts. This week we have the privilege of being joined by Stuart Reed, CEO of Notable. Stuart has spent the last 15 years analysing a hard question from very different angles. How to make sense of vast, mess information in a way that's actually useful for decision making. His prize complex derivatives, used deep learning for market prediction at a quant fund, and now spends the time building infrastructure to help financial institutions search and monitor the web in real time. What I value most about Stuart's perspective is that it's shaped as much by what hasn't worked as by what has. He's very clear-eyed about the limits of prediction and about why having more information doesn't automatically lead to better decisions. At Notable, he and his team crawl and index tens of millions of web pages every day, with a strong focus on point in time integrity and traceability. From investment funds, asset managers, banks, insurers and corporates, pretty much anyone. Every conversation James and I have with Stuart teaches us something new and I'm sure this will be no exception for listeners. Stuart, welcome to the podcast. Yeah, thank you for having me and thank you for such a nice introduction. No worries, you've done some incredible things. First, I'm having an X-Quant on the podcast. Maybe we'll start there. Walk us through your time as a quant. Yeah, I think that's a great question to kick off with. My background is computer science and then somehow ended up as a quant, which I very much enjoyed. Spent time in the back office where I was doing, like you said, pricing of exotic and embedded derivatives that were sitting on Bank and Insurance Company balance sheets. Actually wrote CUDA code way back in the day before it was much, much easier to do than it is now. And yeah, then transition to the front office where I worked at a quant fund where we were using deep learning to try and predict the markets. And how the gardening leaves situation and actually spend a year as the head of data science at a NAG tech company where we were flying drones over orchards and using computer vision to find unhealthy trees. So it's definitely being a journey. But yeah, I love being a quant and every time I hear the word X-Quant, I almost feel like sad about it because I still see myself as a quant. Although technically you might call me like more of a machine learning engineer, I suppose. Yeah, that's fair. My guess is it comes in very much handy with your current position. So back when you were quote unquote a quant, were you using alternative data back then was it more traditional? Yeah, I mean, we definitely started out with more traditional structured data, prices coming from I guess who you would expect in the industry. And what I learned, one of the biggest things I learned was actually that if you are looking at the same data that everybody else has access to, you're probably not going to find any meaningful offer in it. And it doesn't actually matter too much what the methods that you use are like we were using the latest deep learning techniques and back then the cool thing was neural churing machines, which I'm sure a lot of people have never heard of. And honestly, none of it, really made much of a difference of a simple linear regression, at least as far as I could tell. So we did start looking at alternative data, but I feel like we looked at it just far too late. And we didn't give it the attention that it deserved. Yeah, I've heard even at some of the sort of most prestigious quant funds, so talking Renaissance capital, etc. Often when people ask them what they do, the answer is we do linear regression and a ton of feature engineering. And that's basically where the where the alpha comes in, it's picking the data set and then figuring out how to combine those different data sets into something that frankly, the model that's being run on top of them can be incredibly simple. Yeah, I think that's I think that that's probably quite accurate. I mean, obviously nobody knows exactly what goes on at these very very secrets of firms. It's actually one of the challenges we're selling to them is that they're buying your data and they give you somewhat little feedback to iterantly time. But yeah, I mean, that's certainly the experience that I had as well. Like we threw a ton of models with a ton of parameters and generally speaking, all you end up with is something that's overfits and which doesn't actually perform as well out of some solicited in sample. And where we had, you know, where we saw some success was when we spent a lot of time trying to work out like, why would this signal, why would this feature actually matter from like an economic standpoint, to really like having a strong basis for our presumption. And then using relatively simple models to then take that transform that, you know, residualize it and then feed that into, you know, trading strategies that would make predictions. I think that that probably still holds true today. Does that kind of like a fundamental approach at scale with quantitative techniques? You could look at it that way. I think that there are a lot of quants who believe that, you know, if you're processing some data, there should be some meaningful processing it. I think it also keeps you honest. Like if you're just taking every single day of point that you possibly can and you're just, you know, essentially throwing the kitchen sink at some crazy overparameterized model and hoping that you're going to find alpha, my experience tells you that's not going to work, right? I mean, maybe it works for other people, but from what I've heard that doesn't work. I think that having a very robust, you know, thesis behind the data that you're adding to your models and using relatively simple models that are less likely to become overfits and also being very careful about how you do your experimentation is, is the right way to go about things. And that tends to lead to I think some of the more robust results. And we go back to when you were assessing datasets, trialing, putting to production. Let's go sort of look at the sort of trial period. What are some of the questions you would think about or ask the dataset before you started, I guess, rigorously testing? Yeah, that's a good question. I think that there are a few things that matter a lot, right? So one is like, what is the cross section? Like how many, how many securities are actually represented in this data? And how far back does it go? I think that if you the more securities that a particular dataset covers, the less likely it is that you're going to be fooled by randomness. And the more likely it is that whatever signal you identify in it is likely to be true and not deceiving. And I think the other thing that really matters, and I got burned by this and that's why I care so much about it, is is the data actually point in time in the sense that that information people have different definitions of points in time. But for me, what it means is that this information was available on that date to actually train on. I think it's so easy to make mistakes when it comes to point in time data, you know, dealing with time zones is something that trips up almost everybody at least a couple times in their career. Dealing with, you know, lagged indicators which have reported dates, but then they're actually published at a completely different time is another thing that catches people out inevitably. And I think I slept walked into like a lot of these problems. And that is why I care so deeply about them right now. And you know, it's a big part of what we're focused on at the company. I'm not sure what other dimensions I would care most about like from day one, but I think coverage and like trustworthiness are probably the top two. And then as I said before, like having a thesis around it, like, why do I think this data should matter, right? I think that if you have no idea why data set could potentially matter, the odds of it actually mattering, I think shrink. So what I mean by that is that, you know, there should be some kind of like economic justification or economic rationale that sits behind the data set that you're looking at. And you want to exploit that and you want to use that as your, almost like you'll you want to start when you're going into looking at a data set. Like I believe that this data should matter. And that is why I'm going to allocate some time to it. And so putting, sorry, it does to keep asking you to put your ex-quant hat on. But if you were speaking to an alternative data provider, it sounds like from a documentation perspective, you'd probably want them to be providing you with some sort of data profile, which maybe includes the stocks that are covered, the number of stocks, geographic sector coverage. And then maybe even some signal validation, perhaps not a paper portfolio back test, but at least some economic rationale for why they're bringing you that data. Is that a fair description? Yeah, I think that's a very fair description. So I think maybe this is a somewhat contrarian opinion, but I think that back test coming from data vendors, including ourselves, are probably not the thing that wants most want to see.
And the reason for that is because backtisk lie. And it is incredibly easy to overfit a backtist. When I see a set of backtisk results, it doesn't move the needle for me, right? But when I see that a data set, and by the way, we do buy data sets as well, like when we're building and what we're building requires data from like multiple angles. So in some cases, I'm almost putting on, I've even had conversations recently with data vendors who are at exactly the same conferences as us, but we're both on the cell side. And I'm meeting with them after those conferences, because I'm looking for certain data sets right now. But when I see those firms have really meticulously mapped out every single corner of their data, and they know their data better than anybody else. So they can tell you the coverage of security as they can tell you where the data came from, when it was collected, how it was collected, what changes were made along the way in the process of collecting that data, what the lag is between collecting it and delivering it is, the latencies, all of those things I think they'll trust. And that's exactly how we do it. I mean, we provide extensive metadata data sets. We were having a conversation with the fund the other day, and they stopped me while I was explaining a data set, a metadata data set. So a data set about our data set, and they stopped me, and they were like, "Oh, do you sell that as well?" And I'm like, "No, that's free. You go download us and take a look." Like that's what you ought to be looking at and ought to be trying to understand if it is that you want to evaluate our data set. At least that's my opinion. Other firms certainly do want to see backtests and they have arswell backtests, but at least from my perspective, coverage data sets make a lot more sense. Another thing that would matter is the correlations, it's like can you actually observe any kind of statistical significance in this information that could justify really diving into it? - Yeah, I think maybe the much more useful thing you could do, I suppose, as a data vendor, if you're engaging with a quant fund, is almost do all the heavy lifting just up to the point where the backtests might start. Removing outliers, detailing perhaps in a Jupyter Notebook, how you've decided to clean the data with your very, very in-depth knowledge about your own data set. And then going, almost, we stopped here because there's no point, as you say, in running the actual backtests, but provided you agree with every step along the way that we've done, you can get started pretty much straight away with your own backtest as a quant. - Yeah, precisely. And I think that's exactly how we see it. That's how I see it. And that's what we're working towards. And I think that it's a moving target, right? And as you're, for us at least, we're constantly adding new data. We index about 25 million new web pages every single day. So we're adding at least 150 million like individual data points. And as we're doing that, and as we're expanding our coverage, things change, right? So it's as much of a moving target. So it's almost about building the processes around the data set, so that those things that you mentioned, just fall out of the process. That's how I think about it at least. - You mentioned trustworthyness earlier. How are you assessing that? Is it just you told me, A, and I'm gonna go see if A is true, and you sort of go down that list of coverage history, et cetera, how would you think about that trustworthy piece? And how can you prove that to a end user? - Yeah, that's a good question. So I think there are a lot of different dimensions to trustworthiness, and some of them are much easier to quantify than others. When I'm thinking about trustworthiness, I'm thinking about it in the context of the data that we're indexing, which is the web. So we're a search engine, we're trying to crawl index and cross reference, just about everything that's been added on the web. And the web, you probably know this, is full of junk, right? There is a huge amount of junk on the internet, right? And there are also some very perverse incentives that exist for properties on the web, right? Advertising related issues, content farms, bot farms, disinformation from state actors. So there are many, many dimensions to the question that you asked, and maybe I'll just tackle a couple and you can ask any questions that you want. But from a point in time perspective, like we're spending a huge amount of time trying to verify, is there are no point in time snapshots of the web, right? But what we need to do is get to a point where we're putting the web, which is this messy, scary place, onto a solid foundation so that quants can actually use it to run backtests and trust the results that they actually get. So how are we doing that? It's actually one of the reasons why we're a search engine in the first place, across referencing data. Can we find third parties who wrote about the same information at the same time? It's one of the things that search engines are probably the best in the world at is finding similar text, right? So that is one of the methods that we use for identifying and corroborating information. Another thing that we do a lot of now is web tracing. So taking those massive web archives that everybody has heard about because they've been used to train all the large language models that we use and actually repurposing those data sets for something else. We're using them to verify that the text, that XYZ website said that they wrote in 2016, actually exists in web archives from the year 2016. And to the extent that it doesn't, trying to understand why, were these websites not being indexed at that point? Or in the nefarious case, is this an AI-generated website that was created in 2023? And as backfilled 20 years with fake news, right? And speaking to that, some of the other things that we're working on is actually verifying that fact checking, a huge amount of fact checking. Did this website actually exist when it said that it wrote this content? Did the company that it is speaking about exist at that point in time? And that actually maybe alludes to some of the kind of data sets that I'm trying to buy. It's a point in time kind of securities, master type data sets so that we can go and double check that these websites are not portedly talking about companies that didn't exist. Another thing as well is catching them in the act. Like there are a lot of AI-generated websites that are popping up onto the internet. And they're using AI models to backfill their content, going back 10, 20 years. And there are lots of SEO and marketing gurus on LinkedIn who are very proud of this fact. But the reality is that none of these large language models are temporarily sound. Like none of them are anchored. They are hallucinating information from the future in the supposedly-- the supposed article that was written in the past constantly. So what we do is we look for those things. Is it talking about chat GPT before 2022? Is it talking about ObamaCare before 2007? Is it talking about Reddit before 2005? Like there are tens of thousands-- probably hundreds of thousands of concepts, of phrases, of words that have come into existence at a particular point in time. And any mention of any one of those things is a signal that we use to assess whether or not that content is reputable. So that's just-- I know that I just said a lot. Rent on a bit of a rant. Oh, excellent content. That is just one dimension of trust, in my view. That's like the point in time dimension of trust. But then there are other dimensions, right? So like one of the other dimensions is this original content with this website that is ripping off other websites and republishing their content a day later, right? And it looks like it's meaningful, but actually it was meaningful 24 hours ago, just a small example. Another dimension of trust is state-run websites. We know that there are geopolitics around us. We could bury our head in the sand and pretend that everything that is published on the web is true and the free of agenda, but that's not the reality. The reality is that people are trying to influence and manipulate us all of the time, right? Our view on that third point is that we should index everything that there is, and we should give people the tools to work out what of that information they want to include and which of that information they do not want to include. Because there are many different sides to truth. And I don't think it's our job to necessarily be the arbiter of it. There's a really interesting angle because it's actually quite an atypical version of point in time. If you're another alternative data vendor, typically what you do is say, OK, we collect the data in this way and provided we can essentially timestamp that effectively, we can know that we have point in time data. Historically, there's not a lot we can do. We don't have the timestamp. we kind of just have to accept.
that our historical data perhaps isn't point in time and let a fun know. Whereas you've got this sort of unique position where your web source, sorry, your data source is the web. There is a version of point in time in the web. It's just unreliable and reliable to varying degrees depending on source. So it sounds like what you're doing, well, it's obvious that what you're doing is essentially going through all those sources and rating them on trustworthyness. And then I believe, I think I saw from one of your slide decks that you sort of put certain sources in jail, important point in time, jail. And then I assume you just completely exclude them from the data center. So yeah, we can't actually guarantee, in fact, we can probably guarantee this definitely isn't point in time. Is that fair? Yeah, that's fair. Yeah, so I mean, the way that I like to think about it is that we are web detectives, right? Like when a detective comes to solve a homicide, they are the person who goes and gathers all of this evidence and they try and corroborate everybody's stories. They try and make sure that people have alibis. They try and understand like what are the incentives at play? Like why would somebody do what they did when they did it? You know, they go and try and gather as much evidence as they can. And then they compile that and that's their report, right? What we're trying to do that is what we're trying to do is essentially automate that process in the context of trying to put the web onto solid foundations. But it's true, it is a contrary and approach. It is very different to I think what other alternative data vendors would define this point in time. And to be fair, every single data point that we've collected for the last few years would fall into the more traditional definition of point in time. Like we have timestamps before when we crawled that data and we can be 100% confident in all of that. However, I think that in the case of web data, there is so much untapped alpha in it that I think it makes sense to be pragmatic. And I think that it makes sense to evaluate the data intelligently as opposed to simply saying, "I don't have a timestamp and an entry and a database, so I can't be 100% sure." Especially if one of your goals is to get to having point in time large language models or very, very large deep learning models, which as you know, require an enormous amount of content in order to demonstrate any emerging capability. Like the reason why chat GPT is so smart and the reason why Gemini is so smart and, you know, Claude, is so smart is because they were trained on so much data, right? If we overlay the very, very strict definition of point in time and we look at, so we have that as like one, let's call it an immovable object and then we have this other unstoppable force, which is that scale actually matters for training intelligent models. Like these two things do not, they cannot exist simultaneously, right? So what we're doing is we're trying to find a way to make the web more trustworthy and more point in time. And different firms have different levels of trust that they need to see, right? Like some firms are somewhat more permissive. Their strategies are not as dependent on, you know, you know, a strict point in time information as others. And what that means is that we almost calibrate our data sets based on like how how strict that firm needs that data set to be. But what we do certainly do to your point is, is put some websites in jail. We know that they're a fake, we know that they can't be trusted for whatever reason and they get automatically excluded. Either they get excluded entirely or they just get excluded from the quant data sets that we create. So preventing point in time homicide. It could be a time for you. Yeah, essentially. So this is probably a question I should have asked earlier. And the conversation we've previously had, I think I was a bit loose with my terminology and I was almost using web scraping and web search and kind of interchangeably and you picked me up and it's said, no, no, no, there is a really big difference here. So maybe now would be a good time to take the opportunity for you just to kind of high level explain what Nocible does from a web search perspective, what the product is for your quant users. And also maybe the thesis behind why you built Nocible in the first place. Yeah, thanks. So web scraping in my view is a targeted acquisition of data from like particular websites, right? It's like I want to go to Amazon's websites and I want to scrape prices so that I can use that in order to build some kind of predictive model or now costing model for CPI. That is is web scraping. We're not a web scraping company or we don't see ourselves as a web scraping company. We're a search engine. We're a web skill search engine. So I would say that we're very similar to Google or Bing or Brave or some of the other AI search engines that are popped up like Exa and Tably. We're indexing tens of millions of web pages every single day or index contains billions of web pages already and you know hundreds of billions of words of textual information, but it is coming from a very very wide cross-section of websites that cover just about every single topic under the sun. So we have information about finance for sure, but we have information about geopolitics. We have information about you know what's happening in agriculture around the world, weather related news, prime related information. So there is anything that is essentially being added to the web is something that we would consider like in our kind of domain of something that we would like to index. Now where we differentiate ourselves from the likes of Google and Bing and Brave and Exa is that whilst they all focus on offering APIs for you to go and get information from the web, we have specialized in web surveillance. So what that means is that we're trying to push the web to you. So our customers typically come to us and they tell us these are the things that I am very interested in monitoring or knowing about and what we do is we take those interests that they have, we compile them and turn them into searches, and then those searches are continuously executed over everything that we're indexing and any result that is relevant to any of those searches is sent to them either in real time as we index it or near real time or at the end of the day as a very very large kind of park hay file. So some examples of the things you might care about. Let's say you wanted to know anytime any company in the world is the victim of cyber attack or or falls you know or has a data breach of any kind right. That is something that is reported on heavily and what we could do is we could create searches that relate to that particular topic and we could monitor every single thing that we're indexing continuously and and deliver those search results in real time to the firms that need them. But you can track just about anything and we don't just have customers that are in the financial services vertical because there are lots of use cases for this in the corporate side of things in the advertising you actually had a guest recently who was talking about the media side of of alternative data and I found that very interesting because we kind of we kind of live in both worlds. How are you assessing the quality of the search results because I presume it's almost somewhat subjective like how relevant are the top results for the query that someone has asked and we've all used we've all been forced accidentally I think to use Bing at times and then had to kind of remove that as our homepage and go back to Google. So I'd be really interested to know what what is the almost the quality metric do you have test queries that you run? Yeah definitely we have we have actually millions of queries where we have like gold standard answers that we're constantly evaluating ourselves against. We have a very strict like version control system around where we release new algorithms and all of those are rigorously tested across not only different topics but also languages we support 95 different languages so we need to test how good are the queries that are coming back in Chinese and in Arabic and in German and in Spanish and in French which makes which makes it a very very difficult problem to solve and many of the new social algorithms that we were sure would be better have actually failed this process and have ended up in the in the dustbin of ideas at at mostable but to answer your question more pragmatically there are like there are different approaches to search right the more classical approach is lexical search so what that means is that we're going to take this content we're going to take this information we're going to tokenize it so we're going to split it up into different words we're going to stem those words we're going to treat them in different ways and what we're going to do is create indexes which allow us to work out all of the documents that match a certain set of words and then we score around that and that I think is very mature like there's an algorithm called BM25 which we use and I think every other search engine probably uses as kind of a baseline it works extremely
well over the years, over the decades, actually that's how all this algorithm is. Many people have introduced new search algorithms, new lexical search algorithms, and almost inevitably they they underperform BM25. So we use BM25 as our lexical kind of backbone. And then there's another way of doing search, which is semantic search. So instead of looking at the words that you've given me is trying to actually try and understand what their intent is. Like you could have a document that mentions absolutely none of the same terms, but is almost certainly a good answer to your question. So there we rely on embedding models. So how that works is we take the text, we feed it into neural networks, neural networks produce vectors, and those vectors are indexed into our system. And when we get a query from anybody by the API, what we're going to do is we're going to pass it through both of those systems simultaneously. We're going to retrieve documents that have similar words. We're going to retrieve documents that have similar meanings. And then we're going to combine those two. And that's what you would call hybrid search. We also take it a step further than that. So we look at end-grounds, which is like sequences of words that are important to your query. And whether or not they're prevalent in the documents that were returned. And more recently, we're also trying to look at knowledge graphs and trying to understand what are the relationships that matter to answering your query that aren't actually expressed therein. Like let's say you ask a question, like what are the which brands of Coca-Cola are selling really well right now? Right? Like that seems like a reasonable question. But if we treat that as a purely lexical search, we're going to search maybe for brands and Coca and Cola. Like that's it. And maybe doing well, like those arguably might be down weighted because they occur so often in language. If we take the semantic version, we're going to encode those words and we're going to have some meaning of that. But what would be really powerful is if we could do named entity recognition and say, okay, well, you've just asked me about Coca-Cola. Coca-Cola is a company. I know who Coca-Cola is and I know what brands they have. They have Fanta and they have Sprite and they have Minneth made and they have what else do they have? I know that they have a lot of pride. And we pull down all of that information. And when we're retrieving these documents, we're going to say, well, I know it doesn't mention Coca-Cola exactly and maybe the semantic meaning is a little bit off. But hey, it mentions Fanta and Sprite in the same document. So maybe we should boost that in our search results. So those are like some of the some of the ways that we approach search. And I think that search is like a very powerful backbone, four alternative data and four months because it's rarely about the signal to noise problem. It's like, how do I take an intent, a search intent? This is the kind of information I want. And how do I find only the pieces of information from from this massive corpus that are related to that? It's a very hard problem. My catchphrase at work is search is hard. And the more I look at it, the harder it gets. This might be a really dumb question, but naturally, web scraping is facing quite a difficult time in the media as publishers and businesses look to try and protect their IP on the web. Does this sort of data set, is that exposed to a similar issue as these businesses try to retain that IP? That is a very fair question. And I think that I think that this is a topic where there was so much change that if we look at this in six months time or a year's time, the answer could actually be different. I know that new legislations are currently being proposed and that is something that we are very much on top of. But what the law says right now, I think is a good place to start. And what it says right now is that if you're a search engine and you are not redistributing the full content of the web pages, you're redistributing relevant snippets of content that relate to user generated queries. That is okay. It is a transformative use case of that content and it is protected under the fair use laws in the United States and other copyright laws around the world. And that is precisely what we do. So unlike other web scraping companies who's I think are operating perhaps more in the gray area when it comes to this, we do not distribute like the full content of pages. We are not trying to disincommediate publishers. In many ways, I think that we could actually be very supportive of publishers going forward. And we also don't redistribute the raw HTML. What we're doing is we're answering questions and finding the most relevant documents. And I suspect that many of the funds that we work with are then going and pulling down information from those pages to enrich that themselves. And I think that that's probably the way that things should be done. I think that it's a more ethical way of sourcing data from the web. But like I mentioned to you before, it also improves the signal to noise ratio. Like you don't actually want to perceive 15 pages of content from every single website that mentions anything. It's actually too much information and it's not that helpful. I think that as a search engine, we're protected. There are lots of precedents that have been set and laws that have been written and we abide by all of them. And I think that we're also solving the problem of information overload. So I know that I'm talking my own book and some people might disagree. But that is certainly the way that I look at it. And things might change. I think large language models, since that's kind of the elephant in the room and that's why this conversation is very much coming up. I use them. I love them. I use them a whole bunch more than I use the tools that I used to use previously. But the reality is I think that they are to some extent substituted. They are actually replacing you going to those websites and reading that recipe for making cookies yourself. You are just going into voice mode and asking it to walk you through the steps one at a time and not actually looking at those pages. So I think that if we zoom out, what's happened is that the social contract of the web has been violated. And there is a new world order that's being established and it's yet to be seen exactly where the chips are going to land. And I'm not entirely sure what side of all of that I'm on. Yeah, and I would agree with that opinion. Because I mean, my day job, we are thinking about how we can protect the IP. A lot of the company that I work for is IP is on the public domain. To some degree, I can understand that's user generated and there are ways that we could do to protect that aspect. More of that would negatively impact how consumers can engage with that information. I think what would be interesting as Reddit, Twitter, all of the big media, publishers, think about okay, this textual information is really valuable. I wonder how much influence they will have over the laws that will inevitably change that. But it sounds like you're sitting a very different place to that. It seems like as a fund would see that feed come through and we'll talk about later how you identify some of the more useful data points that you collect. But it seems like we'd only encourage them to go and license that directly from a provider if it became quite clear that that that's the source that was solving the problems for the more commonly than not, if that makes sense. Yeah, I think that that does make sense. And I think that it's yet to be seen, like I guess how all of these things are going to shape up. But I think that the position that we're in is that we're a search engine. We're indexing everything on the web that we can get our hands on. We're following all of the rules that come associated with that. And we're not redistributing the full contents and the use cases that our customers have and that we're providing for are not like substitutive in nature. There's also a huge amount of transformation that goes into the process of turning this raw text data into something that is searchable. Like searches is at the core of our business. Making information on the web, which is huge and messy and just far too big. Easy to find and easy to fact check and verify. That's like the core problem that we're solving. And I think that our data sets are somewhat complementary to other data sets that exist. But they're also different. I think that what we're doing is very, very different. That's at least the feedback that we've gotten from everyone that we talk to. And I think we have a very different perspective on that. I agree with that. I agree with that. I think maybe a good question there is that I guess this concept would be quite familiar. Your data set would be quite familiar with different types of your data set would be really familiar with the firms that you're selling with. But there are other areas where you're having to describe to them the point in time nature, etc. How are you? What sort of roadblocks are you hitting in terms of having to educate buyers on this isn't web scraping. This
This is indexing. Yeah, I think that you hit the nail on the head. It is very, very different. Almost every fund we talk to is, this is very different at some point in the call. And what that, you know, that's a blessing in the curse. I think that the blessing part of that means that we're creating a new category of data that is much broader and as massive cross-sectional coverage and historical coverage, right? And that is what Hans really looked for or like some of the things that they look for very early on when they're evaluating different data sets. So that's on the pro side. On the negative side is that it is very different, which means they can't put us very easily into a bucket. And we do have to explain to them what it is that we do and how it differs from, you know, for example, news aggregators. Like news is certainly a large percentage of what we index from the web, but it is not even close to everything, right? The web is so much more than just news, right? But it's also so much more than just corporate news. It's also geopolitical kind of like information and all these different topics. So we do need to educate a lot as to like why it is that searches like a good fit for quantum. That has been going quite well. And I think it probably is going well because I'm an X quant and I can kind of speak to to that. And part of the reason why I built a search engine is because I could see the value of having a search engine as a quant, right? So that's been going quite well. The point in time conversation is also is contrarian. It's very different. It's very pragmatic and some firms love it, right? And they really are very interested in how we take it and how we scale it and what we can add to it. And I think that they have a very pragmatic approach looking at the world. They realize the same thing that we do, which is that scale matters for immersion capabilities and models and that there are no very high quality and low latency point in time web snapshots. So either you accept the fact that you need to be pragmatic. If you want to get to the point where you might have like points in time, large language models, which I think is on every quantum swish list. Or you say, I'm never going to have that. And I'm just going to get by on all of the strategies that we have already. And we hope that none of our competitors find massive amounts of alpha in that other direction. So I think that that's like the way I see it. And it feels like you're still on this sort of journey of getting to a mature go-to-market sort of function. What have been some of the lessons learned as you've built this business over the years talking to different end users that would be interesting to know? Yeah, definitely. I mean, we've been going for five years. Our first four years were pure research and development. As I like to say, you don't roll out of bed one day and decide to go and index the web. It's kind of an insane thing to do. But I'm kind of an insane person. So I like to do that. And only the sauce here have we really started like commercializing it. We're very, very fortunate to have had like an early quant fund as like a, as a very early customer who gave us a huge amount of feedback and guidance along the way, as well as a very early adopter advertising customer and some corporates who use it for, you know, competitive intelligence and monitoring. And we've kind of taken the lessons that we've got from all of them, which took us years to learn. And now we're, we're productizing that and kind of taking it to market. But I think some of the lessons learned along the way is that if you're doing something easy, like there, it's, it's, it's not something that people are going to buy quite frankly. You have to be doing something a little bit insane or you have to have like a really, really interesting and novel data set that is hard to come by in order for it to have like serious, you know, serious potential boom monetization, which is what your guys podcast is about. Like indexing the web kind of falls into that category. But I think, you know, that lesson was learned because I think some of the first kind of adventures that we took as a company were more focused on, you know, slightly easier to acquire data. And I think if it's easy to acquire, there's no reason to buy it. So that, that is certainly like one of the lessons learned. Yeah. I would agree with that. I think the, the only caveat I might add to that is there's a cost ratio, right? Is that, if it's really, really easy to acquire, doesn't necessarily mean it's useless, but does mean that, that your price power is, is a lot less. So there's a lot of data sets out there that are quite easily scrapable today and are in many, many different products. And therefore, the pricing power of that has diminished quite significantly. It would be the only caveat I think I'd add to that. The question I feel is, my guess is it's a huge data set that you have. There's a big, I think a big notion that these funds won't all the data that they can have, but to some extent, this is a massive data set. So like, has there been any feedback that this is too large? How have you thought about sort of navigating those conversations? Yeah, we've definitely had that feedback. I think the feedback for us has actually kind of been interesting. So the first one, when we started commercializing this last year, like we really sold it as just a pure search engine. So you put in anything it is that you want to search for and we kind of like deliver that to you as like a historical data set with like potentially hundreds of millions of records, right? And the first challenge that we encountered was that people when presented with a whiteboard sometimes, so they know what to put on that, right? Or like you presented with the chat GPT input box in 2022 and you're like, okay, what should I do? And I think that that was like actually the first problem that we had was that we had such a powerful capability and you could really do anything with it. We didn't demonstrate it. So now what we're doing a lot more of is we're actually carving out data sets on behalf of customers and we're demonstrating the value in multiple different dimensions, macro data, securities data, looking at product-related kind of like data sets, all content related to certain product categories and products that are on the market. And that makes it more tangible. People can look at that and they can see, okay, I get it. I understand what this is about. And then they can start to think about how they would like to modify that for themselves. Because unlike we don't see ourselves as a traditional data provider, we don't have like one massive static data set that we sell. You know, we sell the capability to find anything that you want on the web, historically pointed on to like put it onto a good point inside basis and use it to develop novel and interesting strategies. Hopefully that are using large language models and neural networks, which is interesting. So we kind of got past that hurdle. But now the hurdle is, okay, some of these data sets are bloody big. And we don't actually just ship the data either. Like we're adding embeddings, which are these high-dimensional vectors that are attached to every single data point. So these are massive, massive, massive data sets. And I think our presumption or our assumption, you know, at the end of last year was that all of these firms would have pipelines internally to process exactly this kind of data. And I think that some of the largest firms certainly do, but a lot of firms don't. And actually the skill set that we've developed that knows more within our team is very, very powerful, right? So we are, you know, increasingly helping them kind of like narrow that down, compress the data, work out what of it actually matters and extracting real signals from a stellar, you know, a tradable, much quicker to test and much quicker to go and validate. So that is kind of being like the evolution that we're on. But at the same time, we also don't want to lose our identity in that process. Like we want to make sure that everyone still understands that, you know, this is customizable. This is not a buy one time series and you have to buy every other time series with it. And that time series is what I have defined and what I've, you know, imposed on you. Like this is a search engine, you know, Google only works because you can go to Google and put anything you want in it and get back relevant results. And you can't just chat to be tea and Gemma only work because you can ask it any question. Like, no, so we'll can't lose this identity and the process of making this easier to consume. So what we're trying to do is find a way to balance the two. How do we productize this and make it as easy to plug and play as possible, but also make it as flexible and customizable as it already is. I don't know if that makes sense, but it does.
the, in summary, what you've done is first you index the web and now you're creating the web for a specifically investment use case. And so that means that in comparison to a web or search engine like Google, you've got the point in time obsession, you've got those very clear kind of demarcations macro into different investment use cases. And that seems to be kind of the real differentiator. You're just very, very focused on that investment use case. Yeah, precisely. But more than that, we've also optimized it specifically for, like I said earlier, bringing the web to you, right? Search engines like Google, I think, I mean, one of the challenges that a search engine has is because you can put in absolutely anything. And I can predict what you're going to put in. I can only do so much pre-computation, right? I can only cash so much information. So if you go in and you type in absolutely any query, they have to take that query, they have to spell check it, they have to process it, they have to tokenize this, but not the words, work out what parts of the index are the most likely to contain the relevant documents, they have to hit a neural network, they have to embed it, they have to process that embedding, they have to query and then aggregate and stitch the results. Like there's a lot of work that goes into that 0.5 seconds between you hitting enter and 10 links coming back. And if you're looking at the large language model or AI, AI-based search, there's even more work that has to happen because then you have to take those links, feed it to a large language model that they're interpreted and then tries to work out how to answer your question and which documents that were returned to site. What we've done is we've recognized that there are some things that you're probably always going to want to know about, like in the case of like a cyber attacks that we spoke about earlier. Maybe that is something that you always want to be notified of as quickly as possible. In that scenario, we can cache a lot of that information, right? We can say, I already know what you're going to search for, I already can work out where in this index it's going to be, I can hit the neural networks and embed everything already and I can pair it. It's like fully loaded bullets in the chamber, like all I need to do is pull the trigger. And then we have the benefit of knowing the last time we ran that search for you, right? So if I ran that search 24 hours ago, I only need a search over the data that we added in the last 24 hours. If I ran that search for you 60 seconds ago, I only need to look over the data that we added in the last 60 seconds. And that is like one of our key differentiators. So the obsession with finance and understanding the pain points that exist for funds and optimizing our retrieval system around that to some extent is a huge thing. But also more than that is like optimizing it for this kind of continuous delivery. And like I say, like we want to bring the web to you. Like I think the idea of you having to like constantly go and hit an API every 15 seconds to query some data is just inefficient and wasteful. You know, we can do that in a much better way. And we don't have to sacrifice the API experience either. You can just go query on demand. But that's not necessarily what we've optimized our entire infrastructure to do. I think this is a really great takeaway. And I want to read the quote that you said to me when we were doing the prep call for this, which was the data that actually moves the market in your data set is a very small fraction of the entire data set. And I think that's translatable to a lot of providers. You've touched on a bit there about being a customer led and learning from how your customers using the data and building things that are more off the shelf. But maybe more broadly, how it's a really difficult question, especially if you're new to the market about, okay, great, how to actually identify what is the most valuable part of my data set and maybe talk us through how you do that outside of being customer led. Yeah, I think that that's exactly where, yeah, I mean, it's my own quote. So of course, I agree with that. I'm not going to disagree with myself. But although I do sometimes on occasion disagree with myself, the fact of the matter is that a very, very, very, very small portion of the web is going to move the markets, right? It is paid attention to the right people at the right time who have the right trading terminals to move the price, right? Or maybe as our people, maybe as algorithms, right? Probably algorithms. But how do we go about, I actually identify in that. And I like to think of what we're doing, almost like an optimization problem. So you start off with the exploration phase, right? Which means you have to evaluate a lot of different options. You have to kind of like, you have to cover everything so that you can get, you can get a map of the terrain and you can work out like, this is where all the mountains are, this is where all the rivers are about. Only some of that is actually going to matter to you. And then you kind of move into the exploitation phase, which is like, okay, I've mapped the space. I have all the data at my fingertips and I have the tools to be able to slice it and dice it and work out what actually matters. And then you use those tools to kind of narrow it down. So I think that we've kind of, we've done the exploration phase. Like we have a very, very good and very high frequency kind of understanding of the web at scale. And you know, the coolest thing about that is we can actually see the web. Like we can see when America goes to sleep, like in the data, like you can literally see it. We can see like when it's weekends, right? Like, like we have like a pulse of the internet. But what that also means is now we can take that system and we can take financial markets, which fortunately the data is quite easy to get access to structured data is not very hard to acquire. We can buy those two things and we can start to draw relationships between them and work out like which are the publishers that publish first, which are the ones that follow on, which are the ones that when they publish markets move, right? And you know what, it might not be the biggest ones. It might not even be the biggest brand name ones. It might be that random sub stack by that guy who is very famous for like having shorted, you know, a whole bunch of stuff. Maybe that's what means. It's atariously only got one right, I think. But that's my point, right? It's like, it's not necessarily like the biggest and the loudest properties on the web that are moving things. It's these niche little crevices that people have learned to pay attention to in that matter. And the only way to find them is to index everything, right? So we're kind of like in the process of now, you know, now that we're we have a very, very strong coverage of everything. Now I'm trying to kind of zoom in and work out what of that is is very material and then extract extract the best possible signals that we can from that. So I guess maybe to zoom out a bit there is understand the data that you have and where you're collecting it really strongly. So mapping to tickers, understanding where you know, even if I think about the reviews data I have, we cover millions of businesses, but in reality there's probably 50,000 businesses that we get significant reviews on narrowing down on that and understanding what aspects of the data could be correlated with KPI's market fundamentals, etc. Is that a fair way of looking at it as a thing? Yeah. Yeah, I think so. And that actually brings us right back to the beginning of the conversation, right? Which is like actually more important than a backtisk is how well do you know your own data? And I think that that kind of that is a perfect example. It's like how well have you mapped it? And for us, you know, it goes far beyond just tickerizing the data. Like we're doing a project right now. Maybe I shouldn't mention it, but I'll do it anyway because I'm sure that it will work out. You know, these things, but what we're trying to do is try and work out like what are the what are the economic structures that exist behind the web? Like who are the businesses that actually own these websites? And what are where are they physically located? And who are they run by? Who are the people? And what are their agendas? And that is that is a project that we've been working on for a while. And it's incredible the amount of information that you can extract from so many different like disparate sources. Like a couple of examples like many websites that are run by whose business model is advertising, for example. Like we have mapped out every single ad sense ID and every single advertising network ID of all of these websites. And we're finding these clusters of websites that appear to be independence of one another, but are actually run by the same entities because they're sharing the same kind of like, you know, ads inside these, just a small example. You know, going to all of these websites and finding their terms and conditions and their privacy policies and finding out who is the legal entity that is mentioned in these things and then trying to find those legal entities and specific systems of record. So when I say like understanding the data like I don't mean just having these metadata data sets which we already have, I mean something far more than that, which is like really trying to understand the reason why that paid was created in the first place. And what that person's intention when they created it was like is there incentive advertising revenue? Is there incentive to push a political agenda? Right?
right, is there incentive completely altruistic? And they just have this blog because they care about this thing that they're writing about, right? That's actually what I'm getting to. And we haven't fully solved that and I think it's a really, really difficult problem and it's gonna take us a lot of work to solve it. But that's the direction that we're on. And I think that the outcome of that is gonna be much better signals because we're gonna have a lens to look at this data that no one else has. - The web detective work goes on. - Yeah, that's an incredibly difficult problem to solve. But as you say, if you're not doing something that's really difficult, then there's not gonna be worth that much. And I think you're right, if you can have the ability to understand the motives and the wise behind certain behaviors and I think that's incredibly valid. I'm not aware of that even existing in any product I've come across in the alternative data market. Well, hopefully I haven't given someone an idea. (laughing) - Well, it's been an awesome conversation. I think there are many tidbits from this that will be valuable to listeners. So, Stuart, thank you for joining the podcast and yeah, we hope to have you back at some point in the future. - Yeah, thank you very much for having me and I've been listening to all of your guys' episodes so far and it's fantastic what you guys are doing. And I wish you guys all the best. - Thanks for your time. - Thanks for coming on. - Yeah, thank you. - Awesome. (upbeat music)
Podcast Summary
Key Points:
Stuart Reed emphasizes the limitations of complex models in quantitative finance, advocating for simple models with strong economic rationale over data quantity.
Key criteria for evaluating alternative data include coverage (breadth of securities), point-in-time integrity, and trustworthiness, with backtests being less valuable than thorough metadata.
Notable addresses web data reliability by cross-referencing sources, using web archives for verification, and detecting AI-generated or manipulated content to ensure temporal accuracy.
The approach prioritizes process-driven data quality and transparency, providing extensive metadata to help users assess data validity themselves.
Summary:
In this podcast interview, Stuart Reed, CEO of Notable, discusses his journey from quantitative finance to building infrastructure for real-time web data monitoring. He reflects on his quant experience, noting that using widely available data with complex models like deep learning often fails to generate alpha; instead, success comes from combining simple models with well-justified alternative data. Reed highlights critical factors for data evaluation: coverage across securities, point-in-time accuracy to avoid temporal biases, and overall trustworthiness.
He cautions against over-reliance on vendor backtests, advocating for detailed metadata and transparency. At Notable, his team crawls millions of web pages daily, employing techniques like cross-referencing, web archive verification, and detecting AI-generated content to ensure data integrity. This process helps financial institutions trust web-derived data for decision-making, focusing on robust, thesis-driven approaches over sheer data volume.
FAQs
The podcast explores how businesses monetize and sell data, featuring interviews with industry insiders to share their experiences and lessons learned.
Stuart Reed has a background in computer science and quantitative finance, including work in exotic derivatives and deep learning for market prediction. He is currently the CEO of Notable, which builds infrastructure for real-time web monitoring for financial institutions.
He learned that using the same data as everyone else rarely yields meaningful insights, regardless of the analytical methods. Success often comes from combining alternative data with simple models and a strong economic rationale.
Key factors include coverage (number of securities and historical depth), point-in-time integrity to avoid look-ahead bias, and having a clear economic thesis for why the data should matter.
Backtests can be misleading and easily overfit. Detailed metadata about data collection, coverage, and processing provides more trustworthy insights for evaluation.
They use techniques like cross-referencing with third-party sources, verifying against web archives, and detecting anachronisms (e.g., references to concepts before they existed) to assess point-in-time integrity and flag unreliable content.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.