What on Earth? Ask LGND AI: Nathaniel Manning with Scott Harley
26m 14s
In this podcast episode, Scott Hartley interviews Nat Manning, CEO and co-founder of Legends, about his journey in geospatial data and the founding of his company. Manning's career has focused on geo-data, from crisis mapping in Africa to roles at the White House and USAID, and later founding Kettle, a company using AI for climate risk insurance. This experience revealed the challenges of processing vast Earth observation data, leading to the creation of Legends. The company builds on open-source projects like Clay, a large Earth observation model, to develop a platform that applies transformer AI architectures to satellite imagery. Legends aims to create "geo-embeddings"—compressed, queryable representations of Earth data—enabling users to ask complex questions about the planet over space and time, such as identifying fire breaks in California. The goal is to make Earth data understandable to both humans and AI, partnering with data providers while exploring future innovations like in-space data processing. Legends positions itself as essential infrastructure for the growing field of geospatial AI, similar to how OpenAI serves language models.
[MUSIC] Everywhere, podcast network. [MUSIC] Hi and welcome to the Everywhere podcast. We're a global community of founders and operators who've come together to support the next generation of builders. So the premise of the podcast is just that. Founders interviewing other founders about the trials and tribulations of building a company. Hope you enjoy the episode. [MUSIC] Hi everybody. I'm Scott Hartley, co-founder and managing partner of Everywhere Ventures. I'm super excited to be here today with really long time friend of mine, Nat Manning. The thread through your whole career has all been in and around geo data all the way back to your role as CEO of Ushahidi, which was a crisis mapping application in East Africa to running open data under the Presidential Innovation Fellows Program at the White House for USAID. And being the Chief Data Officer at USAID to founding of kettle re-insurance using geo spatial data sets across time and geography to better manage risk around catastrophic events like wildfires, given what's been happening over the last few years. So your current role as the CEO and co-founder with our friend Dan Hammer of a company called Legends. Legend is really building Google Earth for the AI generation perplexity across all things geo data sets. Welcome to the podcast. Super excited to have you with us today. Thanks for having me Scott. Could you be here with you? Tell us a little bit more about Legends. I know that Dan Hammer, your co-founder who was running all things APIs and data extensibility for NASA, had a deep geo spatial background as you. You guys were working on this open data project called Clay. Walk us through the genesis of legend and what led you guys to this idea. I was working as he said in the last iteration built this company kettle. The thesis was that climate change is going to break property insurance markets and two, you could use machine learning or AI to be able to run it on satellite imagery or weather data and better predict risk. Wands that were being exacerbated by climate change like wildfires and hurricanes. Both those thesis ended up proven really true and worked really well and in that work we probably sent millions of dollars training CNNs on questions like could we find all the fire breaks in California and then many other ones and then that created ensemble model. And each one of these CNNs you like feed at tens of thousands of images to train and over one query like that. And it worked really well but then this little thing happened to actually came out 2022 and we were in it and that company is sold insurance and it went well and eventually passed the baton because the next phase is really about scaling an insurance company as a technologist that felt and you like a real insurance head is what made the most sense for the next phase. Now I was sitting back just thinking about this technology and what had changed since chat cheaply came out and at the same time then I went and caught up with my boyfriend Dan Bruno is our third co-founder. Dan and Bruno had been building this fully open source model called clay made the same thing right after chat TPD came out they said hey could you apply this transformer model architecture which is the intention is all you need it's the innovation behind everything that's become LLM's and LLM stands for large language model instead of a large language model trained on language could we train it on a large earth observation model. Clay is a large earth observation model fully open source and open weights all done on a nonprofit for the purpose of climate environment would donate compute it's a GitHub project it's out there they'd been working on that then we are all catching up and realizing those two experiences with the two threads of the DNA helix that became legend because people started saying hey Clay is cool we put this to work but most people don't know how or want to go grab a model and do everything for themselves and in our space over here in earth observation. Earth observation geo there's no open AI enterprise tool there's no chain yet there's no infrastructure to put all this to work it's like 2020 and LLM went I became very clear to us that there was a need to build this bridge between these emerging models being trained on earth observation data not language and being able to put them to work easily that was the formation of the company the real goal what we're trying to do is being able to create a tool that lets you query the earth over space in time. Quirying the earth over space in time it reminds me of when I had the chance at Fika ventures offside a couple years ago to meet one of the co founders of SpaceX I asked him what problem are you solving and he said in the simplest terms gravity and I thought that is the most still clear definition I've ever heard. But I'll go about querying the earth data over space in time has a Google ask clarity to it or a space X S clarity to it which I love and it's a huge vision. I remember we had lunch at Tarteen in San Francisco and you said to me I knew that your background had been in and around crisis mapping and a lot of emerging markets and we shared that passion having spent a lot I mean you said to me I'm thinking about going into insurance. It was a mic drop I didn't quite understand the questions that you had or the foresight that you had thinking about all of the things that you learned in around mapping in in around these emerging market data sets in around USA and open data and thinking ahead to these applications for risk modeling and for insurance you know Steve jobs speech about staying hungry staying foolish. He also said in that speech that the dots don't make sense going forward they only make sense going backwards as you guys were building cattle and you were looking at how do you get more granular risk assessment to better model and better under right risk and a wildfire edge case it should be way more granular than a zip code it should be based on which side of the hillside are you on what are the wind patterns where the fire breaks all these earth level data sets it makes sense to me as you explain that that you guys really. Discover this market for legends in some ways out of the problems you encountered with data modeling for kettle is that right yeah I think it's certainly open the door to it I was in a very applied space for the area highly vertical eyes like kettle. >> It doesn't sell technology but we put all of our own technology ground up of user risk with legend allows for what is usually done with these images of the earth I say earth observation I mean satellite imagery for the most part but it also could be lower flying planes it could be drone imagery but it's basically the top down view just think about it that way most of the time when folks are using this observation what you're doing is you're taking these complex. >> These complex pictures and trying to translate them into information information is often semantic right that's what we understand is humans or a grid type format but us humans are really good at doing that we're good at looking at a picture and saying okay in this picture there's a bunch of properties there's a big old eight lane highway and is a forest on the other side and saying oh that eight lane highway acts as a fire break between that thing that's very. >> Burnable and those homes that have property so that's what a traditional underwriter would do that's what an analyst would do that stage one of how you create knowledge and information out of pictures because ultimately these are pictures and phase two was what I talked about earlier where you have a very expensive couple hundred thousand dollar. >> And this is where most of earth observation as an industry has been for the last bit of time where you would try to clean knowledge through doing a very specific trained CNN on one query now take a couple data scientists and an ML ops person and if the structure to be able to do that people talked a lot about counting cars and walmarts and all sorts of other examples this has been done a ton of agriculture it's done and used in carbon accounting it's done and used in government defense work. >> Government defense work but it's still the same thing you're ultimately it trying to create a data set out of these pictures and then what has just happened in applying all the technology that's behind LLM's and into this data set is the ability to do that orders of magnitude faster because you have a pre train model and that's what this pre train model is doing it's able to do the same thing of saying that eight lane highway is what we think of as a fire break. >> And label it and organize it and create that insight build a data set in like milliseconds from a query by saying can you create me a map of all the fire breaks in California that doesn't exist you can't ask chat tp t today can you create me a map of all the fire breaks in California. >> But the answer to that question does exist it exists in pixels in open data that has been available back to what we used to do you know that dots looking backwards only makes sense that's the part that always stuck with me from jobs is speech Dan Brunner I all met back in 2012 14 timeframe Bruno was the chief scientist in that box been I were both working in opening up government data and making it more available that's the point is because this data that NASA or European say she is he is being. >> And making available and then help build some of those APIs and put them out there listed with the fire breaks one the answer to that question exists but because chat cheap beauty or any LM is trained on language the answer that question doesn't exist in language today exists in the pixels that's a were able to bring to light and we think that there's a lot of value and that's been really untapped because it has been such a hard stack to work with it's a complicated set of data I mean to put in perspective I think this amount of data is not a good thing in the world. >> And it's a data I mean to put in perspective I think this amount of data doesn't necessarily equal value but it does equal complexity and maybe some bit of potential value.
All the language in the world, like what all these elements are trained on is sub one petabyte I think Dolly and all the clip image generations are three to five petabytes of data of pictures that have been out there that people trained on and We've collected something like a hundred to two hundred petabytes of Earth imagery But I'm like 200 times more volume which both makes it extremely heavy difficult stuff to work with versus semantics I think that's both why there's so much potential Just the fact over the last decade plus in thinking of what's been transpiring with SpaceX and in bringing the cost per kg Down asymptotically to zero if you think of that It's almost building the railroad to space. It's building the railroad west instead of a West young man It's go up and we're starting to put more and more things into orbit Google Earth was created out of keyhole There were a number of planet labs doing early optical imagery I saw recently Spire Global which runs a lot of different forms of geodata around ADSB which is flight tracking AIS which is maritime tracking GPS R.O. which is a lot of weather data Starting to integrate data sets starting to think about ADSB So flight tracking data Interface with FAA calls and calls about turbulence and mapping and creating context on top of Weather data and on top of flight data to figure out where are actual areas of the earth that are high Probability of turbulence and that's the point of layering context on top of pixels layering context on top of raw data In our portfolio we have a company called Satine out of Poland Which is doing this specifically around asset tagging mostly around military assets But maritime and military to be able to take a synthetic aperture radar image S.A.R. image that can be taken through cloud cover and through darkness and addition to optical imagery Which can only happen in that 25% of the time when it's not cloudy and daytime But taking that sort of data and then being able to tag it with maritime data tag it with ship names tag it with information like that There are a lot of these vertical specific Companies being built around data tagging within specific domains. Are you guys partnering with some of those players out there? Do you see all these raw data feeds as piping into legend or how do you think of this as this ecosystem evolves Starting with the price per kg going down to zero more things going in the space more data being collected in Leo Lower its orbit. How do you think of not just the data sets that exist? This is the tip of the iceberg as this set of data explodes even further as those costs go down But how do you think about that from a market standpoint? Our thesis is this which is that all large data sets are going to have transformer model architecture applied to them We've seen that happen in language. We know the winners there We're seeing it happen in self-driving. We know the winners there We see it happen in image generation and more recent like audio the live kits and chatterboxes out there The largest actual data set of all of those from as I just said there isn't a synonymous name with that for AI today And that's what we aim to be so what happens when you apply a transformer model architecture to data is the output is an embedding and an embedding is in short It's a vector of a string of numbers It's essentially compression and I talk about as a metaphor of like it's like a fingerprint if fingerprint can tell you who somebody is It has super low data But that fingerprint is to one person and all of the information you might have about that person It's a giant compression of identity in one simple thing and that's what an embedding does for language There's more to it. It talks about how it works in context with other things but for us The thesis is that for 20 to 25 years keywords were the primary data source for what I think of as a first order data object for Organizing and making sense of language about the same amount of time Mac tiles or raster data was the first order data object for making sense of earth observation In the last 30 months that has changed from keywords to Language embeddings that is now the first order data object for organizing semantic knowledge We just think that the same thing is going to happen to the earth observation to geospase We're going to transfer from peer pictures and Mac tiles to geo embeddings for all the same reasons We're trying to help us share that in In the article tech crunch put out about us talked about us as the standard oil It's ironic because we all come from climate backgrounds and environmental backgrounds But what was specifically meant there is with standard oils it really well was refine crude Outputs, you know raw or and refine it into something trusted and usable and repeatable That's essentially what legend is aiming to do is compress and refine all of this data that's out there We love all of the satellite companies public and private or drone companies or who are creating this content and imagery content Material and we end up partnering with them and be helpful because hopefully we can help build the on ramps to using them That much more easily and make it speak the language of AI One of the taglines I like to talk about for us is we're trying to use AI to make earth understandable We're also trying to make earth understandable to AI and that's what we do It's really fascinating to think about making things both human readable but also machine readable It seems like there have been a few attempts over the years to compress as you say Observation data around earth or location pinning to specific things one copy that comes to mind from many years back was Trying to paint locations on the map with three specific words, but that was a somewhat human interface to get to Hyper-specific mapping locations, but there's been a number of attempts to consolidate over the years I love that metaphor of a fingerprint and in many ways the world Always the cordians in and out and there's expansion and compression and it seems like we've gone through a boom of expansion Where this level of data acquisition this level has gone so far to the point of using you know 200 petabytes worth of imagery data And there has to be some Means of consolidation to make it accessible extensible able to build on top of and I think that's you guys's backgrounds And the ends in particular of doing this with NASA data sets that becomes really interesting as a new ground zero One follow-on question on this would be the data that's collected from space The subset that's even downloaded or on earth is less than what's actually collected in space because of the compression challenges Because of overflight times with ground stations and limitations on bandwidth There's a number of companies being built to expand bandwidth to enable more downloading of data from space There's also companies like star cloud and our portfolio that aim to do compression and data storage in space And be able to do compute in very low earth orbit rather than downloading to earth As you think about the frontier of where the world goes over these next few years Are you guys thinking about some of those really far-flung ideas like Data centers in space the expansion of more earth observation data going even more bananas than it already has over the last decade I would love to talk to that other portfolio company. Yes We have and it has come up with a couple of the satellite companies out there. It's like yeah, you could run Legend we're not building but you could run it on site on prem on a satellite and send embeddings down Which would be orders of magnitude less heavy Our aim with these models and things is for them to have minuscule data loss right minuscule information loss That would work tremendously well It has come up a couple times that is absolutely a use case that we haven't talked publicly about yet So that's pretty fun The other one I was remembering you were saying earlier about weather data and things too So I talked about our standard oil for these heavy objects into Refined embeddings and that's part of our thesis that embeddings is all going to go in one direction and that We're going to be the shop for building the top-down view of the world into embeddings The second is to talk about the weather data and other things is we're also very much experimenting with how does that interact with other data sources We're not going to try to build our own weather foundation model people done incredible jobs doing that It's not our competitive advantage But what's so cool in the world right now one example would be MCP servers and things letting you access all this data Like how these integrating all this API so we are absolutely experimenting with the other side I would say the front end and being able to integrate other insights as well and back to the fire break example The aim right is we'll say hey help me find all the fire breaks And you have an interface that helps you instinctually know that you build that data set But then you could ask it okay weed out anyone that's had rain in the last three months Or color and blue anything that you think is going to get rain in the next month And then with points of interest data you could say call out any of them that are Under 10 miles from the fire station. I'm still obviously thinking like a wildfire underwriter But it's like some of the stuff that it's all out there I have these incredible conversations with LLMs that can get you a similar experience Asking about all this wild stuff that's Known in language quantum physics or how some of these models work because they're written up But the answer to those three questions there are not in language in the same way and that's because it's pretty cool Are there human data categories in the loop that are helping stitch together these pixels to linguistic applications like you talk about rain impacting something or fire break these words have context have meaning and have an application Or a scope that applies to pixels, but that scope can't be known unless that data is first tagged and Incorporated into the model as you guys build what is really more perplexity for geodata rather than open AI for TO data, as you mentioned.
because it's an open set of ability to query across multiple models and multiple things under the hood rather than one closed ecosystem. But how does data tagging get into this and humans in the loop? One of the reasons why all this works now is because you can build off of all of the tagging that's happened in the last 20 years. So the models get trained off of that. And then what our product allows you to do is do that fine tuning for the rest of the tagging at the end. Let's stick with the same example for all the the lat longs of all the fire breaks in the state. And let's say I want to turn up and you go, "Mmm, really through a direct interface being able to click yes and no, fine tuning the data set at the end." And it's an easy to do tool versus what currently happens, which is like a pretty highly skilled data scientist and ML engineer doing that fine tuning of the model under that CNN example. But here an analyst who's trying to find this answer would be able to do it through an easy interface. And that fine tuning creates labels. A lot of what we're doing here with the models is creating data. If you believe that 99% of the value of what we know about the earth has already been collected is basically where's the Starbucks and how to drive there, which is a problem that has been solved. I give all the credit in the world to Google Maps and Keyhole team who help back us, which is why they're tremendous. Who solved that problem? And then Mapbox came along as well and did it outside of the Googleverse who I also think the world of. But that is points of interest in driving and navigation. And if the belief is that is 99% of the dollar value of the physical world, then we're not going to be that successful. We don't believe that. We believe that there's a lot of value in knowing a lot more about the blue and green blobs of the world over space and time. The mission or engineers and we believe that in that adage you can't fix what you can't measure. And that's ultimately what we're trying to help fix. Shrifting gears just in the last couple minutes of the podcast here, you also run kindergarten ventures, which she started, which enables you to see and talk to a number of founders like yourself and invest in around these themes. And as the CEO yourself over multiple companies, you talk a lot about leadership and how leadership is setting vision and removing obstacles. And those things that sound so simple, but are so actually hard to do in real life as you evaluate CEOs, you're going to invest in or as you think about your own journey. Any takeaways or any learnings from how you do that well. And there's many ways to run a company, first off, for a lot of different approaches and it can work in many different ways. What I've learned is the ability to build a world class team. I ultimately think this is a team sport and then it's the ability to look at how that team works together and diagnose and fix bottlenecks. People want structure, everyone's like, "Oh, I do not want TPS reports. No one wants TPS reports, right? We don't want bureaucracy." But people want structure. They want to know where they're going and how to get there and how they best contribute towards direction. That's a balance that you constantly have to be balancing. And then you need to be constantly trying to fare where the bottlenecks is and fixes it. It's a constant plugging of holes game and that has to be fun. And that is a question of team efficiency, velocity, talent density, and working as a soccer team on the match. We're a soccer team, we're not a track and field team. One person can't win a race and it's a win. And if the team loses, it doesn't work that way. I think CEO's job is to keep that in mind and be operating in that way. And ultimately I think as an investor back to that quote, "Where the dots looking backwards make sense to the next step. They're obsessed with this thing." Amazing. Such a rich set of experiences that you've had through those dots looking backwards. And thankful that many of those dots from the Presidential Innovation Fellows program at the White House in D.C. And to overlapping Kenya together, to investing in Kettle with you, to investing in Legend with you, but had four or five touch points with you over the last 15 or 20 years, which has been a true joy in my life. Me too Scott. It's been really fun. We've gotten into a lot of fun stuff over here. Scott is amazing and has been an incredible resource and friend and advisor all sorts of sundowns. Looking forward to the next decade plus. Finally, working listeners find you online. Legend is LGND.io. For me, it's Nathaniel Manning on LinkedIn or NAT, N-A-T-P Manning on X, Twitter. I still call Twitter. That's where I am most of the time. Those two places. Any book or podcast that you're currently listening to or recommend for us? On the podcast train, kindergarten adventures is me and David Rosenfeld from acquired biggest fan here. I thought the Indian Premier League was just tremendous. It was like such a story I didn't know. Such a hero is journey. So that was fascinating. I loved that. Book-wise, so many good books. Honestly, as a leader, there's a psychology practice called "Parts Work" or "Internal Family Systems" by Schwartz. And I was reading one of his books recently. I thought it was really revelatory. And I think kind of important as a leader as well. You realize it's often in our language. This part of him is really upset about this. This part is affected by this. It's good to know and be able to recognize that. I think it's a helpful self-improvement and self-awareness as we all try to cost everything improving over time. Thank you for sharing that. We'll include the links in the transcript of the podcast. NAT, thank you so much for joining us. Thanks for your time. Thanks for everything you're doing with Legend. And we're super stoked to be small ambassadors and a little part of the journey with you. Of course, one of the first people I call. Thanks so much. Thanks that. Thanks for joining us and hope you enjoyed today's episode. For those of you listening, you might also be interested to learn more about Everywhere. We're a first-check pre-seed fund that does exactly that. Invest Everywhere. We're a community of 500 founders and operators and we've invested in over 250 companies around the globe. Find us at our website everywhere.vc, on LinkedIn, and through our regular Founders Spotlights on Substack. Be sure to subscribe and we'll catch you on the next episode. [Music]
Podcast Summary
Key Points:
The podcast features Nat Manning, CEO and co-founder of Legends, discussing his career in geo-data, from crisis mapping to founding a company that applies AI to Earth observation.
Legends aims to be "Google Earth for the AI generation," using transformer models to create "geo-embeddings" from satellite imagery, making complex Earth data queryable and understandable.
The company emerged from experiences at Manning's previous venture, Kettle, which used AI for climate risk insurance, highlighting the need for better tools to process vast geospatial data.
Legends seeks to compress and refine petabytes of Earth imagery into usable insights, partnering with data providers to make geospatial information accessible for applications like disaster risk assessment.
The vision includes integrating various data sources (e.g., weather, maritime) and exploring future possibilities like on-satellite data processing to reduce bandwidth constraints.
Summary:
In this podcast episode, Scott Hartley interviews Nat Manning, CEO and co-founder of Legends, about his journey in geospatial data and the founding of his company. Manning's career has focused on geo-data, from crisis mapping in Africa to roles at the White House and USAID, and later founding Kettle, a company using AI for climate risk insurance. This experience revealed the challenges of processing vast Earth observation data, leading to the creation of Legends.
The company builds on open-source projects like Clay, a large Earth observation model, to develop a platform that applies transformer AI architectures to satellite imagery. Legends aims to create "geo-embeddings"—compressed, queryable representations of Earth data—enabling users to ask complex questions about the planet over space and time, such as identifying fire breaks in California. The goal is to make Earth data understandable to both humans and AI, partnering with data providers while exploring future innovations like in-space data processing.
Legends positions itself as essential infrastructure for the growing field of geospatial AI, similar to how OpenAI serves language models.
FAQs
Legends is building a platform that functions like 'Google Earth for the AI generation,' using transformer models to query Earth observation data over space and time. The goal is to make Earth understandable to both humans and AI by converting complex geospatial data into accessible, actionable insights.
The idea emerged from the founders' experiences with previous projects like Kettle, which used AI for climate risk modeling, and Clay, an open-source Earth observation model. They identified a gap in infrastructure for applying AI to geospatial data, leading to the creation of Legends as a bridge between models and practical applications.
An embedding is a compressed vector representation of data, similar to a fingerprint, that captures essential information from large datasets like satellite imagery. Legends aims to transform Earth observation data from raw pixels into geo-embeddings, making it easier to query and analyze.
Legends focuses on applying transformer model architectures to Earth observation data, creating embeddings that simplify complex geospatial analysis. Unlike vertical-specific companies, it aims to be a foundational tool that refines raw data into usable insights, partnering with data providers to enhance accessibility.
Applications include risk assessment for insurance (e.g., identifying fire breaks), environmental monitoring, agriculture, carbon accounting, and defense. The platform allows users to query Earth data intuitively, such as mapping fire breaks or integrating weather data for predictive insights.
Legends addresses data complexity by using AI to compress and refine petabytes of imagery into embeddings, reducing the heavy computational burden. This approach enables faster, more efficient analysis compared to traditional methods like training individual convolutional neural networks.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.