The discussion centers on web scraping's misunderstood role and the origin of Browserless, founded by Joel Griffith. Initially a jazz musician, Griffith transitioned to tech due to personal circumstances and a freelance web project. His "aha moment" came while building a wishlist app that required parsing JavaScript-heavy sites, revealing a gap in scalable browser automation tools. This led to Browserless, a bootstrapped company addressing the complexities of running headless browsers in the cloud, such as system dependencies, memory leaks, and maintenance burdens. Use cases for the service range from data scraping and testing to asset generation and security audits, often with innovative applications. The conversation also touches on the ethical landscape of web scraping, emphasizing a balance between open information access and respecting data ownership, with Browserless positioning itself as a tool for legitimate automation needs while navigating technical and legal challenges.
this thought around web scraping. Every time I say web scraping to most developers, you can kind of see them curl up a little bit like, "Oh, don't say that around me." You know, it's kind of a dark part of the internet, but it serves a big purpose. And I think bringing that to light a little bit more would do it well. (soft music) - Hello, welcome to Deb Tools FM. This is a podcast about the developer tools and the people who make 'em, I'm Andrew and this is my co-host, Justin. - Hey everyone, we're really excited to have Joel Griffith on today. So Joel, you are the founder and CEO of Browserless, which is a pretty cool company. I've actually used Browserless a few times, so it's like fun to have you on. Browser-on-mation kind of sucks and y'all do a good job at it. So, you know, browse through that. (laughs) We're excited to chat more about what you're doing today, but before we dig into that, we'd love to just hear a little bit more about you. So, would you like to give us an intro? - Yeah, I'll give you the TLDR hopefully. Thanks for the intro and the kind words. I'm really happy and excited to be on the podcast. So, yet Joel, founder and CEO of Browserless, I/O, wasn't an engineer when I started 12-ish 13 years ago and then became an engineer, you know, wanted to solve my own problems a little bit with websites and whatnot for bands and stuff I was playing in. And then similar sort of thing happened with Browserless to be honest, I was working on a few things and I was like, "Hey, this really sucks." It could be better and then kind of had an aha moment when I was building something else at the same time and kind of reran into another web automation problem. I was like, "Why is this still suck?" So, yeah, kind of did a pivot then and, you know, Browserless and it was a good fit, you know, as an kind of an engineer myself and knowing the language and the people and, you know, where they hang out and all those things definitely made it a lot easier to kind of like, you know, find first users, having easier time talking to customers. I think especially in tech, you get used to this like weird metaphorizing of like technical problems such that, you know, laymen or lay people can understand what you're saying. And I don't have to do that, which is great. It's not my best skill, I try, but anyways. So yeah, that's the quick, very, very quick about me and yeah, we could go any direction from there to be honest. So yeah. - So before you were a developer, you were a jazz trumpet player. That's pretty out there. What, like, how did that switch happen? It seemed like a seismic shift, almost. - It is. Yeah, there was, so there was two things. One of them is kind of sad and one of them is just kind of like opportunistic. So at the time as a jazz trumpet player was about to have my first kid and, you know, we go in 20 weeks for this ultrasound and realize there was some like abnormalities, you know, and they weren't good ones to be quite honest. And so we kind of, my wife and I were looking at this long, long list of things that we needed to do. And I was like, I don't know, as a musician, I can, we can afford those things that might need to come up. And so it was somewhat out of, you know, external pressures or external things that had happened that kind of like, I don't say forced me into technology, but like, strongly encouraged me to get back into technology. I always loved hacking on the side, you know, even like high school was writing, programming, calculator functions for friends and selling those for like, dollar or two or so or something. So there was that kind of like tray I had for a long time. And then, you know, this happened with, you know, a kid. And so that really kind of was like, okay, I really should see, you know, what I'm capable of here. And if I can make it happen. So that was the big function. And then the second thing was the store that I was working at Music Store. Actually, you can go look it up on my LinkedIn profile. You can see, I was director of internet services at a Music Store. They were getting pitched to website and they were getting charged a lot for a site. And I remember looking at this with the owner at the time and I was like, man, if you could give me like a couple days, a couple of weeks, I think I could probably build this, you know, if I'm allowed to on company hours. So it was kind of those two things kind of fusing together at the same time. The timing was fortunate. And yeah, so did that. This was 2012. So we didn't have Flexbox. We didn't have much no AI. So it was a lot of like, you know, the Peter Griffin gift of him like trying to get the blinds to go on the right way with, you know, horizontally and vertically centering things in CS. That was like my life for the first five years in a nutshell. So yeah, but I loved it. It was fun. I came to it, you know, every day after work, I was still like, well, let's do some more. Like, let's figure this other thing out. So, you know, having that, I guess, part of you is really, really, really important in tech. Even with AI, it's still, I think it really, you have to be kind of like a tinker or a hard or curious by heart to get into. So yeah, anyways, that's, that's the longer form version of how I got in the technology. Yeah, I wouldn't wish block or table layout on my worst enemy. Do you remember having to like float left and then clear both, I think was the hack? Even nowadays, I'll start writing a class and it's like clear. And now I got to stop that. Joel, we've got Flex, just use Flexbox. IE6, IE8, support, awful, awful. Anyway, moving on to brighter topics. So you've done a lot since then and it props you for that. So you've got, you've got browsers now, which is a boot strap company, which I think is incredible. That's the hard mode for building a company. Good on you for taking that on. Can you kind of walk us through like, you'd mention your aha moment earlier. What was that? Like, what was the inspiration? It's like, ah, I need to do browser automation. Yeah. So I was working on a site that lets you like create a wish list for birthdays or anniversaries or whatever. I'm sure most people on listening will understand like, hey, it's your birthday. What do you want? Like, oh, let me go to Amazon or whatever and create a wish list. But it was like, I don't want you stuff from Amazon. I don't love Amazon. Like they're great and everything. But like I'd rather like venture out, you know, to other things too. So anyways, I was building an app that kind of just did that. So you could just paste a link into something that you wanted. And it would, you know, grab price pictures, made data, all those things. And then at the time, it was target.com was running as like a spa single page app. And so none of that made a data was available unless you parsed and ran JavaScript to get that data. And so this was back in like 2018, 2017, 2017. And so I was honestly at the right place in time, ran into this problem. I think what was it that was before like puppeteering headless chrome? There was a phantom, phantom JS just like that maintainer has said, hey, see I'm done. Goodbye. You know, Google's got it this now. And so right time and place, probably a little early to be honest. So the first few years were pretty slow growth. I mean, I ran the company by myself for the first four years, roughly, and worked full time. So yeah, it is kind of hard mode. You either got to like subsidize your living with either ramen profitability is, you know, bootstrappers like to say, or you know, you work at another job and fund your living with that. So yeah, kind of worked two jobs for a long time until it got to a place where it was, you know, predictable enough, profitable enough that I could, you know, do my own thing. But yeah, I think that wish list was like the turning point, you know, thinking about if I ran into this, other people are going to run into this. And then I was even looking at like puppeteer issues, like sorted by the most commented. And the first one was, how do I get this to run in Linux? And I think that was actually the aha moment was like, oh, man, if this is the number one pain point, maybe I could do that and start a service that, you know, that's all it does is just browsers as a service. So yeah, that was the, that was the start and, you know, the big realization. So, and then, you know, looking back now on eightish years, like a lot of other things just kind of happened that I wasn't even aware of or, you know, working towards it just happened organically a little bit like love developers, kind of like how I started the top of this was like, I spoke the language. I knew they were. I knew their struggles. I knew what it was like to be in the trenches with them. So that helped a lot, just kind of like knowing what the struggles are and, you know, speaking the jargon and all that. >> So for devs who might not have like tried to run browsers, like in the cloud at scale, what problems are there? Why can't I just string together the tools that are out there now and do something similar? >> Oh, man, where to begin? There's a lot. Like out of the box, it's gotten a lot better out of the box. Like you could grab play right now. It even has like a sample Docker file and that'll get you like 90% of the way there. There were some funky things a few years ago where it was like you ran the Docker file and in line the code that you wanted to run. So there's just kind of like some weird API, you know, boundaries. You know, my mental model for like how this was going to work was treating it almost like a hosted database in a sense like, you know, you get an URL, username password, you connect it, but the database doesn't do everything for you. Like you still need to tell like, hey, I want this structure grows. Those sorts of things. And even today, like the, you know, the playwright Docker files, other ones don't feel like that to me. They're still kind of like very opinionated on how you use that Docker file and I wanted it to be a little more open and, you know, generally available and just kind of have some like best practices and cut it into it. But so that was the biggest problem, I think. And then, you know, all sorts of other things like running into system fonts, other system packages you might need for this weird part of Chrome to Work or WebGL or
or whatever it is. There's just a lot of like gotchas, a lot of sharp edges that you run into having done it for a few years. And if, yeah, like over a weekend, sure, you can get up and running perfect, but then you're gonna be like constantly maintaining that baby and it going for the next, I don't know, I'd say six months probably, until you've kind of like, okay, this is not fun to deal with anymore. I should not be thinking about this problem, but, yeah, that's just a start, to be honest. And then there's new binaries that come out all the time and so you're constantly having to update that. And do you have a CI to test that, the new binary runs the same way that the prior one did and we do, we have several hundred specs now that run on every binary change. So, yeah, I guess it's like, how much of this problem do you want to own or do you want to outsource? So that's kind of like the ultimate tradeoff, I think is deciding I'm done with this, I'll just outsource it. That's kind of where we are, a natural fit and a good fit. - Yeah, I imagine there's like, all these other sort of like system properties, like browsers are very complicated. You know, you're talking about switching out binaries. I mean, I know that you've sort of written about like, JavaScript memory leaks and like other like challenges that you have at like running these at scale. And curious, can you talk like in more depth about some of like concrete challenges? So, you know, bit memory leaks or you know, process isolation, whatever, like what are the big meaty things that's like been hard to work through? - Yeah, it's honestly like, if I could put it into a sentence, the hardest parts are the things you don't easily find on the internet, you know, there's like weird combination of Chrome CLI flags that help it run better or in a more like user-ish type way, you know, and you know, they're really even Chrome CLI flags there isn't a centralized dock site for Chrome launch flags. There's like, I can't remember the guy's name Peter, something Peter Lou, it's like some developer tried at one point to do a description for all of them, but so many of them still have like, you know, comments. So a lot of it's trial and error. And like, you know, again, you'll get an update to play right. There'll be new, you know, Chrome binaries or whatever, and they break for whatever reason you don't know. And so then you gotta go figure out why that is. And usually it's like, oh, this right combination of node, you bun to and play right, there's a missing package that just isn't there yet. So you're like too far on the bleeding edge and you gotta like downgrade node or downgrade, you know, bun to or whatever the base OS is. So just like hundreds of those kinds of things really, as they just pile up and it's like, oh, this again, so you gotta go, you know, dig it back up. But honestly, that's one of my favorite parts. You know, I love taking something that maybe shouldn't work or, you know, isn't supposed to and then like coming through, you know, and being excited about that and having that, I'll have them like, oh yeah, I did get it, you know, just took a lot of trial and error. So those parts, but then, you know, that's the big picture. Node memory leaks are always a fun one. Like as an example, we had this interesting bug where we were splatting objects into an array and, you know, splatting from one JavaScript object into a new one and like probably creates a shallow clone and then there's a memory like there. And so it was like better not to do that in a like map or reduce function. It's better to just keep the original one and mutate it. So there's a bunch of stuff that you run into like that where it's like, this doesn't look good. It's not a good practice, but man does it perform well and I don't run out of memory when I'm doing it. So that was a pretty big one. It was like, okay, all the new fun array methods, sometimes they don't work out. They look good and they, you know, run well, but like the compilers or whatever, haven't caught up to, you know, handling those performantly. Yeah, that was a big one. A whole bunch of other, whole bunch of other worst stories too. We've had fun with engine X and trying to get it to do things sampling and, you know, I don't know. Yeah, it can go on and on and on. Yeah, lots of not fun issues to deal with there. Being on the bleeding edge has, it's pros and it's cons, but there's usually a lot of cons. It's funny 'cause now I get people come up and be like, "Oh, like you're so, you know, so much about these things." It's like, "No, man, I just like, somebody gave me a hammer "so I learned how to use a hammer and I just did a bunch "a hammering and have bruises all over it." And now I'm the expert at hammering is like, "No, I just, somebody gave it to me." I just had fun with it and now I guess I'm an expert or something. So yeah, that's how it goes. Yeah, so you made a pretty useful hammer and it's a pretty generic hammer as well. So like, what are like the types of use cases that you've seen people use browserless for? And like, have there been any that just like, completely surprised you? You're like, "Oh, I would never have thought of that thing." Yeah, there have been and there's ones that people still haven't used it for. And I'm like, "Man, that would actually be kind of a cool "product to build." So maybe I'll give one out here in a minute, but like the three big use cases we see, I mean, like data scraping, which is what I sort of had built it for initially, you know, automating third party systems that are like either antiquated, don't have a REST API, you know, those kinds of things, which seems to come up a lot in like third party logistics. Oddly enough, I never thought I would be in third party logistics, but here I am doing stuff with third party logistics and shipping ports and everything. Testing's a big one. And then, you know, like asset generation is what I call it, but like screenshots or PDFs or videos, like we had an ad tech company that was taking like HTML5, ads and converting those to video. And so they would like screen recorder stream those. Honestly, it's like every day, I'm like, "Oh, that's cool. "I don't know you could do that." Like that's actually kind of a smart idea or a good way of going about that. But the one, sorry that I kind of like lead it up with was, you know, there's been a big kind of a problem recently with, you know, security like NPM security and like package takeovers and, you know, things of that nature. And there's like a really good amount of like dev tools in the browser. Like you can, fire up a browser, go to a site and, you know, just by all the APIs that that has, you can look at the package chain a little bit more and see like, "Hey, you, you're running this version of, you know, whatever, "which had a takeover at some point and you should probably fix that." So there's like a lot of really cool security things I think you can solve with with Headless Chrome. And there's been a few, you know, companies that have done that and more of like a different way. But just like, you know, that's a company I think in my mind is like, loading up a website and just doing like a quick security posture snapshot and like, okay, you're using these versions of this library which are bad or, you know, you're not using SSL, HTTPS cookies or, you know, whatever it may be. And that's just, it's kind of low hanging for it almost, I think, because it's really easy to see that, you know, puppeteer playwrights. So yeah, all sorts of stuff, load testing websites to, you know, governments testing for like compliance on, you know, federal sites and it's, it's nuts, it's crazy. Every day it's something new, which honestly keeps you coming back. It's a lot of fun to see what people are building with it for sure. - Well, going beyond the use cases a little bit and just like to, you know, how the technology works. I think there's this interesting challenge you're probably facing today. As AI started kicking up and, you know, companies with LMS are crawling more and more sites, like folks started locking down their sites a lot more. Services like Cloudflare started offering, you know, bot protection solutions, scraping protection solutions. There's been a long time of like folks like, oh, you know, we don't want people like scraping our data or whatever. I worked for Food Network for a while. We had this problem with like recipe data. As like, oh, do we like want people scraping? Like do we try to block them or whatever? And I'm sure as you're building a product for this, you kind of go through that same thing. You're on the other side of it. It's like, we're enabling people to like get information. And then we want to allow that. But the companies are trying to fight us because you know, the data is their mode and they're like really holding on for your life. So how do you, how do you navigate that? Like what does that mean? - It depends as the answer. Yeah, I mean, I've done a lot of soul searching on this topic a little bit just personally. And you know, like I kind of go back to like the founding of the internet, you know, like when the internet came about, like the whole point of it was like free open information, you know, like it wasn't locked behind. Encyclopedias or libraries or books you had to buy, like it was accessible to just about everybody, you know, and it was out there for you to take into what you wanted. And so that's kind of like the motive, like I sort of like try to base assumptions on is like, you know, trying to lock portions of the internet down feels almost antithetical to the reason why the internet is the way it is. And so I definitely, I mean, obviously it runs browserless. So of course, I'm going to be more on the like, yeah, it should be free and open, accessible blah, blah, blah. AI does bring it particularly interesting spin now, because now it floats on from like, okay, if everything's open and free and accessible, like who owns the ideas? Is there an owner, you know, and it's the negative, you get even further out into the philosophical weeds, like, who are we? What is is, you know, what do we, you know, we come with nothing, we leave with nothing. So do we actually own things anymore? But, you know, anyways, so trying to bring it back a little bit. Everybody does it. Everybody does web scraping at some point or other. And it's just, you know, people say they do or they don't. And a lot of people just say they don't, but like, I can't think of a single business that at some point didn't do some kind of like, you know, competitor analysis, you know, obviously like e-commerce.
is big, you know, they want to make sure that their pricing is competitive. And so in order to do that, you got to see what other prices are on the internet. So it's just funny. Like people will pull that lever to get data, but they're like, oh, we were not going to put our data back into that pot. No, no, sir, we'll block all that stuff out. And it's like, well, you're saying one thing and doing another now. So like, which one is it, you know, which side of the fence are you on, I guess? But anyways, I'm all for free and open. And for the last kind of like example, I'll use for this, which is interesting. We're starting to see like certain, you know, federal bodies, mostly in the EU, where they actually do need to get through cloud flare on like fake or, you know, false consumer sites. I can't remember what country, but they have a whole program where they literally scrape bad actors on the internet. And those bad actors use cloud flare dust to bot block people check you them. So it's like, it does flip on its head at some point. It's like, okay, now you have like legal bodies that are trying to use this technology to get through to see if, you know, something is legitimate or not. So anyways, it's kind of a fun example to kind of contrast against because usually, yeah, it's always about like breaking through cloud flare or whatever. And you're not supposed to. Well, it's like, well, no, there's there's times where that case is actually legal and justified. So I don't, I don't think it's all totally black and white. So I guess the next question after that is like, do you think there's like a too far in that? Like, there's lots of legitimate use cases for the product. Do you view any use cases as like, ah, that we should be doing that. Yeah, I mean, intent is always, you know, an interesting one. So like, trying to suss out at the end of the day, like, what is it your after? You know, does that? Is it good? Is it bad in like a moral sense, illegal sense? You know, and so like, there is a bit of a KYC process for us just to kind of like know what you're up to if you're going to do bad things. Fun example was like one of our first high paying users, like when we started, this guy was trying to like swing the bets on, I think like a poker site. And so he was firing up bots to do it. And I was like, ah, no man, that's, that's not cool. I don't, I don't think that's good. Um, so yeah, no, you can't do that on the platform. I'm sorry. Go go find business elsewhere. Uh, and uh, best up luck, I guess. And in a sense. So, um, yeah, it really comes down to intent. And that's, you know, that's where it gets tricky and more nuance. You know, like, what's a good thing? What's a bad thing? Um, so yeah, we try to weigh each one in kind and then, you know, make sure that the lines to, you know, like our own moral and ethical code. So now like, now we've got to go figure out what that is too. So I don't know. A lot of nuance, you know, some of it's good. Some of it's bad. Not all of it is bad. I think the biggest takeaway, I would stress to people that are like kind of in this space or idiom for whatever reason. So, um, yeah, I think that, I don't know, to be honest, like, that's a big thing for me this year is to kind of like try to flip a little bit. The, um, it's kind of like this thought around web scraping. Every time I like say web scraping to most developers, you can kind of see them like curl up a little bit like, oh, that don't say that around me, you know, it's kind of a dark part of the internet, but it serves a big purpose. And I think, you know, kind of bringing that to light a little bit more would do it well. Yeah. And I think it's worth acknowledging that like, they're just like a lot of companies that don't have API. It's like, you have your data, your user, you've got stuff in there and you just can't access any of it, you know, be it from a technical narrative like the, the company or product like just doesn't have the capabilities to build the APIs or, you know, they don't want you getting your data because it's valuable to them or whatever the reason be. Like I can definitely see all like a wide range of reasons why this is like good, necessary. How use one example that my company uses is, um, we're a legal tech company and we watch the SEC marketing rules. It's like when they publish a new update, we like, you know, need to know, you know, and, um, you know, sometimes it's just like, we'll just hit the page and kind of scrape some of the data and get that. And I feel like that's a very like normal okay, like use case of that. And if they had better ways of us getting that data, we do it, you know, and like, just like sometimes, especially if it's government entities or whatever, you just, that's the only way you can save the day. It's only what you gotta do. What you gotta do. Yeah, they're not gonna move fast enough to have rest APIs and into your point, you know, like, it seems easy for engineers like, oh yeah, I'll just make a API to do that. But like there's so many things you gotta think about with that, you know, like, are you, you know, doing tokens the right way? Are you making sure that those are all secure? You know, it's the whole stack ready for it, you know, I don't know, there's just, you open yourself up to more potential security vulnerability things with those kinds of things. So I get it, you know, and it is funny. I have actually, there's been a bunch of new like legal e-startups and they're in it to save hours. They're trying to save like human effort, you know, by scraping these sites and, you know, yeah, it makes a lot of sense to me. But yeah, it's hard to sometimes justify it to the person depending on who you're talking to. For sure. I wanted to sort of like tie all this back to a point that you made earlier, which I thought was really interesting. You're saying you're like, you're thinking about browser listen initially is like, oh, well, I'm thinking about this as like a database service. And it's kind of cool that you have this feature called browser QL or like this kind of query language for scraping or like more like turning the internet into an API, which I like kind of like frame it as, which is really cool. I'd love to hear more about that. How did that come around? Like, was it look like kind of, yeah, this more detail there? I had the thought of this, you know, because that was our kind of like company, thematic was, you know, turning the internet into an API, you know, and then we tried to like make the other part of music, oh, let's make getting data on and off the internet is easy as like a SQL statement, right? Like, the other technologies, databases in particular, it's pretty easy to like, base upon like what you're asking, you kind of have an understanding of what you're going to get back. And you know, that didn't, doesn't kind of didn't exist until, you know, browser QL or BQL as we call it. And so that was the whole point, it was like one, I want to make it more structured and easy and like, it's based on GraphQL, which sorry, if I immediately lost half of your audience when I said that word. But the nice thing about GraphQL is like the shape of what you ask is what the shape of what comes back, you know, there's, it's strongly typed, there isn't a language or framework barrier. And so there was just a bunch of like nice things you get off the shelf with it that I was like, man, it feels like a pretty interesting thing. I wonder if I can mix, you know, headless automation with like a structured language. And that way we could offer it to, you know, people that have to run PHP or Scala or, you know, go that don't have either like a really nice language to run or nice framework to, you know, do this type of work with. And so it just kind of opens the doors for those folks, was a big, another big reason. But yeah, there's also just like a lot you have to, you know, puppeteer and playwright to their points are really good for what they're kind of focused on, which seems to be more and more towards testing. But there's a lot that, you know, those libraries could handle for, you know, users that they don't today. And it's because they take much more agnostic approach to it, you know, why you're using them. And so kind of having our own first class language was our way of like trying to do a little bit more and kind of getting out, but, you know, outcomes people wanted without having to be so kind of like overly dogmatic about, we're not going to do that because it's too opinionated or whatnot. So yeah, I think those are like the big benefits. And then like having a great editing experience is nice. Like if you play with our editor, it's like you write the statement, you run it, you see what the browser is doing, you got dev tools, all these things that tries to like help coach and recommend a little bit like what could be better, what couldn't be. Obviously it's a new tool, so it does take a little bit of a hurdle to get over to learn it. But I think it's worth a trade off because we do try to simplify a lot of things and not have to like make you think so much about like promises and, you know, race conditions like that that all's kind of handled for you, which is nice. So yeah, I love it. It's my favorite part of the product to be honest. I love like that's the first place I go when I need to like figure something out. So it's kind of nice to have it all in one place and docs and everything and you like one stops you don't have any like context which a bunch. So yeah, yeah, that's a lot. We could dive into any of any of those things. But that's yeah, the kind of the big TL very long did read part of it. Yeah, I find it super interesting. You guys created a whole language around this like an IDE like how far do you guys go? Do you have like a language server? Like how how how language is it? Oh man, I wish. No, thank thank God we didn't write Alexa or an apparcer. Though I'm kind of debating if we should there's some like bottlenecks, especially in the node implementation for like GraphQL JS just, you know, being single threaded and some of the I think some of the Lexi and the parsing is still ran on the main thread. I don't know if they do it in the background by they may have a sea library. Now that does that anyways. Neither here nor there. It is that stuff is thankfully off of our plate. Yeah, those are those are really challenging to solve sometimes. And I think you know, anytime you adopt a new technology, it's like how much do you want to like put in this to yourself? How much ownership do you want versus how much can you just get off for free? You know, if it meets your needs. So yeah, kind of another reason for going GraphQL is like their editors will componentize already and react. So it made it really easy to like build this remote viewer for the browser and the dev tools part was really easy because we could just swap out those components. So yeah, big props to like MEDA and their
dev team for like getting the fundamentals right and getting the building box right such that we can come and do this pretty easily which is amazing to me to be honest. Yeah, I think it's a really interesting use case for GraphQL. It's all a random PR for someone wanting to do or an RFD or something for like GraphQL and Chrome DevTools as a protocol. I don't know. It's something like that. I don't remember exact context. It's just like cropped up on Twitter and I was like, oh, that's interesting. Yeah, that is interesting. We'll see if that happens. Maybe relevant and useful. Yeah, I don't know. It's GraphQL somewhere. Some layers GraphQL all the way down. So yeah, I kind of want to talk about your Braselus 2.0. So you know, it's done a lot of work. I spent a long time kind of working on this like revamp of Braselus. It's been a few years now, I guess, since that point. And I'm curious like what sort of like lead up to that and like where you're all at today. Have you done another major version since then or are you still sort of like rolling on those two and like, I don't know what's going on. What's new in the ecosystem? Yeah, so 2.0 came. We wanted to have support for other browsers just besides Chrome and Chromium. So we do support WebKit and Firefox as well now. And supporting those made us rethink a lot about like how routing worked. And then like, you know, what would need to change to in order to support those two. And like, you know, making it work in the, you know, 1.0 branch would have been probably possible. But there was just 1.0 was so focused on Chrome and Chromium that it felt like odd to have these other two vendors as well in there and to be like, I just feel kind of like left handed and feel like sort of like intentional. That's what we wanted. So I was like, I'm not a huge fan of breaking changes. I feel like you should spend the time, you know, think about it a long time. The 2.0 release probably took me eight months to just like going back and forth on like, how do I want to route into work? Which of those look like file wise, you know, what should the system take care of for you and what should it not, you know? And then like, honestly, like building that up, using it and being like, nope, this isn't it. This isn't feel right. I'm going to go back and rethink that part of it again. And so, yeah, I mean, it's, it was, you know, the culmination of all the years of building the first product and like going back to the drawing board knowing that we have big breaking changes. We want to make like what other things do we want to do? So yeah, we, as one does, we built our own router. It supports WebSocket and REST at the same time, which was a big one that was hard to find and still kind of is even in like next JS. I think it still feels a little awkward or left handed to have a WebSocket server because you kind of like bind it to a global object if I remember right. And then it's just kind of like accessible via, you know, outside references to get access to that object. And it's like, no, that doesn't, none of that feels right. So, anyways, just seeing like all of that and then wanting to use TypeScript and really like leverage a lot of TypeScript. So we have this fun build process where, you know, routes can define what they accept, what they return and like query parameters to. And at build time, you know, we look at all the routes, we take all those definitions and we convert them into open API schemas. And then those are verified at runtime. This one took quite a while and actually that build file is kind of insane for that, but it gets us a lot of stuff for free. So like if we have a property on a post body that is like directly corresponds to maybe like a puppeteer API, any puppeteer documentation comments they have on that. That API actually get passed through all the way to our route definitions and in our open API build as well. So you can like go and see exactly like the puppeteer notes for that particular property, which I thought was like, oh, this is sweet man. Like, I got this for free. Like, I don't have to go and like, pain painfully like annotate like, what are, what is a cookie in, you know, puppeteer parlance? Like it just passes through for free, which is, which is awesome. And so yeah, anyways, all those things, you know, multi token work as well, like wanted to just do a bunch of different, you know, features and stuff, which was just really hard into it to get into 1.0. So 2.0 is like every birth of everything and like building on what made it successful, but also like keeping it more open for, you know, browsers and other, you know, better debx performance boosts. And yeah, honestly, I'm happy with it. I like, it's been like you said, I think a couple of years now, maybe year and a half. I don't really envision a 3.0 to be honest like, nice. Yeah, don't break chain, don't make breaking changes unless you really have to. I think we actually have a blog post about this coming out soon. It's like, yeah, like don't, don't break if you really don't need to. And if you do break all the things because you might as well just get it done, you know, when, when that event happens. So, yeah, fingers crossed, they'll be, they'll never be a 3.0. Is my, is my go to. So we'll see what happens though. A good goal. It's a, a lofty goal. It's surprising we work on a platform, the web that does that. Like it's kind of crazy to think that like they try to make no breaking changes forever. Like I can't even imagine as a live job. That is. Yeah, that is a good point. You bring up a very great point. Like the fact that you can run JavaScript from IE five days today and it still parses and executes is like, that is a miracle that that ever happened. So, yeah, very true, very true. Yeah, it kind of goes back to the data access stuff we were talking about. Like the web doesn't feel like a platform that can die as easily as some like iOS apps. It's like Apple can update the SDK and then eventually you have an app that can't run on anything but an emulator that you can't run unless you have an old computer. It's like multiple layers of like I can't do it anymore. Yeah, no, that, sorry, I had callbacks to running like the IE VMs from back in the day. Like virtual box, much of it. And if you've had to do that, it's like run the IE six VM and it's like, yeah, and then those stops working after a while Microsoft pulled pull the plug on those and it's like, well, shoot, now what am I going to do? I have to go to like a secondhand store to find a NEM three or whatever to run this operating system again. Like, yeah, it is crazy. It is crazy. Cool. So moving on to the next topic. We like to talk about open source and business models a lot on this podcast. You have a lot of open source out there. It's seemingly you can run your own browser list. So can you walk us through like how you think about open source and how you think about your product in relation to that? Yeah. So I, you know, I as an engineer myself, I've got tons and tons of value from open source. Like, you almost can't give back enough for the amount of, you know, benefit there is to open source. And so, you know, thinking about that and also, you know, like a personal sort of skill trait motive that I love is like showing, not telling necessarily. And so that was another reason for like trying to do open source. So like, I get it. People are in very different stages of their journey where it's like, I just need to do this thing on a personal project or whatever. It's like, great. Yeah, don't even, don't even, don't even talk to us. Like, just grab the doctor file and go go have a good day. You know, like, I don't want to have to like loop you into a bunch of sales calls and like do licensing and blah, blah, blah. So the motive here, you know, for me to do open source was like one to show people that like, yeah, you can just, this is how it, how it's done, how we do it. And so if that's all you need, awesome, I'm glad that we could have like helped you and hopefully, you know, someday you might need us or whatever. So great. We've had a good experience. The second part too is like, I was, I was never an engineer, you know, I never went to school for this. I kind of want to show people like, hey, you can do really cool things, like really important things. And maybe they're outside of your reach and like, use me as an example. And like, this is kind of like, I don't know, a way of showing a story or telling a story to that's like a little more at least grounded for engineers to like, you know, Joel, he, he went from never having like, opened up get or CLI to like doing this. So I think it's just another way to like illustrate what's possible for people. And especially now with AI too, it's like, it's even more accessible to just about any audience. You don't have to be an expert. I think they're probably still a reason to be an expert, but like, you can get a lot done and not know a lot, which is pretty cool. Obviously, be careful with that. But yeah, so those are kind of like the motives. So yeah, the open source, stalker image is free. You can download it. You can run it. It's never going to ask you for a license. It'll never bug you about it. It doesn't call back home. It'll log. I think like a couple of links. It's it. So you can feel good about it that way that you don't have to like drive, you know, go to a website, sign up for a license, keep put it in, you know, all those sorts of things. But yeah, so on the commercial side, it's like, yeah, if you're going to use this for like something proprietary, then that's when you want to contact us and get a license or, you know, it's hard to put that in like legal terms. There's always, I think we've tried two licenses at this point. I like the second one. Okay. But I think there's still some nuance. I would try to tweak, but you know, not a lawyer. It's really hard to get that out and like not piss off the wrong crowd, so to speak, because yeah, you know, at the end of the day, it's like obviously got bills to pay and stuff. So, you know, it helps you be successful in your journey with whatever we would love to partake in that success with you. So, yeah, that's kind of the overall, I guess motivation and thematicness around open source. But yeah, I'm actually kind of curious. Is it varied very much from folks that you've interviewed or has it been, it seems like a crazy, you know, to me, it still is messy. It's like kind of a problem to be honest. It's like funding open source projects doesn't really have like great answers. There's some, but they're like, you know, they're too far on one side of the spectrum or another. both of you are either
if you've seen changes in that. - Yeah, I mean, we've seen all sorts of things. I do think like you see more and more of the open core model, like collapsing a little bit. This has been very classic around like database companies, for example. So you use the MongoDB server side public license, which came about because their business was being cannibalized by other cloud providers, right? And of course, like, you know, the license drama around licenses continues to be a thing where people feel like they deserve. Yeah, and to any engineers listening to this, sorry for, you know, probably this couldn't doesn't come off wrong, but like, we are a privileged class and broadly can feel like we, you know, that if a maintainer decides to change a license from one to another for their personal, like benefit to make a project sustainable or whatever, people take that personally. You know, like, I mean, it is especially true if it's done by a company because, you know, people feel like it's a rug pull or whatever. And like, also like all understandable and empathize with it, but it is, it's fraught, you know? And I think the really interesting thing is, we haven't actually had anybody on the podcast to talk about this recently, but like, there's a lot of open source projects that are actually closing to contribution because like a bunch of like drive-by, vibe-coded PRs and like just massive amounts of like new contributions coming in and they're like, look, I'm not even gonna try anymore. Like open an issue, open a discussion on GitHub or something like that and then maybe we'll sort of like talk about it, but like, wow. I'm gonna close your PR, I'm gonna close your issue or whatever. And that's also the really interesting thing for me that I'm seeing a lot more of. - Wow, I did not. That is crazy. Yeah, we haven't had a ton like that necessarily, but we have had, you know, like punctuation changes in PRs, it's like kind of gaming, like how many PRs have you opened up in the last month a little bit, there's like a little bit of a gamification going on. So it's like white space changes, punctuation changes, like what the heck, fortunately for me, it hasn't been too much of a problem just yet, but man, that would be, yeah. I always think it's always just a good practice to open up an issue and talk about it like, hey, you know, this is what I'm seeing or would like to see, here's the strategy I'm gonna go down, thumbs up, thumbs down, you know, before I go and like waste my time building out this pull request that you're not gonna even review them. So yeah, I guess I'm not shocked really to hear that, to be honest, but kind of sucks to have to like close down all of that, you know, just because of the amount of abuse. I think the other thing I'll add, and I know that you've talked about this in other forums is that like, we're also seeing the sponsorship model is just like not scaling. They're very, very, very, very few, and usually only big names, like maybe like Tanner and maybe like Evan U of you, and like, you could probably count them on one or two hands, the number of people that are like really successful doing sponsorship, and they're also really, really freaking good. Like top, you know, top point, oh, one percent or something of like developers. So it's like, I think that's an incredibly hard thing to do well to get any sort of like actual value. So yeah, I agree. Even back in 2018 or whatever, when we was trying to think through some of these things and like go sponsorship route and do like, you know, consulting or whatever, like even then it wasn't on a good trajectory, it was my thought of the time. And to your point, it's like, yeah, you really had to be in the top point, 1.01 percent of, you know, authors out there to make it work. And then then you got to deal with like, consulting potentially to like subsidize what it doesn't do. Yeah, I don't know. I even thought about like, oh, it'd be cool to have like, a marketplace of like apps, like, Dockerized apps, and then you could charge, you know, I'm actually kind of surprised that Docker didn't do this, to be honest, because we needed at one point, like a way to host like, kind of like private images. And like, oh, yeah, like other users on the licenses from us, so like, why can't Docker just like have a paid model where, you know, like they're verified, they get those security stuff for whatever, and you don't pay for pull, or I don't know how the model would necessarily work, but like, it seems like there's a need. I know I not have it, of course, but I don't know, just kind of an interesting business tactic. I thought would be, you know, and it would help with other people that like, yeah, I built this really cool database engine, it's over here, you can pay for it blah, blah, blah, blah, on, you know, and like, yeah, I guess it depends on what part of the stack you're working software wise, but if it was like a fundamental layer, or like a service of some kind, and that route seems to, could have been a lot more sustainable, I think, going in the future. But we'll see. - Yeah, I've often thought the same thing of like, MPM, like there's this like kind of shadow ecosystem of like packages I can pay for, but like never know about. It's, you just kind of like, happen upon them, you're like, oh, wow, that here's a private registry thing that I could maybe do, but that it's not baked in, it seems like a obvious miss for developers and open source in general. - Yeah, I mean, just think about the amount of duplicative effort going on right now for like, just supply chain verification, you know, like anytime we get, you know, get through a company and we're going through like a vendor onboarding, it's like, they have to do that all over again, and it's like if there was just some like supplier or whatever out there that just said, okay, like, these are all verified, you don't have to think about it, here's like a certificate or something for it, and then like, just like, get you through a lot faster, it would be super nice because every business like has to do the same process over and over and over and over and over again for every new vendor and it just like, this could be consolidated in some way and mastery lined, but I don't know, if you're listening, go build that. I'll be your first customer, maybe. - So we said the magic words a little bit while ago, AI, it's changing the industry, really changing everything. I wouldn't have said that a year ago, it's kind of crazy for me to come full circle on this, but how has AI played into browser lists and like, what changes have you made to kind of accommodate this new world that we live in? - Yeah, it's affected us on two fronts. I mean, we're consumers and producers in a sense for AI and models, a lot of those things out there that need real time access to the internet need a web browser at the end of day to do that part. So we are, I guess a cog in the wheel, so to speak, for a lot of AI ambitions going forward. It does have a lot of same symptoms to like the data access, probably I know we kind of touched upon that a little bit earlier, but it's kind of the same questions I asked myself at the end of the day, like, you know, who owns these ideas? Do we own ideas? You know, who gets the dollar for it? And you know, those are big like bigger than me philosophical questions to a degree. But, you know, I think the same time Jeannie's out of the bottle, you know, this is like the internet again, but potentially even more so. And so like, it's open, it's out there. So, you know, what are you going to do about it or how are you going to react to it? And so, you know, kind of looking back on my history with, you know, getting into dev, getting into like business running, you know, I'd really stress to people that are like listening that are kind of scared about it, because you know, I've went through that like, depressed sorrow of like, shoot all my skills I've worked so hard for, kind of like replace them like, no, you still need those, you still need like that compass that internal like, I don't know, we called it like, your spider sense for bad code or whatever, you know, there's that like that thing that's inside of you that like, this is good, this is bad, like that is still very much needed, because you know, obviously you're not going to get perfect code from an AI as it stands today. But the amount of just like boring problems it can solve for you is just crazy. And so like, once you go through that value of depression with it, you come out and like, oh man, I can do so much more. I don't have to think about like wiring all these things up. And you know, as long as you're okay or good at reviewing code, like you'll get a lot of benefit out of it. So, you know, in terms of like the, the discipline of software engineering, I think it's still a tool. I don't think it's going to replace anybody soon. I think, you know, there's still going to need somebody to like, contextually understand things. And you know, there's a limit in my mind of like language, you know, you can only like articulate so much about how a system can work. And sometimes it's like, you just need to know a little bit and it's hard to put into words exactly like behaviors or whatever. So at the end of the day, I had never met a PM that just couldn't tell me in perfect terms to get a prompt to work. So I feel like, you know, your drives are still safe because you can like, hey, what about this? What about this? What about this? And, you know, AI's not going to replace that anytime soon. So yeah, anyways, people are doing really cool stuff with it. I think, you know, it does feel like early internet times to me a lot. So, you know, I think trying to leverage it is best you can and being part of the narrative of what's going to happen with it is more important than just sitting there frozen like, oh God, what are we going to turn in the matrix? You know, like, is everybody a machine? Am I married to a real person anymore? Like, no, this is a chance. Like you have the time, you can go do something with that, you know, for good or for bad or for nothing and just let happen to you. What happened? So, anyways, I always like having a little bit autonomy and like dictating like what happens and what doesn't. So I think that's kind of where I view is like, I'd rather make use of these tools, leverage them well, talk about them well, and like, good in a good sense how to use them too. So, yeah, I mean, it's crazy. It is a fast, it is fast pace. Every day I wake up and there's like a new N8N or a lovable or something like that that we're, we suddenly got to be a part of and I was like, oh, that's cool. How come nobody's always about this? And we're like, adaptor into their system. So I guess I better figure out how this works. So, yeah, it's fun. It's a lot of fun. I hope it continues to be fun and I hope it does.
if you're in a good direction societally, otherwise, we may all be robots someday, I don't know. - Yeah, so many unknowns, like I've thought about this so much. And like my brain goes back to like the printing press and like all these other innovations that kind of like expanded access to information, like the internet as well. Like, so many things that comes from it, but like the dark side, the bad things that we can't predict and how it'll change our society. So it's like, and the speed this time is just like unprecedented too. So it's like, it's a crazy time to be alive honestly. - Yeah, no, I think, yeah, the speed of which it goes, and like I just keep coming back, it's a tool. At the end of the day, you can use a tool for a good thing or a bad thing. So, you know, but like never using the tool and letting it, you know, it or somebody else do something with it, you know, I think is not good either. So I don't know. Yeah, use the tools for good things. I guess is the answer there. Okay, well that wraps it up for our questions. Thanks for coming on Joel. This is a fun conversation about everything browserless and all the data access and our lives in AI. It was a great time. So thanks for coming on. - Hey, appreciate it so much. Thanks for having me. It was great chatting with you both. And we'll see you on the flip side of the AI storm. - Yeah, thanks again Joel. It's really cool to see what you're building. I think this is a hard space. I have a lot of appreciation for it. I haven't had to do a few of these things in the past. And yeah, keep up the good work. Awesome, thanks.
Podcast Summary
Key Points:
Web scraping is often viewed negatively by developers but serves important purposes.
Joel Griffith founded Browserless after encountering web automation challenges while building a wishlist app, identifying a need for a scalable browser-as-a-service solution.
Running browsers at scale involves complex technical hurdles like system dependencies, memory leaks, and constant updates, which Browserless manages to simplify for users.
Common use cases for Browserless include data scraping, testing, asset generation (e.g., screenshots/PDFs), and security analysis, with applications often surprising in their creativity.
The ethical and practical challenges of web scraping are navigated by balancing open information principles with respect for data ownership and legal boundaries.
Summary:
The discussion centers on web scraping's misunderstood role and the origin of Browserless, founded by Joel Griffith. Initially a jazz musician, Griffith transitioned to tech due to personal circumstances and a freelance web project. His "aha moment" came while building a wishlist app that required parsing JavaScript-heavy sites, revealing a gap in scalable browser automation tools.
This led to Browserless, a bootstrapped company addressing the complexities of running headless browsers in the cloud, such as system dependencies, memory leaks, and maintenance burdens. Use cases for the service range from data scraping and testing to asset generation and security audits, often with innovative applications. The conversation also touches on the ethical landscape of web scraping, emphasizing a balance between open information access and respecting data ownership, with Browserless positioning itself as a tool for legitimate automation needs while navigating technical and legal challenges.
FAQs
Web scraping is the automated extraction of data from websites, often viewed with caution by developers. It serves a significant purpose by making information freely accessible, aligning with the internet's original goal of open knowledge.
Browserless is a browser automation service that simplifies running headless browsers in the cloud. It addresses challenges like setup complexity, maintenance, and system dependencies, allowing developers to outsource these tasks efficiently.
Common use cases include data scraping, automating third-party systems without APIs, testing, and generating assets like screenshots or PDFs. It's also used for security analysis, load testing, and compliance checks.
Challenges include managing system dependencies, handling memory leaks, updating binaries, and dealing with obscure Chrome flags. Maintenance requires constant troubleshooting and testing to ensure reliability.
Browserless emphasizes the internet's founding principle of open information. It navigates ethical concerns by focusing on legitimate uses like data accessibility while acknowledging the need to respect website terms and bot protections.
It was inspired by personal frustration with web automation tools while building a wishlist app. The 'aha moment' came from recognizing widespread pain points, such as running Puppeteer in Linux, leading to a browser-as-a-service solution.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.