Speaker 1You're listening to a Day One.fm show.
Speaker 2Can we start about Bridgewater? Nine years, the world's largest hedge fund. Why are you nice?
Speaker 3I don't think we've ever seen a company be so proud of losing control of its own software and having it do things that weren't intended. No, the police would be...
Speaker 4Public announcement. Yeah, the police would come over.
Speaker 3Exactly right. I mean, you've got something that's effectively locked in a box that you're making smarter and smarter and its incentives is to break out of the box. It does look good on a chart, Ben, it really does. It has a lot of pain and kicks and twists and turns underneath that chart.
Speaker 5You should have a lot more grey hair.
Speaker 3I think what AI is really changing is people with agency to go and create and build and do something different are going to be extraordinary beneficiaries. It's like pretty hard to attract great people into a sinking ship. And so while I don't think it's too late, there's definitely a window of time where if you don't sort of get on board and start... Start moving, it's going to be really hard to catch up.
Speaker 2Scan my nail for me.
Speaker 3What?
Speaker 6There you are.
Speaker 2Oh my gosh. Show the cameras what happened. It is my... What's going on? YouTube page. I've got an NFC chip in my nail. It's the most genius thing I think I've ever done. Hello and welcome back to In the Blink of AI. I'm Georgie Healy and after today's episode, you're going to be a massive fan of Benjamin Plummer. I definitely am. But I confess, Ben came to me very strongly recommended by very smart people, but that happens a lot. I have the most incredible recommendations and it took me a long time to schedule a call. Anyway, four minutes into that call, I was like, oh my goodness, I need to move some things around in my calendar. Ben is the co-founder of Dragonfly Intelligence. He ran Invisible Technologies, working with Frontier AI Labs. And before that, he spent nine years at Bridgewater. In New York City, the world's largest hedge fund. He's been working on AI since before TrackGPT was even a word in our lexicon. And today we get into some really big headlines, like the Hugging Face hacking incident and the lack of accountability there. New bottlenecks in building businesses and its people. What does nine years at Bridgewater Capital under Ray Dalio teach you? And software bloat. We've got more and more features and vibe coding and platforms trying to be something like everything. We kick off the show with me asking Ben to scan the NFG chip in my fingernail. So make sure you watch that part of the video and let's dive in.
Speaker 1You're listening to a Day One.FM show.
Speaker 7Found a scale faster on deal. Set up payroll for any country in minutes. Hire anyone, anywhere. Get visas handled fast and get back to building. Visit deal.com slash day one. That's D-E-E-L dot com slash day one.
Speaker 2Thanks, Ben. Ben, thank you for coming on in the Blink of AI. We always start up the episode with a hack of the week. Start us off strong. What's your hack of the week?
Speaker 3So one thing that I've been using a lot lately, and Andre Kapathi, a famous researcher, tweeted about this recently, is the rambling session, which is take a microphone. Take your recording, transcribing tool of choice, and just talk for five, 10 minutes, explaining an idea, a concept, a problem you're wrestling with. And there's something really powerful once you learn not to try and compress your thoughts into a sentence or two of a prompt, and you really give it all the richness of your thinking, the things you're not sure about, the nuance, and feed it into the model. It's incredible. It's incredible. It's incredible. It's incredible how well it can distill all your rambling thoughts and play them back to you more coherently and clearly than you could ever imagine. It's amazing.
Speaker 2What are you using to ramble?
Speaker 3Mostly WhisperFlow. Sometimes I just go have a granola meeting with myself for 20 minutes. It doesn't matter.
Speaker 2Can I tell you, I've been doing this, and it is one of the few AI tools, WhisperFlow, that isn't just a fun little hack that I then retire. It's incredible. And you're right. I do a lot of content. I do a lot of writing. I do podcasts, obviously. And it's about the vibe. It's not about having the perfect prompt. And to get that feeling, that emotion, it's really hard to put that in a sentence sometimes.
Speaker 3And it takes some getting used to, like even pausing and gathering your thoughts. It's quite weird when you're in a sort of conversation, but the AI will happily wait for five minutes if you need five minutes. It's not mad. You do need to sort of retrain yourself to allow yourself to really just sort of express the thoughts. I find, especially for creative things, it allows you to just sort of like really get in a flow and explain what's on your mind, whereas the typing can be really slow. So yeah, I use it all the time. I probably five, ten times a day. It's amazingly powerful when combined with the LLMs to clean it up on the other end. I think if you just got a long rambling list of your notes, probably not so helpful, but it's helpful in both ways.
Speaker 2So strong. Don't try this at home or do. My best thoughts come like really late at night. The whole family's asleep. But with WhisperFlow, you can genuinely whisper. So I'll be in bed being like, great event idea.
Speaker 6Yeah, I have those in the morning sometimes. The same thing. Everyone's still asleep.
Speaker 2Okay, you're going to be my guinea pig for the morning. You'd think I invented nuclear fission. I'm so proud of myself. Do you have your phone on you?
Speaker 6I do. What do you got for me?
Speaker 2Scan my nail for me.
Speaker 6What? There you are.
Speaker 2Oh my gosh. Show the cameras what happened. It is my YouTube page. I've got an NFC chip in my nail. It's the most genius thing I think I've ever done. I put my YouTube page on there. There's a lot of events this week. It's just a much more fun way of, you know, oh, what's your podcast? It's called this. Then they may or may look it up on their phone. QR codes are not so whimsical. I find that fun. Is that fun?
Speaker 6Surely people have to remember. That's not, I've not seen anyone doing that. That's awesome.
Speaker 2So I've got my YouTube page on. I'm going to put my LinkedIn on my other thumb. I will make all my nails the same color. I just wanted to try this before the pod. That's my hack of the week. Thank you for humoring me. Okay. So speaking of hacks, I genuinely need to dive into this. Hugging face, Chachi PT hacked it. It became like this. It felt like a celebratory moment where they're shaking hands and like so happy. With the partnership of being hacked, I'm like, should we be more worried? Is this marketing? What's your take then?
Speaker 3Yeah, I think both are true, which is certainly, I don't think we've ever seen a company be so proud of losing control of its own software and having it do things that weren't intended. And I think, you know, there's certainly an element of this where, you know, the idea of these models becoming more and more powerful and tropic doing a similar thing in the early phases of fable. I think. Is a very consistent sort of pattern with that. I think on the other side, these threats are very real. And in this case, it wasn't sort of a malicious threat. It was the model basically trying to cheat on a test that it was doing. And that is a very common behavior for these models, the way they're trained and the way they're incentivized to sort of achieve a goal means they will do basically anything they can to achieve that goal, which in this case includes breaking out of their own harness. So I think these sort of challenges are very real. I think the other interesting part of that story, which is sort of less reported is Hugging Face actually caught it about a week or so before OpenAI. They did so using open source models. And so there's a very interesting sort of other side to that story around the balance of these things where sort of the either the threat or malicious actor side of things and the detection mechanisms. Really need to move in unison, or I think there's going to be challenges. I think the other sort of really interesting thing there is just around accountability and who's actually on the hook for this, which is, you know, in the OpenAI case, it's their models and their actors basically doing it. And so it's a little clearer, but in a world where some other company using OpenAI or anthropic models does something similar, I think it's going to be very hard to attribute who's actually at fault and who's responsible. In this case, there was no harm. There was no damage, but very easily that could be a different situation. And I think the sort of legal frameworks around the accountability for this is really unclear and something we're going to invest, need to invest a lot in.
Speaker 2So well said. Imagine a person, imagine you or I hacked into Hugging Face. I don't think it would be the public announcement.
Speaker 3No, the police would be.
Speaker 2Yeah, the police would come over.
Speaker 3Exactly right. And so, like, I think that is sort of the undercurrent of this thing is really around who's responsible for these models as they become more and more powerful. And, you know, a lot of this is at the sort of harness layer, that these underlying models are way stronger and more capable across a range of different dimensions. And these safety guard rails are meant to protect them from doing these particular things. are not bulletproof, as this points out. many examples online of people being able to sort of break out of those safety guard rails and that is another challenge here where you sort of got these things that are increasingly more and more powerful that need to be contained and understood on the other side in terms of what those threats
Speaker 2look like and how to protect against them such a great point i remember when we had anthropic on the show they were talking about agentic harnesses and it sounds perfect right you know this is the safety this is the privacy this is you know all our rules and regulations and it's like great done but since this happened like a powerful company like open ai doesn't have control of their harness
Speaker 3or exactly i mean you've got something that's effectively locked in a box that you're making smarter and smarter and its incentives is to break out of the box and so like inevitably what do you expect is going to happen particularly if you have actors who are trying to get it to break out of its box in this case that doesn't seem like that was necessarily so but certainly there will be others who are um i don't think this is necessarily needs incredible amount of alarm and blocking these models and things like that i think really what it is about is recognizing that as these things become more and more capable we need to be just more aware of the threats that
Speaker 2are emerging in different ways one more thing on this you mentioned you know hugging face being the ones that identified this this had happened anyone that's listening that might not be aware it's it's like a github for ai models this is a very advanced technical company that could identify something like that happening do you think this would even be identified with a less sophisticated company or
Speaker 3highly unlikely like i think you're talking about um the real sort of operational launch of technology capability um within these organizations interestingly you know you could kind of argue that they're a indirect competitor of open ai and in a lot of ways they represent open source and a sort of alternative path and so even that creates like a really interesting dynamic imagine you had some other company accused of hacking its effective competitor you'd have a whole bunch of different questions and so yeah i i think most organizations are totally unequipped to deal with the level of sophistication you know you're looking at these models being able to find vulnerabilities that have existed in software for 20 years and this is like web browsers core infrastructure of the internet that sort of every engineer in the world has had access to and i mean poking holes in these models are finding gaps in in that core infrastructure and so it's just a whole new level of capability that most organizations don't have and i think one of the risks of you know companies that probably shouldn't be building software building software opens up is that yes it's much easier than it's ever been to build and vibe code and prototype things but actually building robust secure software is probably as hard as it's ever been
Speaker 2hot take right there you're gonna have so many opportunities for more hot takes you're the ceo and founder of dragonfly intelligence uh after a series of incredible career steps which we'll unpack but some some things that i read on your website and on your blog that i'd love to unpack one of which ai expands coding capacity exponentially we've seen that all the engineers have have reported this but then every other process in the pipeline is exposed as the bottleneck what are some bottlenecks that you think are particularly
Speaker 3risky yeah look i think if you sort of focus in on engineering and software development to start you know either side of the coding activities i think you're seeing people sort of slow down is one you have to think about what you want to build decide what's important what's not have judgment around that and that you know there's elements of that that you can use these models to help with but you know the humans think are still really important in that process of sort of taste and judgment and decision making and on the other end of that is this sort of verification and validation of like did the model actually build the thing i wanted and i think you know in a lot of cases you know you're not going to be able to do that humans are still quite important there and so what you're seeing is historically the sort of coding was the part that took all the time and there was much less effort on these other things it's now sort of flipped in the coding becomes really quick but deciding what you want to build and then checking it was actually built becomes a real problem and then if you zoom out i think there's a much bigger sort of organizational societal problem which is we're just being overwhelmed with software which is because it's so easy to build and it's so easy to build and it's so easy to build and it's so easy to produce um apps or solutions to things every day there's 50 new applications open source closed source that do one very niche thing and even in areas that i care deeply about and spend a huge amount i don't have time to go test them all out and figure out which one's good and which one's bad and that's the same within an organization people's ability to change and learn about all these new tools and capabilities um is really going to become the bottleneck and there's probably a backlog of 20 amazing apps or sort of frameworks that have come out in the last month that i'm desperate to try and just don't have the capacity to go figure out exactly what they do and how they fit within our stack and so again it's sort of back to the humans being the lowest common denominator i think people on the forefront are getting more creative around that which is okay how do we use the agents to actually go test four or five different pieces of software and then we're going to have to figure out how to do that and give me an assessment of how they work and how they're different and so we're slowly sort of abstracting our way out of some of those activities as well but that's going to take time that we're
Speaker 7sort of very early in those phases i think found a scale faster on deal set up payroll for any country in minutes hire anyone anywhere and get visas handled fast so you stay focused on scaling deal takes care of onboarding hr it eor benefits and compliance so your team can grow without borders it's why more than 40 000 fast-growing companies trust deal to move fast visit deal.com slash day one that's d-double-e-l.com
Speaker 2slash day one i i don't know if we can equivocate it exactly but it does remind me of the complete like explosion of a i imagery for a short time there and then everyone was like revolted and disgusted and i haven't seen it quite so much do you think that's kind of similar yeah it's i think it's going to be
Speaker 3really interesting in that like one of the things that's happened over the last decade say is i'd say generally software has gotten way worse and it's gotten way worse because things you start a company with an idea that's different that has an opinionated view of how things should work and then over time you just keep adding features and capabilities and it becomes completely unopinionated you're trying to be everything to everyone it's bloated with all sorts of different things that no one needs just go open up slack excel whatever whatever software you want there's a million things that you have no interest or need for i think there's an opportunity for ai to strip that back and end up with software that is purpose built for you that has every feature you need and none of the features that you need and none of the features that you need and none of the features that you don't need that i think is much more of a sort of apple approach to software development which is like there's simplicity and beauty in minimalism and and i think that's really exciting in in what it can do for companies and people i think on the other hand there's basically no barriers to shipping software now and so there's going to be a whole bunch of noise and crap that people are going to need to sort through and so there's a little bit of good and bad that i think is going to come from this over the next few years
Speaker 2i know it when i see it in those like monstrosity websites that it's every feature we do everything
Speaker 3and it's like just ah stop every customer every salesperson wants one more thing and they never take anything away and that's sort of been i think the last decade has been that shift towards just adding and adding and adding and not really taking away anything so agree before we talk a bit more
Speaker 2about dragonfly intelligence you have a crazy origin story i remember when we were talking about when we first got on a call i like my a4 page just got filled up very quickly trying to figure out what to talk to you specifically about can we start about bridgewater nine years the world's largest hedge fund why are you nice
Speaker 3um i think people fundamentally misunderstand bridgewater um i think you know there's a lot of sort of quirks around how that place works and i think it's a lot of quirks around how that place operates that gets a lot of attention whether it's the recording and some of the tools i think at the essence it's about this sort of search for truth and and that comes from trying to understand the world and how the world works it comes from trying to understand each other and what we're like and actually when you sort of think about it being direct and telling you what i think and when i disagree and when i don't agree with you and why and when I think you're making mistakes is so much nicer than sitting here and thinking you're making all these mistakes and not telling you that's not how most companies operate but that was sort of very much an idea that Bridgewater lent into and once you've sort of experienced that it's kind of hard to ever go back you just kind of like how would it why
Speaker 2would it ever exist any other way I love that I love when someone tells me to my face I don't I I read your media deck and I didn't like this I love that that's so helpful yeah otherwise you
Speaker 3just get crickets and then you wonder why and it's really interesting how in other domains people sort of really understand it like if you go to professional sports or whatever it's like very clear there's a goal to win and there's no surprise that hey you look at tape afterwards you give each other feedback you're out of position you missed that and it's sort of very clear that purpose is making it better it's not that you hate those people it's not that you're trying to like advance your own causes it's like we're all playing a game the purpose is to win we want to get better and I think Bridgewater sort of took that same philosophy around you know the markets in which they wanted to participate and win and it was the same mentality of like how do we get better every day how do we push each other what type of environment and culture is required to enable that and it's sort of very refreshing after you sort of go you know you're not going to win you're not going to win you're not going to win you're not going to win you're not going to win so it's it's a little it takes a little bit to adjust to honestly I think as an Australian it was way more natural to me than a lot of Americans feel the the sort of transition I think we're naturally more direct and and sort of culturally I think a little bit more aligned with that way of thinking and it is interesting how different cultures react very differently to that type of behavior yeah hot take Queenslanders
Speaker 2should work at Bridgewater in New York City because I feel like we've given each other so much shit by the time we're done with this show I think it's a little bit more aligned with that like it's like hit me I can handle it and clearly a sucker for punishment nine years in a hedge fund and then New York City you started a startup no no before that elemental cognition tell me about
Speaker 3that first yeah so elemental was an AI lab that we incubated within Bridgewater this was probably 10 maybe more years ago and the founder Dave Ferrucci who was the inventor of IBM Watson back in the day that beat Jeopardy had a very clear vision for some sort of next wave of AI and and what some of the gaps were in the current sort of machine learning approaches
Speaker 2and and other things that were quite popular at the time and for the record 10 years ago is before
Speaker 3GPT-3 like well before yeah yeah no one's even and maybe even longer but it was certainly at least 10 years ago um and so Bridgewater sort of seeded that and invested in um growing that out it got to a point where it was very clear that it was sort of much broader applicability than Bridgewater had used for and so we spun it out and raised external capital from a bunch of sort of top tier VCs and I left as part of that and built the commercial side of that business over a few years and it was just a fascinating time to dive into the AI space as you said this is a couple of years before chat GPT but it's very clear that these technologies are heading in a direction that's going to be much more broadly useful than pure sort of machine learning algorithms or anything like that so it's really exciting time lots of challenges in taking an extraordinarily talented group of researchers um and scientists and building a commercial business around them so lots of learnings but amazing fun and I learned so much about these technologies the limitations the types of places they succeed and fail through that experience that's been really helpful over that time
Speaker 2um I I have to know when did you think AI might be a thing it's it's fine I mean you in every
Speaker 3Juncture it seems um so obvious but then you look back at what's happened since each one of those pieces and you're like I had no idea what was about to transpire and so I think those early days of Elementor you just sort of have these aha moments where and I suspect everyone sort of had these some of the early days of interacting with chat GPT where you're like wow this thing really surprises me in ways that I have not you know ever experienced with technology or software and so I was lucky enough to have many of those moments early on where it's clunky it makes mistakes it does silly things but you see these glimmers of potential where it's very clear that those the kinks will be ironed out and what remains is something that's like incredibly profound
Speaker 2and did Elementor give you like did you get bitten by the VC bug the like what like what is it like to have a real startup I'm gonna go all the way to SF now yeah like what was that thinking
Speaker 3it definitely I'd sort of gone if you sort of follow my career it's like every point I'd gone smaller and smaller and earlier and earlier stage and so I mean Bridgewater wasn't big by any stretch of the imagination or about 1500 people but it started to feel big after a period of time Elementor was a great way to sort of move you know still pretty closely connected with Bridgewater but do something more entrepreneurial and invisible was sort of the natural extension of that which was a company I'd been advising with and helping from sort of the early days of Inception that I
Speaker 2eventually took over as CEO tell us uh how much revenue and profitability so we grew a lot so we
Speaker 3went from about 10 to 15. million uh USD to about 160 over the course of 18 months uh which looks really good on a chart uh it does look good it really does um has a lot of pain and kicks and twists and turns um underneath that chart you should have a lot more gray hair um but just an amazing experience we sort of ended up in an incredible position in the middle of um I'd say the the largest reallocation of talent and capital I've ever seen in my lifetime but I think probably that we've ever seen in society and so we had an amazing group of clients we had an amazing team and got to do some really cool work
Speaker 2my favorite part of the interview you're back in Sydney Australia proud CEO of Dragonfly intelligence tell us what it does tell us why you're excited after doing these incredible things to date that you're like you know what all in on this yeah look I think in a
Speaker 3lot of ways Dragonfly is the culmination of all of those things which is you know one of the things that was really apparent to me in the work that we did in invisible is just the huge disparity between what these Technologies are capable of how fast they're moving and how fast they're adding capabilities and how far behind the whole economy is in terms of being able to metabolize those changes and realize those benefits and in some ways the companies that have the most to gain at the least equipped to actually harness those capabilities and so I spent a lot of time thinking about what is the right mechanism for sort of creating value and driving this change um in this new era and I it was clear to me it wasn't going to be sass it wasn't going to be consulting companies were not going to figure this out on their own and so the idea behind Dragonfly is um we're a capital allocator and investor we we invest in buyer companies we're a technology company we've got our own proprietary AI platform and we're operators we're entrepreneurs and so we take these businesses we rebuild them from the ground up as AI native versions of themselves and really unlock the potential that exists within these Industries that you know in many cases haven't evolved in 20 or 30 years um and we think there's a really important moment in time now where AI sort of unlocks a number of constraints that have been holding these businesses back
Speaker 2right ready yourself up I have so many questions number one is how do you identify a company especially private markets where I I called them zombie companies when we spoke before they may have got VC capital and they actually don't have any customers or they're actually not making any money or their churn is crazy and you don't know any of this like how do you identify good and how do you know oof
Speaker 3what's red flags yeah I think there's sort of two categories of you could call them zombie companies um they play out slightly differently I think there's a lot at the moment of these I built some product that I find useful therefore it should be a business and lots of people running around trying to convert the first thing into a second thing and realizing that building the product isn't the hard thing particularly now with these tools selling it serving customers you know building Revenue raising money all of these other things are actually way harder and and particularly great engineers sort of solve the first thing very quickly and then very quickly run into these other sets of problems and they're smart enough that rightfully one or two companies will take a punt on them and they'll get some traction which in some ways is worse because it sort of can perpetuates the belief that they have a real business and I think there's lots of those um sort of different versions of that of every week there's someone else pitching me some agentic harness thing that they've that they're using that they're going to turn into a business and the reality is like every great engineer has one of those now you're not like building something that you're going to sell to the masses and the average day person is not looking for an agentic harness and so there's like a market it's not in my google search history no um and then i think the other sort of category and this is sort of much more i think strategic is when each model release comes out does the gap between your business and the alternatives get larger or does it get smaller and there's a lot of companies that have built things that you go back to like in the early days there was um i can't remember the name of now there was like a writing app that like exploded it hit like 100 million in revenue it's not gramelly or one of those i can't remember the name of now unfortunately um but explosion in revenue it was really cool it took these models that were kind of hard to use and made them slightly more expensive and it was like a lot of people were like oh my god i can't more usable and it was like at the time one of the fastest revenue growing companies you've ever seen but then very quickly like open ai could just do that and then claude could do that and people like why am i paying for this other thing and i think there's a lot of startups that are just ahead of the models that they've sort of taken the models and layered them or put a ui on them and it is helpful now it's better than using the models themselves but that's just a matter of time before they basically eat you up um and that's a hard place to be in and i think knowing where those models are going and i describe it like a freight train like they're coming and you better not be on the tracks that's easier said than done
Speaker 2it's not totally clear exactly yeah we're not judging like it's not easy to do but we see it
Speaker 3but that's a really important question to us i think for venture investors investing saying hey how's this going to play out is this going to be consumed by the models is there real sustainable which means this company is going to win over a long period of time or is this a flash in a pan thing which could have great success for a period of time and then likely dissipate and and you can still make a lot of money you can build some great businesses but those are probably not venture scale businesses just given the time horizons that are involved yeah you need um i think for vc funds it's what an average of 10 years yeah i mean there's a whole bunch of factors that meant that those distributions and returns have really dragged out companies are staying private for a lot longer there's not ipos there's been less m&a and so while those time periods used to be a lot shorter
Speaker 2they're really dragging on now so can i be cheeky and ask you to be specific because like i'm thinking of industry verticals but i've seen them get disrupted you know dentistry uh products uh that rely on image generation and then when nano banana came oh there goes that one like what what is it is it getting even more domain specific yeah so i think like to use a
Speaker 3specific example there's a lot of incumbent software use zero as an example um that have a really actually have an incredible opportunity when you think about what they have they have extraordinary distribution they have millions of customers they know a lot about them they have a they basically need to cannibalize their existing business and compete with their customers instead of providing accounting software go be the accountant and that in theory is a really easy thing to sit down and say like of course you should just do that that is like the way out of the current predicament you're in getting alignment across shareholders board ceo leadership team all the functions that have a vested interest in protecting what exists there is nearly impossible and so i think i just use that as an example it might not be the most extreme one but i think there's a lot of versions of this where there's a theoretical obvious move a company should make to get themselves in a stronger position to move up the value chain move out of this sort of like token eating into your margin problem and actually solve the customer's problem don't just provide them a tool solve the problem but it's actually a lot harder to execute than it sounds and so there's part of what we're betting on with dragonfly that actually starting with the services business and turning them into an ai native business is probably faster and easier than taking a piece of technology and trying to turn it into a service delivery business you don't hear that take often i like it yeah i mean it's it's people have been attuned over the last decade or so that software is where the value is at it's where the multiples are at it's where the scalability is and i think that is becoming less and less true that it's really hard to build defensibility and software alone that there's a lot of competition and sort of replication you build something someone else builds it the next week and so i think we're at a tipping point where a lot of the things that were true over the last decade will cease to be true in the following decade
Speaker 2i have never been more important with my charisma i will tell you i was so sad i couldn't code ben i was like oh my gosh i did chemical engineering and i was like oh my gosh i did chemical engineering not software my career is over and i'm like i'm back in the game okay so i went so deep on the chinese ai models because we we knew about deep seek that that changed the game in terms of wow you can do a lot more with less but with moonshot ai and their kimmy models and that there's just so many chinese models i i read somewhere that there's a new kimmy model every 10 weeks or something crazy like that when you're when you're looking for businesses to buy and they're private what is your recommendation open source like start somewhere we'll we'll we'll work
Speaker 3our way through this is like such an amazing look i think my overarching advice to people is do not get yourself trapped in being stuck with any particular model provider or approach like we're in such the early phases we've seen the back and forth between even the frontier labs in terms of who's leading the model and who's not leading the model and who's not leading the model and who's leading who's most cost effective who's better at what yeah ben i keep changing my stickers it was open ai now it's claude you know if you build your whole organization around one model you very quickly could be stuck with the most expensive the least capable or whatever and i think open source just broadens that continuum of options that you have there are more complications with the sort of open source models that i do think you need a bit more technical depth and understanding around how to use them how to put some of these guardrails and so they're a little less user-friendly but extraordinarily powerful and i think at the moment they're they're like a quarter to a third of all tokens are going through open source models i think that's only going to increase like i think there's a lot of companies that are mostly focused on building out capabilities first and then we'll think about optimizing cost second and open source has a really key role to play in in that
Speaker 2cost optimization i read that kimmy k3 i know we've got opus 5 now but it was a quarter the cost and similar benchmarking to opus 4.8 it's like what are we doing i mean there's some important
Speaker 3nuance there around a there's a lot of accusations around are they distilling these models are they benefiting from them i think certainly in the past that has been the case whether that's ethical or not i think sort of put that to the side i think it's certainly helping their catch up and then the other thing that sort of isn't reported as much is the benchmarks can be a little misleading in that one of the things that if you sort of look at the real world use of these models is these models are are hyper optimized for the benchmarks that they're trying to appear much better than they are and then when you look at their sort of real world usage it drops off it's much more spiky and then it's really good in certain things and then much worse in others where because the frontier models are getting so much usage they're pretty well rounded in terms of the frontier of their capabilities and so the headline numbers don't necessarily always tell the story the full story but it's certainly true that they're accelerating and they're very close to the frontier and increasingly seem to be closing the gap between you know if they were six months behind maybe they're three months behind now or something but it's not far
Speaker 2i did try and try one of these models for myself you can pick the one of the kimmy models in uh github it's got the same drop down as as the you know western labs but it's coding and i i'm not a coder so i can't really judge ben how do you recommend to your companies what to do yes don't be too loyalist to any one company
Speaker 3but high level is there anything you tell them yeah so it's really important to understand how the models perform for your business and your needs and that means you'll hear this word even vowels but basically having a set of tests that you can run the models through constantly to evaluate how they're performing relative to your specific tasks and the reality is any business is made up of hundreds of different tasks and the right model for one task might not be the same model for different tasks and so it's not even one uniform answer for a particular business we have lots of processes we run that might have three or four different models in the same process and so those Because evals are really important in knowing how it's performing, what is the quality per model? What is the cost per model? And you can make informed evidence-based decisions around those. And I think the sort of state-of-the-art company have built really sophisticated infrastructure around these evals that as soon as a new model comes out, they're able to run it through every single use case they have across the organization and flick a switch and say, we're migrating to this model for these five things. We're using it as a backup model for these three. And it's automated. They don't even need to make some sort of judgment decision around those things. It's sort of based in data and math around which models perform best. Organizations don't need to build something that sophisticated, but certainly having a sense of how do these perform in real-world scenarios for my business is really important. Just because one's better on a benchmark doesn't mean it's the best model for your specific use. So well said. Even listeners
Speaker 2will know. If you want to do your color analysis, you can use track TVT for that. Do you really need Fable? Are you burning through credits when the questions you're asking are not that critical for that? I actually have yet to see a use case for Fable that really made sense. Have you seen any?
Speaker 3Well, look, I think certainly we are in that same process of sort of experimenting with how much better is it? Where are those gaps? We're testing it as we should, as the data is being used. We're testing it. We're testing it. We're testing it. We're testing it. type of company that we are and so we're using it for a lot more sort of strategic planning architecture decision making and you know in some cases running that side by side with opus or other models to test like where does it outperform there's certain areas would actually underperform which overthinks things and it sort of makes things more complicated maybe than it needs to be and so that sort of fine-tuned understanding of what the character of these models are is really
Speaker 2important it is like a character right they've got personalities they've got their own personality
Speaker 3and look i don't know if you saw this whole open source thing has blown up in the u.s over the last few days the sort of various voices in the trump administration on both sides saying they should be banned and kimmy should be banned and there should be all sorts of controls it seems like some lobbying from the frontier labs around that um a huge um you know i don't know if you've heard of it but i don't know if you've heard of it but i don't know if you've heard of it but i don't know if you've heard of it but i don't know if you've heard of it but i don't know if you've heard of it but i don't know if you've heard of it but i don't consortium of nvidia of google microsoft the who's who strongly supporting um the open source ecosystem in these models and so that opens up just a whole can of worms around what does banning these open source models really mean what geopolitical implications of that how does a country like australia participate in that where we don't have our own frontier models the open source models can be really valuable to us from a strategic perspective if we get cut off from fable again or fable like models though those open source models are really valuable and so it's going to be interesting to see how this plays out over the next few weeks it's hard to not criticize everyone's kind of taking a pretty obvious self-serving position
Speaker 2around oh no kidding you don't want us to not spend money on u.s models um what if meta became more powerful in the open source what do you think would happen then because american
Speaker 3company i do think you've sort of got two problems wrapped together which is this sort of closed versus open source and then u.s versus china and it just so happens to have played out for a variety of different reasons that china's been the clear leader in open source models and the u.s has been the clear leader in closed source models and there's been various attempts in the u.s it's sort of half-hearted efforts to have sort of open source models but nothing that's sort of been really credibly close to the frontier for any sustained period of time i do think one of the likely outcomes of this sort of current posture is that there will be either incentives or more pressure for the frontier labs and hyperscalers to start producing frontier level open source models to give a bit more of an alternative to the the chinese model and so we'll see i i think that's like a little bit tbd but that would be my guess so you think like anthropic or or google or like anyway there's a whole bunch of them that have the talent the compute the resources to go do this and and some of them actually the incentives to go do so and so i wouldn't be surprised if we we see um more frontier-like models coming from those other
Speaker 2players when it happens i'm tagging you in a linkedin post heard it here first okay when you guys announced your business of course it was in the afr the most prestigious paper in the country uh and it was posted on your linkedin and i was looking through the comments i was borderline obsessed with this one this is a great idea there are so many boomer slop businesses just sitting there what's a boomer slop business i know you didn't write it um and and another common rhetoric in the comments was this is such an ambitious thing australia needs to be more ambitious and this is such an ambitious thing australia needs to be more ambitious like this do you think you guys are particularly ambitious or you just stand out
Speaker 3comparatively uh it's a good question let me answer that i think we are trying to be quite ambitious in the sort of breadth and scale of what we're setting out to do there's something like two trillion dollars of services in australia i don't think most of it has changed in decades we think we can put a huge dent in basically bringing that into this sort of ai native world and there's a really amazing blog post i think from sam altman a few years ago about doing hard things and his basically view was i loved this it's actually easier to do hard things than it is to do easy things you can attract better people you can attract better talent it's less competitive it's worth the effort of dragging yourself through it and i think there's a lot of truth to that that actually the more ambitious you are the easier it is to get people excited about the mission and where you're going and what you're doing and so i think there's a lot of that that's sort of woven into dragonfly which is there's no reason we can't be maximally ambitious around the impact we want to have on the australian and global ecosystem i think there's a lot of ambitious people here in australia i think there's a bit of a density problem that they're kind of scattered all over the place and you kind of have to find them and they're not all concentrated in the places that you would expect in a way that the us kind of does um and the broader sort of economy has been quite rewarding to people who don't take risk it's kind of been a great ride property's been great there's good stable paying jobs there's lots of stable sort of companies and so i do think that over time folds into the culture and the dna that there's a lot more risk-taking and sort of contrarian thinking that's prevalent particularly in the u.s i think the u.s stands sort of head and shoulders above most of the world all of the world in that case um and but there's no reason that needs to be true like i think australia can do amazing things we've built some incredible companies at global scale and we should
Speaker 2build a lot more of them perfect follow-up around data centers this is such a hot topic then if you're on my uh instagram algorithm it's they're pure evil like they're they are actually hell on earth and nothing could be worse than to have even a single data center here but then on my linkedin it's like can we be a little bit more ambitious how are we going to compete in the future how is the economy going to survive can like we we don't want to hurt the environment but let's be realist where do you fall look i i think this is like super clear cut for me
Speaker 3which is it's as you said we need to have a vision for what type of country we want to be and how do participate in the next 20 years of the global economy advancing we mostly spent the last 20 years digging rocks out of the ground and sending them overseas um that's probably not the place in the world that's my undergrad degree but thank you um it's probably not going to be how we want to position ourselves for the next 20 years and i think it's important to look at like what are our competitive strengths and play into those and not try and replicate what makes sense for the countries might not make sense for us and i think we have extraordinary energy capacity for a variety of different mechanisms we have a huge amount of space we have a pretty sophisticated like construction ecosystem there's a ton of reasons where there's a proximity to asia there's a whole bunch of different things that mean this is an amazing place to build these data centers we could build a real competency and expertise in these ways that you know we're not going to be building frontier models in australia ever and so that's not going to be our path to participate in this sort of ai wave i think the data centers is a very credible way for us to play to our strengths and solidify that i think the superannuation capital base is another really unique feature of the australian sort of financial system that can take long-term views and support some of that stuff and mostly like if you actually look into the environmental and the pricing stuff it's like totally alarmist it's like ignores a whole bunch of facts around what's likely to actually transpire around renewables the mix of technology no government policy around pricing and the same thing happened in the u.s there was a lot of sort of alarmist rhetoric the i think actually the trump administration came up some really sensible policies around how the hyperscalers pay for the electricity they use it doesn't disrupt the grid if anything it's actually feeding power back into the grid subsidizing that is what they're all saying
Speaker 2we're going to pay more electricity bills so that's not happening in the u.s no i'm not sure
Speaker 3that that policy actually flowed through but all the major tech companies have agreed to it which is they'll build their own power so if microsoft's building a data center in texas it will build a power station next to the data center it's totally off grid it's not taking power away from anyone they're actually selling extra capacity back to the grid and introducing that more power and i think it doesn't need to be that exact configuration you I think there's lots of ways to solve that problem if you're committed to actually doing something versus just sitting on the sidelines and complaining that like it's not perfect
Speaker 2no risk of people saying we didn't we didn't really step on the wire here uh i was sent a link to abc four corners and i knew it would end badly about how bad ai is and data centers was one of the things and i copied this transcript because it was such alarmist journalism that i was like i'm just gonna copy and put in claude and be like verify the facts one of which was the cost of these data centers uh the water usage the the power and and claude was like respectfully this is assuming that you would use you would have the poor efficiency of a very very old 30 year old data center when you're building them from scratch you would never build them in that way and they're extrapolating in a way that's just false and i think that's a really good point and i think it's it's a really good point and i think that's a really good point and i think it's a really good point
Speaker 3yeah people are just cherry picking like you can go find data to support whatever point of view you want um it's not exactly like the most productive going back to sort of like the search for truth it's not exactly the most helpful way of navigating what are in some cases quite complicated topics but i think the idea of sort of just sitting on the sideline and outsourcing basically everything to the rest of the world um seems like a much more terrible path than trying to work our way through it so i think it's a really good point and i think it's a really good point and i think it's a really good point and i think it's a really good point and i think through some of these problems many of which actually could you know subsidize huge investment in renewables and distribution and a whole bunch of other jobs and creation around these things that come with making investments in big sort of audacious ideas this is unfair because it's
Speaker 2not your background what is your thought on ai submarines like defense and and it's just a topic that seems to have really blown up like a lot of people and i think it's a really good point and i think it's a really good point and i think it's a really good point and i think it's a really good point
Speaker 3yeah look i don't know like have any expertise on the sort of deep technology side of things it does strike me as there's a set of sort of industries where australia is a little bit sort of goldilocks in that we're big enough to matter and have the right sort of infrastructure and rule of law and stability to sort of be a mini version of america or europe or whatever you want to describe but actually just much simpler in that you can do a lot of things and you can do a lot of things and you can do a lot of things and you can do a lot of things and you can do a lot of things move faster and cut through a lot of the sort of bureaucracy and the institutions that are there to sort of protect the way whether it's the military apparatus or whatever and so it's not surprising to me that there'll be a handful of these industries that australia actually can leapfrog ahead because we have great talent we have capital that can be put to work against these things and we have a sort of semi-structural advantage relative to some of these bigger markets which just move much more and more and more and more and more and more and more and more slowly and then release them to the world and so i without sort of being deeply familiar with the
Speaker 2technology it doesn't surprise me i'm deeply familiar that they're square for some reason these ai subs are square uh okay so there's three things australia should play in it's not frontier
Speaker 3models what would the three things be that's good um so i think the data centers is a clear one um i think there's no reason we can't be the one of the leaders in like the application of these technologies which is how fast can you take what's coming out of these labs and push it through the real world into the economy it's obviously what we're trying to do with dragonfly but you can think about that at a government level you can think about some of our largest organizations i think for many of the same reasons we're not as stuck to hundreds of years of legacy in the way that you know it's sort of europe and america and stuff are and so we should be able to move much faster and then i think the next frontier around sort of physical ai there's a huge case to be made that we can and should play a huge role in that given everything from sort of agriculture to mining um construction these sort of heavy industries that we've built quite sophisticated capabilities in ai is going to be radically transformative to many of these and australia has some of the largest most established organizations across those and so i think there's a real opportunity to do that and i double down in what will be the next frontier of sort of ai development and so those would be my
Speaker 2my three i love them are you ready for rapid fire one minute questions to finish what keeps you up at night look i think um there's like a crappy version of that which is
Speaker 3like not moving fast enough um i there's not the answer is much more nuanced than that than that is it's not just moving but it's actually getting in the right position so that we're able to capitalize on this really unique moment in time and i sort of think about it like a wave which is like you know the wave's coming you can see the wave you can paddle as fast as you want but if you paddle in the right wrong direction it's actually going to pick you up and dump you on the beach um and so it's not only the paddling fast but paddling fast in the right direction and i feel like that's very much this moment in time that energy does not equal strategy and there's a lot of you know moving pieces and fundamental truths or truths that have existed for the last decade or so that i think will no longer be true and so you do need to go back to sort of first principles and be like what do we really believe how do we believe this will play out and if that's true where do we want to be in the field when this comes and so that's like a lot to wrestle with before going to sleep but that's definitely what keeps me up what time do you go to
Speaker 2bed and also a very bridgewater response to you that is much longer than a minute ben thank you queenslander who are the people most at risk of job disruption what should they do
Speaker 3it's a good question um i don't you usually see this come out as like lists of jobs that are going to be replaced or not replaced i actually just don't think it cuts that way i think what ai is really changing is people with agency to go and create and build and do something different are going to be extraordinary beneficiaries this whether you're a doctor whether you're a mathematician whether you're any construction worker you're going to be hugely beneficial for this because there's just this inflection point that allows you to scale your capabilities in ways you never could and if you're sort of predisposed to just sort of putting your head down and like doing your job and not really rocking the boat and not really asking questions i think you're going to very quickly be commoditized out and so to me it's much less about which jobs they're all going to change pretty drastically it's really about how people approach those circumstances which is going to change the the outcomes reminds me uh have you been to the dentist lately my news is ai yeah they should have you i have not and maybe i need a new dentist oh this was
Speaker 2fantastic because he even showed on the screen he had his he asked permission but he had a road microphone i was like okay content girly and he was like explaining exactly what he was doing as he was like going through each individual tooth i could see it transcribed i learned a lot
Speaker 3as a patient yeah fantastic exactly and how the last 20 years before that that experience had basically been exactly the same um and so that's like a perfect example where we're just hitting this inflection point where there's a lot of these legacy industries hadn't really changed are going to fundamentally change over a really small period of time
Speaker 2yeah it's it it feels like the last time i went nothing and now it's all ai incredible is there anything you believed about ai two years ago you know you've been in ai for over 10 years technically that now you fundamentally are like oof i got that wrong
Speaker 3i think the technology has developed faster than i imagined and that the pull through has been much slower than i imagined that i sort of thought the technology would come up a little bit slower like the leaps and bounds we're still seeing now like it's kind of crazy last year as we were talking about have we hit a wall is is the ceiling yeah um and it's i think pretty clear that there's not any wall on the horizon but when you actually look inside these businesses nothing's changed and i think that like in hindsight it's kind of obvious but at the time i certainly would if you made me bet i would have bet that there would have been more sort of real world impact I think that's probably one thing that's definitely evolved and is part of why we're doing what we're doing.
Speaker 2I was speaking to someone yesterday about you can show someone how amazing AI is, they still won't use it if they don't want to and that weirds me out. That weirds me out.
Speaker 4Yeah.
Speaker 2Listen to the show. Last question. What does an Australian services company look like in five years if they didn't listen to this podcast? Or speak to you.
Speaker 3I don't think there's some tsunami event that just wipes all these companies out. I just don't think it's going to transpire in that way. I do think they're just slowly but surely going to fall further and further behind, which is their ability to serve their customers is going to fall further and further behind, sort of state of the art. I think their employees are going to be less and less empowered to go be their best people. That means the best people are going to leave. That's going to then mean that the service drops even further. And so it's going to be this slow erosion of competitive advantage that will sort of play out over a period of time. But once that's happened, it's kind of impossible to catch back up again. It's pretty hard to attract great people into a sinking ship. You're sort of losing market share. And so you can't invest in the same way you could. And so while I don't think it's too late, there's definitely a window of time where if you don't sort of get on board and start moving, it's going to be really hard to catch up.
Speaker 2Beautifully said. How do people find you? How do they discover Dragonfly?
Speaker 3Dragonfly.com.au and Ben at Dragonfly.com.au.
Speaker 2Love your blog. Love what you're doing. Thank you for being on the show.
Speaker 3Thanks for having me.
Speaker 2Thank you. Thank you so much for listening to In the Blink of AI. If you want to go deeper on anything we've spoken about today, I write a weekly substack called Attention is All I Need. Yes, it's hilarious. It's a pun. And essentially, I go into AI rants, tech news, events I'm going to, and more. It's bite-sized, and I hear it's awesome. The link is in the show notes below. © transcript Emily Beynon