Go back

204: “Ship a prompt”

63m 16s

204: “Ship a prompt”

The podcast hosts, Johnson Dell and Mr. Guillrambo, address their long hiatus, explaining that the show has become an irregular hobby project due to busy schedules and a desire to produce quality content. They reassure listeners that the podcast is not canceled and will continue with episodes on personal projects and programming topics. The main discussion centers on their expectations for WWDC 2024, especially Apple's AI initiatives. They are hesitant to build features using third-party AI APIs due to risks like cost, instability, and rapid changes in the AI landscape. Instead, they hope Apple will offer competitive on-device AI solutions, emphasizing privacy, lower costs, and integration with Apple's ecosystem. They speculate on hybrid models where AI tasks are processed locally or in the cloud, and discuss the potential for Apple to subsidize advanced AI APIs for developers. The hosts also explore innovative ideas like action models that could automate app interactions without explicit API support, leveraging Apple's deep OS integration. Overall, they are cautiously optimistic about WWDC, planning to align their development strategies with Apple's announcements to create sustainable, cost-effective AI features for their apps.

Transcription

11314 Words, 60774 Characters

English
Hi everybody and welcome back to stack trays. I say really welcome back there because we know it's been a while but we are really excited to record the show again and are excited for you to listen to it as well. My name is Johnson Dell and I'm of course joined like always but my good friend Mr. Guillrambo. How's it going Mr. Rambo? Is this thing on? Hello? Does it still work? I did actually have to check my audio equipment before we joined the call here just to make sure that everything was still working. After a bit of a hiatus if you will. Yeah, I'm still podcasting regularly so I still know that all my stuff works. However, I am always afraid that I'm going to turn on the mic and try to talk to you and that I will have forgotten how to speak English because I haven't been speaking that much lately so it's good to get the practice on. Yeah, absolutely because all the other podcasts you do for your podcast network they're all in Portuguese right right right makes sense. I haven't forgotten I still have you know client meetings and everything in English plus you know here at home we speak like three languages Swedish Polish and English like all mixed every single day so it's that's the only danger I guess for me is that I will all of a sudden start saying some words in Swedish or Polish but other than that I think we're good to go. So maybe we should just like quickly address like because we have gotten quite a lot of messages from people like what's going on with the show how come there's no new episodes of stack trays what's going on with swith by some Dell and so on and you know I think we've mentioned it before that once we kind of stop doing the show on a regular basis and once we kind of stop having sponsors and these sorts of things. We just started treating the show more like a hobby project and just something we do kind of when we have time and when we feel like we have something interesting to talk about that is not just kind of updating you on the tech news because if we do a podcast like once every two months or something like talking about the news is not really that relevant right it's more interesting then to talk about like learnings from our personal projects and development and talk about programming topics in general and so on. So that's kind of what we're focusing on and you know we mentioned then that you know the show is becoming a lot more irregular as a result and we know that for the past few months it's been a bit too irregular I would say we haven't done many episodes and we do want to try getting back to doing episodes more frequently but like like we say life gets in in the way right like we have so many other things that we're working on and that's going on and it's sometimes hard to find a time where we can both sit down and record an episode and we know that we can do that. And not just record the episode but also had time to edit it and make sure we have a good show because we also don't want to do a show just for the sake of it like we want to do something that we feel proud of and something that we think was a good discussion and that people enjoy listening to so that's kind of why there's no like we haven't gotten tired of doing the show or we don't want to do it anymore and like Swiss and Dallas not canceled and that is not canceled is just that you know we have so many other things to work on and we do hope to release more episodes and and to do it as often as we can but sometimes there might just be a few months where not much is going on but don't worry there will be a new episode until we officially say that you know the show is canceled which we hope will never happen or at least maybe happen when we're super old I don't know right yeah rest assured if we ever decide to not do the show again we will let everyone know so don't worry if it's been a while since the last episode doesn't mean we've given up it's just irregular but we are we are keeping the show on and we'll keep going until we officially cancel it and then we'll let you know don't worry yeah absolutely and I just said there that we're not going to talk about the news but since we are right now before WDC 2024 I thought it would be a good kind of topic to start with here to just do a little bit of a check in around how are we feeling about this years WDC like are you hyped Rambo are there any expectations you have any wishes like how are you feeling about the conference this year yeah I think it's been a while since there was a WC I was looking forward to this much and that's mainly because I have been doing quite a lot of stuff around AI which we will definitely talk more about soon and we especially with Chibi studio we've been doing lots of prototyping around different ideas and things we can leverage generative AI for and we haven't shipped any of those yet and that's because I'm worried that we'll have like this whole thing based on like an open AI API or something similar and then come IOS 18 there will be built in stuff that we could leverage instead and then we'll have all of this cloud based stuff that we worked on and that's not based on that and that costs a lot of money and things like that so we're kind of waiting to see what comes out at WDC in order to decide like if we realize that we can't do much with what Apple announces then we'll just ship what we have based on third party APIs but if there's at least one thing we can leverage from whatever Apple announces that then we'll probably work on that and then ship it later this year and keep working with what's built in because I think that's going to be a big advantage if we can figure out ways to use what Apple is going to provide on the vice or on Apple's own cloud services rather than going with these third parties. Yeah, absolutely because we have seen especially in recent years kind of some of the dangers of building against a third party API with things like you know the reddit API the Twitter API and so on like either shutting down or becoming really expensive or hard to use in other ways and those are just some examples and those kinds of things happen maybe not all the time but they do happen right is it's not impossible they that is a real concrete risk that you're always taking when you're incorporating some kind of server based third party component into an application and especially if you're treating it as like a really core part of the application which you know the features that you've been talking about and you've been showing to me that you're building. It seems like they are mostly additive to the products like for Chibi studio and so on but still you want them to work and you want them to be maintainable going forward and you probably want to keep expanding them and making them even bigger parts of the applications going forward. And then it makes total sense I think to be very careful about you know what kind of technology am I incorporating because especially with things like AI it feels like it's changing so incredibly quickly even by like technology standards you know we're used to a new version of I was coming out every year and a new version of Swift coming out multiple times per year and things changing and we have we're having to update our code basis and relearn things and so on but with AI like this modern iteration of like generative AI basically didn't. It's not like I basically didn't exist a few years ago like it's so new and it's still being iterated on so quickly that I absolutely understand your hesitation to kind of lock yourself in to a third party vendor especially when you know apple are very likely to really something AI related in just a month right yeah and I imagine it's probably not going to be the case where we can just drop everything we have and go fully and on apple stuff. But if we get a good vision from apple on what I've been working on and we get a sense that it's going to keep improving year over year or even within individual OS releases then I think it's something where we can wait a little bit more and ship less than what we've got currently with the third party providers. But in a way that's more sustainable going forward and that we know that at least apple is not going to cut out the API without notice or start charging a million dollars per month for the API even though I know they would love to but services revenue exactly at least in terms of APIs that's not how they usually do things so yeah I think if this was like a year ago and we had all of this cool stuff with our parties built and and ready to ship we could like ship it and take advantage of the hype and like the gold rush but I think the gold rush is kind of over in the AI space not fully done yet but there's still some stuff coming out that I feel like is gold rushy but I think the main meat of the gold rush where you could and I'm doing scare quotes here just release an AI app and get a million dollars I think that's over so I think we can afford to basically just wait and see yeah and I think another key components was something that you mentioned earlier that I think a lot of people us included are expecting at least to a large extent apples AI implementations and models to be local and be running on device because they have been investing so much in core ML the neural engines in all their devices and so on to kind of set up this infrastructure for on device AI and you know they've been talking about it quite a bit it's funny how apple they seem to now how So now I have a question. fully embrace the AI lingo where before they were always saying like on device intelligence or on device machine learning It seems like now is it okay. It's all AI right so so everybody knows what we're talking about And that also is a key component right because then you are not relying on servers You're not relying on those kind of subscription costs and you as the developer taking that on but rather it being distributed on your user's devices and That has of course privacy aspects and that's I'm sure something Apple will emphasize a lot when they talk about it at WWDC but also like cost There's a lot of cost implications there as well and maybe it's going to be a lot simpler for you as a developer to implement when you're already building apps for Apple's platforms So that's I think something also that is something I would definitely keep in mind when making these decisions Right and I think the cost side of things is interesting when you think about Apple because that might be The way they differentiate their offerings from others because Realistically unless they've got something hidden that we never heard about and there were never any rumors or leaks They are not going to have stuff that's significantly better than say chat GPT and things like that It's probably not gonna be as advanced especially if you run on the vice like it's not like you can do GPT 4 level model on the vice that's not currently possible I would even have the expectation that it would be considerably worse because it seems like Unless like you say they've been really working in some secret lab on this for years They are quote unquote behind it seems like it seems like they are Showing up a little bit later to the race if you will and that's not to say like Apple is doomed or anything But it's just that they will probably have they will be earlier in that iteration face So they will probably have something that is less capable, but yeah, like you're saying if we run some device That is a differentiating factor. Yeah, and at least in terms of LLM's right Specifically on the LLM side of things they have some papers that came out they're interesting and They have even an open source model. They released a couple weeks ago and from what people are Saying they are not of course as good as the other stuff that's out there But again if you can run on the vice that's definitely a differentiating factor But then comes the partnership aspect and there has been rumors of Apple is in talks with Google and open AI and things like that So if Apple offers a cloud-based solution even if it's Mainly just a proxy to these other services They could maybe have Cost as a differentiating factor even if they lose some money in the short term Where they kind of subsidize these APIs to their developer ecosystem In a way that say apps on iOS or the Mac or other Apple platforms can offer advanced cloud-based AI solutions Add a more affordable price because Apple is basically subsidizing parts of it at least for now Apple has a lot of money famously so Maybe if they can't win on how much better or how good their models are maybe they can subsidize usage of these more advanced models by leveraging some of that pile of money that Tim Cook has under his Metris or something. Yeah, that must be a very big mattress very thick mattress if that's gonna be the case Yeah, I agree and maybe it can even be something like a hybrid solution where you as the developer Maybe you either have control of it or you by default just ask for inference into a specific model with specific parameters and leveraging technologies like async await you would just await on the response and maybe it will resolve it locally or go to the cloud Depending on what the query is and what the nature of of the thing you're asking for is like I can totally see them doing a lot of Generative AI things on device while if you're asking for some kind of knowledge-based question or some something about the news or something It might actually go off and and and fetch that from the internet because your device of course cannot cash all of the information from the whole internet on device But it can probably cash like information and models related to generating text in specific languages and so on right Yeah, like maybe the most popular languages or the languages that you have set on your device so if your phone is in Brazilian Portuguese for example, then maybe it has that cash locally But then if you're asking for something in Swedish or Spanish or something it has to go to the cloud like I can see them Create something like that where perhaps From our perspective as third-party developers. We're just asking it asking it for a prompt or we're giving it a prompt But how it resolves that prompt is kind of you know dynamic right and We already have some APIs that work like that one that came to mind is the I think it's the Speech recognition API so you can have it basically transcribe some Audio for you and it can do that on the device But depending on the settings you will go and do it on the cloud on your behalf and it's exactly like I mentioned It's just a naysynchronous API and you don't even have to know if it's going to the network or not But you do have control there is like a flag you can set to say basically I only want this processing to happen on the Vice which I think they would probably do as well if they did something like you mentioned so that if I have maybe privacy Concerns or some other type of concern that I don't want this stuff to go to the cloud no matter what then you could Just say hey do this on the vice and and then of course it wouldn't be as good but You know what you're doing at that point and then you can opt into having it go to the cloud or not right exactly and maybe there's even something where Using Apple's favorite API the one API to rule them all NS user activity. Oh no they can do something where Perhaps they can like train Models based on app data app contents but in a completely like anonymous private way and then when you download an app You also download like an extension to that model specific for that app which means that you will be able to ask it things About that app without having to like train it yourself without it having to learn it like on device Re-learn it and so on like I can see a lot of kind of hybrid solutions like that where they could enable a lot of local AI features But without having to always go to the cloud for every single request. Yeah, exactly and Also something like what the rabbits are one device is trying to do not very Successfully at the moment based on what we've seen from the reviews and whatnot But the idea is very interesting where you train this action model and it can use Websites and apps without having Knowledge of PIs that you can call in order to do things so you could think of something like what if you could automate apps on your iPhone without the app having to explicitly support say shortcuts and things like that So basically what would be happening is the system behind the scenes would be actually opening the app and tapping buttons and doing stuff and you don't have to like Record a macro like you would do with some automation apps on the mac and things like that You can just have a model that knows how to use a user interface as a user would and you could automate things That way because the main issue with automation on iOS and iPadOS is that apps have to support it and Unfortunately the APIs to support it have never been that good and Now they are getting better. They're still not very good But they're way better than the List files like the shortcuts The finish on files you had to to use before But there's still a long ways to go for them to be something that Developers can just implement really easily and in order too much about Details which unfortunately is the case right now So if they could use a model that would just use the app as as a user that would be very interesting Yeah, absolutely and I think here again, there's some interesting kind of technological Opportunities where since Apple controls the OS and they know a lot more about an application than something like an rabbit R1 Which is just interacting through VPN or something through a website and clicking buttons, right? Which is like a very opaque way to train an AI like that But if you have like insight into what does the view hierarchy look like? What are some of the accessibility attributes? We have things like an as user activity like I mentioned earlier and you could even imagine there being some feature in Xcode where You as a developer could even train actions as you are like QA testing your app or developing the app and you could say Xcode just observe what I'm doing here now and just learn from it, right? So as you are like working on your UI or as a QA tester is testing the UI Then the model just keeps learning or the AI system keeps learning and then again It can kind of generate some extension file or something that can then be ingested by the system and then extend the action model to be able to perform actions for that specific application or it could be like crowdsource from users because Again, Apple has insights in not just at what point on the screen something is being tapped, but what is actually being tapped, because they're running the run loop inside of the application. They're running the event loop. They know what control is here, even if it's a completely custom control. It's going to be a SwiftUI view or it's going to be a UI kit view and so on. So I feel like there's a lot of cool things that could happen here and really enable some transformative things for users too, because like you mentioned, having to either write your own automation using a limited set of automation dictionaries or something or having to just record taps at specific positions and things like that. Those things will always be quite fragile and be just a secondary system to interacting with the app by itself. But something like this could presumably be just like interacting with the app just in a completely headless way where the UI is not being launched or anything to you as the user. You could just ask it to set a reminder or order something in an app that doesn't actually have direct integration with the personal assistant, but you're just doing it anyway through that model, which I think is pretty exciting. Right. Because if you think about it in a simple way, you could do it like you mentioned just like record a macro where just tap the specific points on screen where the user is tapping and then record that as a set of actions and then just repeat it later. But of course that's fragile. Like you mentioned because if a button was on the right and then it's on the left or there's a new button to the sides and that breaks very quickly. Then you can be more fuzzy and maybe like instead of using the position of controls, you actually look at accessibility attributes of those controls, but those can also change. So it's not very good to just do like a straight comparison. But even then like what if the app throws up a AppStore review dialogue in the middle of your automation? Is your automation going to handle that? So that's the sort of thing where having a machine learning model do that I think can address those tricky situations which are not even edge cases because like every app at some point is going to pop up an AppStore review dialogue. And if your automation can't handle that, it's going to break when that inevitably happens. Yeah, exactly. And you know, we've seen in recent years like the application of AI models when you have that additional metadata that we're talking about here can be just incredibly powerful. Like one thing that I really enjoy using is Nvidia's DLS technology, like deep learning super sampling for games where game developers they're integrating this technology into their games and it does upscaling using AI, but it's not just taking a lower solution image and just like stretching it up to a higher resolution is also using like game engine data like motion vectors and color data and things like that over time. And the result then is like really incredible. Like sometimes an upscaled image can look better than if you just rendered it natively because you also get like things like anti aliasing, you get like lots of artifact removal and so on. Of course, there are still some artifacts that can be introduced, but I would say overall even if you zoom in on a picture that's been upscaled, it can look incredibly good. And I feel like this is something that could also be applied in these other fields that we're talking about where when you combine the just like kind of learning aspects of like doing these many, many iterations and training these neural networks with that additional metadata that is coming directly from the application, you can just achieve like very, very good results with very few kind of false positives. Yeah, definitely. Cool. So that's a AI part of WWC anything else that you're excited about or anticipating? Well, that's a good question actually. There has been some talk about a more customizable home screen for iOS at least and that sounds interesting to me. I'm not big into super customizing everything about the device, but I think I'm interested in seeing what that actually means in practice and also if that happens, what we as developers will be able to do to facilitate that and maybe there will be a second coming of widgets that happened when they first introduced widgets and people went crazy with the widget customizations and things like that. So I'm really interested in seeing where that goes and I know that for example, the whole wallpaper system on both lock screen and home screen on iOS is based on extensions using the extension kit APIs. So I wonder would they maybe allow the developers to offer wallpapers to the system? I think that could be really interesting where say, Chibi Studio, you could create your own wallpapers using your creations in the app and things like that. So I'm interested in seeing how far they go with those customizations. Yeah, and I think a theme if you look at what's been going on in the most recent years, like the major features that have been introduced on iOS specifically, it has been a lot about system integration and it's been a lot about this idea that we've actually been talking about for years, not that we invented it or something like Apple, listen to stack trees and that's why we have all this stuff. No, not at all, but I think we've been really wanting the like the whole kind of operating system to become a bit more like different plugins and that's really what we have with this extension architecture where there's so many more integration points now into the system and you can really build an app where the app itself is mostly just a setting screen. And then it just integrates through widgets and all these other ways with the operating system to just provide functionality kind of in the background. And this I think is super cool and especially as we're moving more towards this kind of idea of ambient computing where you know, you have data being kind of fed to you without having to explicitly ask for it. So instead of having to launch an app and do a task, you can just do something like through the personal assistant on the device or maybe something pops up automatically like proactively, which we're already seeing with suggestions in spotlight and suggestions on the home screen and so on and smart notifications. So that I find very, very exciting and I think like having some some of that stuff also come to the home screen there with wallpapers like you talk about where you don't just you're not just able to vent like an image, but something more like with layers and different metadata that the system can use to render this in a very cool way. Like I think that all of that stuff is just really, really cool. And the more iOS I feel I can become this just like kind of series of plugins that all plug together and that's the end user experience. You get like really good customizability. You get the third party app opportunity to be more integrated into the platform and to be able to offer more functionality. I think it's just all great. Yeah. What's cool is that they've been developing this whole infrastructure for quite a while. At least since iOS 8 where you can do all of that stuff and still retain all of the privacy and security aspects of the system. That's why so many of these things take so long that there's both that aspect and also like battery life and things like that. So it's interesting that they could technically achieve that and still maintain all of the expectations about security and privacy and all that good stuff. Yeah. And this is where also we've talked about it before, but like the whole idea of declarative programming whether it's in the context of SwiftUI or whether it's in the context of you handing a function to the system that gets run that you don't just your app is not just walking up in the background and you just get your app delegates call back and then you can run whatever code you want. But something that is more constrained that has so much potential when it comes to this kind of architecture and we've already seen it being deployed for things like widgets where you know we this is you as the app developer, you provide like a timeline of snapshots to the system that the system can then render over time. And I can also see similar things with the previous thing we talked about with the AI stuff. Like if you're able to vend like a Swift function to the system that is not just arbitrary Swift code, but it uses some kind of DLS like domain specific language similar to SwiftUI to kind of build up building blocks around some kind of prompt or query. And then that can be kind of serialized and running the background whenever the system has some excess capacity and it can be delivered like in a very smooth way and very efficient way. Like I feel like that is also building blocks that are being being put together over a few years and also has a lot of potential and it's also very interesting from a technical point of view. Yeah. So from my perspective, I feel like I'm just going in with an open mind like I don't have so many expectations when it comes to like system features. Of course, I expect AI to be a big focus like we talked about. But more than that, I feel like I'm really hoping that this will be a year where Apple will address some kind of long running stability issues with their developer tools. And specifically what I'm talking about is things like the, you know, with Xcode and the integration with the Swift compiler integration with the debugger and so on. And you know, I had this joke that I used to make like it wasn't much of a joke actually. It was pretty much the truth. It was pretty much like every WWC I would say my wish list for this year is just, you know, an Xcode with a very stable editing experience and everything just works. And you know, that continues to be kind of. of my number one wish. And I realized that Xcode is a very complex beast. And it has probably a lot of legacy code in there. And that the teams that are working on this are doing the best they can to kind of make things better. And things are getting better. I think, for example, the new build system that was introduced a few years ago has been a huge kind of positive improvement, right? To the overall experience of using Xcode, like now the build tasks are a bit more predictable, faster, things like dependency management with package manager integration is better. Like things are improving. And we're getting fewer and fewer of those like source kit crashes. It's just when I'm like really deep into generics. I usually get them. But other than that, I usually things are working pretty well. But what I'm mostly talking about is, for example, I think one issue that's been plaguing Xcode now for at least like two or three years are these ghost errors? I'm sure you've seen them as well. Oh, yeah. I was experiencing this yesterday. And it's really annoying. Yeah. And it's like these things happen when Xcode will surface some build there that you fixed like two days ago, but it will still be in the cache. And even if you like clean drive data and do all the classic tricks, it still is there. And I know that caching is one of the hardest problems in computer science or caching validation, I should say. And this is definitely an issue of caching validation it looks like. But this is just one example of like the tools seem like kind of a little bit too decoupled in a way. And I'm not talking about like the good sense of decoupling where you have decoupled systems. But more that one tool doesn't know what the other one is doing, where Xcode as the IDE, like there's something in the build system there that hasn't really been surfaced up to the editor. And similar things with syntax highlighting an autocompletion, I still think could use a lot of improvements. And this is where I'm excited about all this AI stuff. And one thing that I am expecting also to come to Xcode, maybe this year event, is some kind of co-pilot feature. But I feel like the more they add these kinds of features, the more kind of pressure it puts on the core editing experience of the IDE to the point where it easily breaks. And we've seen this many times in the past where as soon as a new feature comes along, whether it's Xcode previews for SwiftUI or when Swift was introduced and so on, like Xcode stability takes a big hit. And I'm really hoping that they have not just added a bunch of features this year, but also worked on the core editing experience. Because I feel like that at this point, Swift is going to turn 10 years this year. I feel like this stuff just needs to be rock solid. It needs to be similar to when they had that problem with the butterfly keyboards on the MacBooks, where sometimes you type a key and you get two letters or no letter at all, that cannot happen. And I feel like the same thing is true for the IDE, where when you're writing your code, when you're building your code, when you're testing it, debugging it, you cannot have any false positive. So you cannot have any positive negatives. It's like it's all leads to be the reflection of reality. Like I don't want to see builds succeeded and then there are compiler errors because then I stop trusting the tools. And similarly, when the tools say that there are compilers that I fixed two days ago, also really erodes a lot of trust in the tools, right? So I'm really hoping that this will also be a year of stability fixes and that the core editing experience of Xcode will just keep getting better to the point where it feels like really, really solid, which I really don't think we're there yet. If I was in charge of Xcode for a while, let's say I've become the Xcode manager for a month or something, my rule would be you cannot add any AI to this thing until I can type and the letters appear with at least as low of a latency as Visual Studio Code, which is an electron app. And for some reason, sometimes especially in larger projects, just typing in Xcode is really slow. And like sometimes it will lag and it feels like it takes almost a second for a letter to appear. So please fix that before you try to integrate an AI into this thing because I feel like if they integrate AI into it, it's just going to get worse. Yeah, and that's the thing. That's what I'm talking about. Input lag is the most core of core features, right? And have you tried, there's a new text editor, which is called Zed. Have you tried that, Rambo? I have not. So I would say, I don't know if I would recommend it as an Xcode user because once you try it, you really notice how bad Xcode is in terms of input lag. Because they advertise it as being way faster than even Visual Studio Code, which we say even, and we're saying that's an electron app, which, again, a native app should be faster than an electron app, I would say. That's one of the trade-offs you make when you go cross-platform like that is usually things like native integration, such as input lag, suffers. But I think the team's working on Visual Studio Code have done just an amazing job because it's just an incredible editor, like it's just so good. And it's one of the electron apps where I'm not constantly reminded that it's electron because it's so good, you know? But that being said, like a native app, especially a native app coming from Apple, that also makes Mac OS, should just have so little input lag that it should just be the fastest editor on the market, right? Like that should be the gold standard, but it's not. Like Xcode's input lag is just horrendous, I would say. Like it's really, really slow, it's just rendering characters. And then again, that is a really, really core feature. Like we are excited about AI. We're excited about cool, refactoring features with animations that fold your code into origami. That's very, very nice. But just make the letters appear on the screen in a predictable way first, right? Like fix the core fundamentals first before you add more high level features is usually sound advice. And I think that's-- I'm hoping that this is something that's been prioritized by Apple this year because Xcode is a pretty good editor, I would say a pretty good IDE overall. It has a lot of good features and it's nice and easy and nice to use and pleasant to use. But it's again, it's the fundamentals that I think could use a really a lot of polish and optimization. Yeah, like whenever I'm doing one of these little prototypes where I just go file a new project and it's a small project, it works perfectly. Like it's really good. I really love the experience overall. It's when you get a real life project open that is start facing these issues. Although there is one aspect that I also wish they would fix, which is the way the vices are integrated into Xcode, which changed with the iOS 17, basically the iOS 17 train introduced this new core device framework and protocol where you can basically have your devices communicate with Xcode seamlessly, no matter whether it's plugged in or wireless, you don't have to configure something separate in order for them to work wirelessly, which I think is good. And also the experience of using devices wirelessly has improved tremendously with this new system. Yeah, it's so much better to the point where I'm actually using it now. Yeah, exactly. However, it's not always good. And the main issue that I have at the moment is that often, I will have my device plugged in, but it will use the wireless connection instead. So it's bad at basically switching from wireless to wired. And I'm not saying like in the middle of a debug session, like I wouldn't expect that to work. But sometimes like I have a test device, I plug it in, I build and run, and it takes longer than it should. And there's also another thing that happens with wireless debugging in Xcode where the first frame of the app basically on screen, it takes a few seconds to start responding. So when I want really fast iteration, I plug in the device, and sometimes it'll still use the wireless connection. So basically I have gotten into the habit of turning off Wi-Fi on the device, then plugging it in, starting in the debug session. And then after it's established, I can then go and turn it back on. And also because it's basically no longer configurable, whether a device should work wireless or not, all of your devices are always there in Xcode. And I have a lot of devices. I know most developers maybe don't have as many test devices as I do, but I do. And sometimes I notice that Xcode is doing like a lot of stuff behind the scenes with every single device you can find nearby. And I've even experienced crashes in Xcode where I look at the crash, and it's in the core device thread where it was trying to talk to my iPad. That's upstairs. And I'm not even debugging on that iPad. I'm not doing anything about that iPad, but it crashed because something broke in the communication. So yeah, I wish they would make that basically give us more control over this communication with devices. Yeah, absolutely. And I feel like sometimes also Xcode kind of falls victim to the curse of distributed systems, which is that Xcode is definitely a very distributed system. Like the core device framework you just talked about is definitely like a subcomponent. Even if I haven't seen the source code for Xcode at all, we kind of can see from the stack traces in the crash reports and just from how Xcode operates in general, like everything is its own little service, its own little sub system that is running. And usually the problem when you have something like that. You're working on a project that is distributed like that or very modularized. We usually talk about that from the good aspect. It's really nice to have things being separated out into separate modules and to have the interfaces between them be very well-defined and have separation of concerns. Those are all great things. But when we're talking about a big project that has worked on, presumably by a lot of people, you can sometimes end up with this issue where every subsystem is optimized to do what it does without much concern about what the other systems are doing. I talked a little bit about this before when we talked about the problem with the build system and the IDE UI and not being quite in sync. But you can imagine being the developer, for example, working on the device management code. It would be a great idea to just kind of in the background, synchronously, just make sure all the devices are up to date, download the symbols from them, ensure that they're set up for development so that when the user goes and select the iPad in the list and they press run, they just run immediately. That is a good idea in general by itself, to do things proactively to make things faster when the user wants it. But if you then put that into the environment and the context of all these other services doing things and all these other things happening, you end up with this fight for resources. Then all of a sudden it's not a good idea to do this work proactively, especially if the device is upstairs. I can definitely see how it's a challenge building something like this. You need to have good orchestration of tasks. I feel like in a project like this where it's almost like an operating system, where each app or each subsystem cannot just do whatever it wants. It needs to ask for permission, which as third part of developers on iOS, sometimes we find annoying. But it's a necessity when you're building something as complex as an IDE or as an operating system in order to have good scheduling of tasks. Yeah. In the case of Xcode, maybe this particular problem. Maybe even the editor lag we were talking about could actually be improved by making it even more modular, where you do all of this device management stuff on a completely separate process. A lot of it already happens on a separate process, but not all of it. There's still a lot that Xcode does within the Xcode process itself. The code editor as well. Everything is running in this shared thread pool. Maybe this should be a separate thing that runs on an extension. Of course, that can introduce its own problems. It's nothing is perfect, basically. Yeah. I've been thinking quite a lot about this recently, actually, especially in the context of Swift and the main actor, where I feel like one kind of blessing and a curse of apps on Apple's platforms is this like main thread synchronization problem slash good thing, which what I mean by that is that a lot of the things that happen in an application are synchronized on the main thread. Like all the UI runs on the main thread for the most part. You have the main thread usually hold a lot of data and state related to the UI, but also maybe coordinating access to certain resources like the file system or databases, not running the queries or things on the main thread, but actually just going through to make sure that there are not data races in there and making sure that everything is updated predictably in a predictable order. If you just analyze any given application of any size, the amount of work that happens on the main thread, I would say is quite large. Even if you're taking a lot of work to split things up on multiple threads and really run things in parallel, you still have that main thread that is like the core of the application. I feel like that's also an issue with systems like this where even if you do separate it out and you do run it in the background and you do maybe run it on a different process, you still have to go back to that main thread and jump over all the time to update the UI, to maybe check in from some other service and that needs to be synchronized to the main thread in order to be predictable. You end up with this bottleneck essentially with the main thread. I've been doing some prototyping recently. Maybe going to talk about that on a future episode, but related to this where I've been thinking about what would it be like if an application did not have that main thread concept, but rather was just completely distributed on different actors and different threads. That of course has other problems, but it's interesting that this legacy of the main thread on Apple's platforms, which has existed for years and years, is something that sometimes gets in the way of really trying to parallelize things and distribute them across threads and processes. Yeah, especially when you consider, which is not something you have to consider very frequently, especially for beginners, but I think every iOS developer or Mac developer who's been doing stuff on the platforms for as long as we have at least, you also have to consider that the main queue also has a run loop, which you do not get by default on background threads. Any other queue, basically, that's not the main queue. Sometimes there are APIs that will not work if they are dispatched to another queue. If you don't know about the run loop and you don't know that that particular API requires an active run loop, you are in for a lot of debugging because these aren't just not going to work. Yeah, exactly. And then you end up having to go back to synchronize on the run loop, and that introduces this bottleneck. And it's interesting because sometimes you can end up with, especially when you're working with things like Swift concurrency, when you're just dispatching tasks, and you are, you can end up in a situation where you think you're writing very concurrent code, but it's actually not concurrent really because it has to go back and synchronize on the main thread. And that means that even if you have a thousand tasks running, if they all need to talk to the main thread throughout the task, well, they're all going to be scheduled on top of each other on that main thread. So they might be asynchronous, but they might not be fully concurrent. And yeah, it's just kind of interesting to think about. Yeah, definitely. All right, Rambo. So we've talked about AI a bit on this episode so far. And you alluded to that you've been doing some work around this. And I know that we've talked about it before on the show, and you've shared some things with me as well. So I kind of know a little bit what you're up to, but I'm really curious to hear more. And I'm sure that our listeners are as well. So could you first perhaps just give us a bit of a recap around how you are currently kind of integrating different AI features into the projects you work on and expand a little bit more on kind of your overall strategy there, apart from the thing we discussed about earlier around, you know, perhaps using something that Apple will offer or something, you know, from OpenAI and so on. So tell us a bit about your AI adventures. Right. Yeah. So like I mentioned, Forchie Bisturio, we've been doing a bunch of work mostly around understanding the contents of images and also the opposite, which is generating images and even combining both where basically it's basically a fancy style transfer if you think about it. So there are these ML models called style transfer models where you basically take an image and you run it through the model and it speeds out the same image, but in a different style. So you can think like of training one of these models using paintings from a specific artist and then kind of like the texture of the painting is transferred to the image that you put in. So that's a style transfer model and you can train those very easily using like create a mail on your Mac and they run pretty quickly on the neural engine and they are usually fairly small. So that's very basic, but you can do something like that with something like the GPT4 Vision model from OpenAI where you take an image and you instruct the model to describe in as much detail as possible that image and then you add another prompt where you say, okay, now generate another image as described above, but in the style of 3D movie or something and then you can be as creative as you want with the description of the style and you basically get like in the case of Chibi Studio, we can send it an image of this 2D very cartoony kind of character that the person has created in the app and then have it describe the character like the the hairstyle, the clothing, colors, things like that and tell the chat GPT and then it will generate a prompt to Dolly which is yet another model but an image generation model and it will like create a version of your character that instead of being in this cartoon 2D character looks like 3D or looks like a real person or something like that. That's super cool and on a technical level then it sounds like you're doing something similar to what we discussed earlier which is like combining generative AI with known metadata because presumably you are generating these descriptions and so on based on the items that the user has picked in the UI and you know what those items are so you know if this is like pink hair or a black t-shirt or jeans and so on and you could then like generate the descriptions for that based on that is that kind of how it works. We can use that to augment the description so the main description is completely based on the image itself so GPT4 vision you can basically you give it an image file and and you will describe what's in the image. Of course, in order to get it to describe it in the way you want and with the level of detail you want and like what type of features do you want it to focus on, then you have to basically steer the model by basically telling it what to do, which I think I've mentioned here before. It's kind of like programming in a sense. I'm not treating the model like I'm chatting to someone in this case, I'm basically treating it as a programming language. So instead of doing statements and commands, it just phrases in English and they don't have to be in English necessarily, which is really cool, but in our case it is. And I'm also, like I mentioned before, completely abstracting all of this stuff and the user doesn't even have to know that this is using OpenAI or whatever. And that's also in preparation for maybe Apple does something like this and then we basically plug this into Apple's backend and the user doesn't have to know about it. However, we are doing some of what you mentioned where, yes, we know since you create your character by selecting from pre-made items, we know that there are items that this item is hair or is clothing and we also know which item pack it's from. So maybe that can hint to a theme if you have many items from the fantasy pack. Maybe you want the image to be more fantasy focused. So we can basically take those inputs and send them to the model in the way of keywords so that it can steer its description more towards those aspects. And one very important aspect that we found out is difficult to get the model to describe just based on the image is skin color. And that's, of course, very important because we allow, of course, people to change the skin color of their character. And what I found during initial testing was that the model basically ignores the skin color. It doesn't include that in the description. Even if you say very explicit, please remember to include the skin color of the character in the description. It does not do it. So-- And you even asked nicely by saying, please-- Please. Yeah, or offer it a reward. I'll give you $10 if you include this. That's actually a technique you can use with these models, which is crazy. But-- That's funny. I hope that they're not keeping track of that on the OpenAI side. And then they send you a bill for millions of dollars so you've promised the OpenAI to give them. That would be bad. Yeah, so basically what we do in terms of the skin color is we know what parts of what items represent the body and the face. And we know if you have customized the color. And if you have customized the color, we know what color it is. So that is also given as input. And that was the only way we found to actually get it to not forget about the skin color of the character. Yeah, that's really interesting. And it feels like that's a common theme with these AI models that they have a difficulty with for various reasons. Like humans, all the different shapes and sizes and forms and colors that humans take on. There's been so many problems with that with models already. And that's probably going to be a challenge going forward as well. It's really interesting to hear you talk about this. And this combination, again, of manually added data and derived from the model, not the AI model, but the data model in the app, right? Like your actual metadata. And then actually running the prompt, the running division as well. It's really cool to see all those things combined into a single product. And I'm wondering a little bit here. If I can take a little bit of a sidestep here, which is like, if you think about how AI might fit into software development overall going forward, not just like in terms of offering products to end users, but how we might interact with the as developers. I'm thinking that perhaps something like what you're doing right now will be something that we will do way more as our day-to-day work as developers. And what I mean by that is if you think about something like a SwiftUI view, for example, right, which is where you just declare I want this to be a button, I want this to be a text, and so on. That is a lot more higher level than if you would write the same thing in UIKit with kind of imperative commands, or you would say let UI button, and you will set the constraints on it and so on and so forth. And I feel like AI is one step further, or perhaps multiple steps further, up that abstraction chain. And perhaps we will have something in the future, perhaps sooner rather than later, where we can describe the UI so we want to render, for example, with natural language, but then interject specific commands, specific programmatic code into that that would then do specific tweaks. So if it's able to generate a list from a JSON file in a very straightforward way, but you want to customize how the rows look in, what the disclosure indicator will look like or something like that, you could then inject that specific override to say, actually, I want the disclosure indicator to be exactly this UI code, and you would then overwrite that. What do you think about this idea? Do you feel like that's something that could actually happen? Yeah, I do think that it could happen. And it is something I've done using Visual Studio Code when I'm doing JavaScript or TypeScript stuff, where I use call pilot. And it's not uncommon that I will think through a problem. So basically, I need a function to do whatever. And then I think about the problem in the way I would think about it if I wrote the function. And instead of just thinking about it, I actually write it down as a comment in the JavaScript file or whatever. And then I press Enter and call pilot writes the function, basically the first version of the function, because it's rarely the case that you can just use its code as is, but frequently it gets you going in the right direction, especially if you're not as familiar with the environment and the programming language. I am fairly decent with TypeScript and JavaScript, but there are still some things I don't really know how to do sometimes. And then I can just ask call pilot. And it's basically the same thing I would be doing if I was using Google or the go to search. But I feel like it keeps me in the flow of coding, because context switching is really bad for me. I don't really do well with switching between a lot of different things at the same time. So just having to switch between my edit or my ID and Safari or something to Google or search for something, that kind of breaks the flow a little bit. So being able to ask the question in line and have an answer in there, or if it's more complex, there's even like this little sidebar now in VS code, where you can basically chat with call pilot. And then you can give you code snippets. And oftentimes, if I don't like the initial offering that it renders for me, I can delete that and give a little bit more guidance. Like instead of, oh, I don't want you to use this API. I want you to use this other one instead. And then you can do it with the API you mentioned. Or you write parts of the code and you can fill in the rest. And especially with repetitive code, like sometimes there's code that's just repetitive. It's just boring. It's just a bunch of code that you have to write, like a bunch of boilerplate. And I find that co-pilot is really good at filling that in. Like you give it two or three examples, and then you can do the rest for you. So it's basically a very advanced version of auto complete. Yeah, exactly. I agree. That's really, really useful right now. But what's interesting is that that is very much like an offline process that happens during the development. And the artifact of that is still source code. But what I find interesting to think about is some potential future where the artifact is not source code, but rather an actual prompt that you will check into your GitHub repo. And that will be run at runtime and resolved by the system. And it can include like mixes of like either imperative or declarative code, but it's like a combination and that gets resolved at runtime. And that sounds I think kind of scary when you think about it. Like you just like ship a prompt and hope the best will happen on the user's device. But that's kind of what we're doing with SwiftUI, right? Like you are not coding the pixels for this button that you just created. You are just saying button. And then what will actually be rendered is up to each user's device. And the frameworks running locally on those devices. So again, like this idea of having an AI kind of built UI that you've just described locally as a developer, not just with something like co-pilot, but you've actually like shipped the prompt. Like that I feel like is like just a few levels up on that extraction chain. And it's just kind of interesting to think about that maybe that's going to be the way we program, especially UI's like in a few years from now. Like it's kind of wild. Yeah, now that you mentioned one place where I feel like that could start out is something like shortcuts, right? You have people like VTG making these really advanced shortcuts to a point where he's basically a developer right now and his programming language of choice is shortcuts. And there is no built-in way in shortcuts right now for you as the creator of a shortcut to basically create little pieces of UI. There are certain actions like you can ask for text input, you can ask for photos, but you can't like create a form with check boxes and pop-ups and things like that. So I wonder if that could be the first place where this sort of thing could show up where you basically write a description of the UI you want in within the shortcut and what inputs and outputs you want from it. And at runtime it basically generates the UI for you. Yeah, absolutely. I could totally see something like that happening. And I want to emphasize that I it's not that I'm saying that I'm hoping that this will happen because I do also find this scary. But if you think about it, if you chart the progress of software development from the let's say 1950s till today, what is kept happening is that you leverage some kind of system more and more and you work at a higher level of abstraction. So very few people these days are writing like assembly code directly and fewer and fewer people are writing like lower level languages like C, for example. And a lot of development has moved to more higher level languages like JavaScript, like you mentioned, or Swift, or Python, or other things that are like more high level running in some kind of either like a runtime or some kind of managed environment or it compiles down to machine code or assembly code or C or something like that. So, you know, if you look at it from that lens, like writing a prompt is really just like a different kind of compiler that is being run through, right? And you're just getting a different kind of result and you're working at a different kind of higher level. And I just find it fascinating to think about that, you know, especially when we're talking about what will the future programming look like? Like, will it be just like increasingly higher level languages or will it be more like declarative programming, more like things like Swift UI and and Jetpack compose on Android, or will it be more towards like the the AI side of things and with prompts? And yeah, I really don't know where the future will go, but it's going to be interesting to find out for sure. Absolutely. And I think that's a good place to wrap up this episode. So thanks so much for listening, everybody. And we will try to talk to you again soon. We will probably do an episode like right after WWC, right, Rambo? Definitely. We have to do some kind of reaction there to see what actually happened and how much of AI did we actually see? How many times did Tim Cook say the word AI throughout the keynote? That's going to be the most interesting to find. Yeah, from the recent iPad event, I think it was like 10 times. So not them by many, but it was just like 30 minutes. So we'll see. I suspect at W will be way more. Yeah. And I mean, just think about how many times we've said it throughout the episode, right? Like it's it's one of those things you can easily do super cuts of, right? We're just like AI, AI, AI, it's all very funny. So we'll talk to you then. And until then say goodbye, Mr. Rambo. Goodbye.

Podcast Summary

Key Points:

  1. The hosts explain the podcast's irregular schedule due to treating it as a hobby project, focusing on personal projects and programming topics rather than news.
  2. They discuss their anticipation for WWDC 2024, particularly regarding Apple's AI announcements, as they are waiting to decide on using Apple's built-in AI features versus third-party APIs.
  3. Key considerations include risks of relying on third-party APIs (cost, stability), the potential for on-device AI from Apple (privacy, cost benefits), and hybrid cloud-local solutions.
  4. They explore possibilities like Apple subsidizing AI APIs, on-device action models for automation, and leveraging Apple's OS-level insights for AI training.

Summary:

The podcast hosts, Johnson Dell and Mr. Guillrambo, address their long hiatus, explaining that the show has become an irregular hobby project due to busy schedules and a desire to produce quality content. They reassure listeners that the podcast is not canceled and will continue with episodes on personal projects and programming topics.

The main discussion centers on their expectations for WWDC 2024, especially Apple's AI initiatives. They are hesitant to build features using third-party AI APIs due to risks like cost, instability, and rapid changes in the AI landscape. Instead, they hope Apple will offer competitive on-device AI solutions, emphasizing privacy, lower costs, and integration with Apple's ecosystem.

They speculate on hybrid models where AI tasks are processed locally or in the cloud, and discuss the potential for Apple to subsidize advanced AI APIs for developers. The hosts also explore innovative ideas like action models that could automate app interactions without explicit API support, leveraging Apple's deep OS integration. Overall, they are cautiously optimistic about WWDC, planning to align their development strategies with Apple's announcements to create sustainable, cost-effective AI features for their apps.

FAQs

The hosts treat the show as a hobby project now, focusing on personal projects and programming topics rather than news, and life gets in the way with other commitments.

No, the show is not canceled. The hosts will officially announce if it ever ends, but for now, they plan to keep going irregularly.

They want to see if Apple announces built-in AI APIs at WWDC, which could be more sustainable and cost-effective than relying on third-party APIs like OpenAI.

Risks include APIs becoming expensive, shutting down, or changing terms, similar to issues with Reddit and Twitter APIs.

Apple's AI could run on-device for privacy and lower costs, and Apple may subsidize cloud-based AI APIs for developers, making them more affordable.

Developers might use a single async API that resolves AI prompts locally or in the cloud based on the query, with options to restrict processing to the device for privacy.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.