Go back

#27: Recommender Systems at the BBC with Alessandro Piscopo and Duncan Walker

87m 44s

#27: Recommender Systems at the BBC with Alessandro Piscopo and Duncan Walker

The transcription discusses personalization and recommendation systems at the BBC, a public service media organization. Key themes include the need for personalization to help users efficiently access content, while maintaining impartial, high-quality outputs. The BBC faces unique challenges balancing algorithmic recommendations with editorial values, as content appropriateness can shift with current events. Experts argue that scientists should not decide the trade-off between diversity and engagement; instead, product domain experts should guide these decisions. Collaboration between tech and non-tech teams is vital, requiring common language and understanding. The BBC offers numerous digital products—iPlayer, BBC Sounds, news, sports, education, and children's services—with recommendations active in iPlayer, Sounds, and sports, while news recommendations are paused. The personalization team, founded in 2018, built its first in-house recommender for short video clips, emphasizing a cross-functional approach with editorial and product roles. Guests Alessandro Piscopo (lead data scientist, background in arts and classics) and Duncan Walker (principal data scientist, background in particle physics) share their career journeys, noting serendipity and purpose as drivers. They stress that personalization connects users with relevant content, especially given the BBC's vast daily output of over 1,000 items, while aligning with public service values.

Transcription

14007 Words, 77928 Characters

English
Personalization is expected by user persuasion, you know, allows them to get what they want in the small amount of time they have. Outputs should be impartial, high quality and distinctive. You don't want to end up in a position where you have thousands of metrics at the point where you have effectively no metrics because you can't look at them and aggregate them and understand them. But it is about using a number of touchstones to tell a story about what's going on and to really get a better list to understand. It's tricky, right? It's like trade off between trusting in the statistics, but also this kind of softer understanding of what your algorithms do. I don't think we as a scientist, we are not the domain experts, we are not declaration experts. We shouldn't decide what the good balance between diversity and engagement should be. It should be those who are the experts in the product, those who are also set the kind of product roadmaps and say, okay, they go all the way to the ground. There are things that may not be inappropriate today that might start to be inappropriate tomorrow, particularly with current events breaking news and so on. We could learn how to work alongside non-tech people. Just you have just the entire people there used to work to do manual curation. And sometimes you have to find common ground, you have to find language, you know, that both sides can talk and speak. [Music] Hello and welcome to this new episode of RexBurts, Recommender Systems Experts. For today's episode, I have invited two guests to the show. And these two guests are working for the oldest public service media organization. And some of you might already guess what we are talking about. And you are right. It's the British Broadcasting Corporation, short the BBC. And for this episode, we are going to talk about how the BBC does personalization, where the BBC provides recommendations. We will also talk about editorial, grises, or with algorithmic recommendations, and also about the challenges of providing recommendations as a public service media organization, which is a bit different from private media organizations. And for this episode, I'm very happy to be joined by two experts from the BBC. My first guest today is Alessandro Piscopo. And along with him, I welcome Duncan Walker to the show. Hello. Hello. Thank you very much for having us. Yeah, nice to have you on the show. And I will soon hand over to you and first start with a short introduction of both of my guests today. So Alessandro Piscopo joined the BBC as a data scientist in 2019 and is now a lead data scientist in personalization and search. He originally was coming from a totally different direction with a master in arts and classics. And later on, moved to data science, obtained his master of science from the University of Amsterdam. And later on with a PhD in computer science from the University of Southampton, working on Vicky Data. He is also a core organizer of the queer workshop at Rex's and sick IR. And he has also published papers at the Rex's conference. My second guest today is Duncan Walker. He joined the BBC in the same year as Alessandro, so also in 2019 through a R&D graduation scheme and later on transitioned to working on recommendations research. And today works on recommendation systems in production at the BBC. He has obtained his PhD in theoretical particle physics from Durham University and is today a principal data scientist in the I player recommendations team. This should just provide the short introduction, but I guess both of you can say a bit more about yourself. So please share with our listeners what makes you excited about recommendation systems and what made you end up in that domain and working for the BBC. Maybe other some draw if you could go first. Thank you. Wow, that starts with a big question. I would say just what makes me excited about recommend resistance, especially the BBC. The first thing, well, just possibly just it's not needed. This podcast said it's a lot of data available, not many other content and especially BBC, we produce kind of more than a thousand newbies of content every day. And so like there might be lots of things, you know, treasure troves, you know, very interesting things, you know, very engaging relevant stuff for users, our users in our case, but for user, I think this is one of the cases where technology can really help, you know, to connect, you know, people who have limited time with the right stuff and you know, use their time the best they can. And finding interesting stuff that they will found otherwise, and especially in the cases of BBC, it's also about, you know, doing that in a way that is aligned with public service values of which I suppose we'll talk about later. So I think it's, you know, one case with technology can really help and, you know, just give users a better experience of online media content. What made you changing the direction of your career because you first started with that master of arts and classic and have been working in a different direction and then later on you went totally into these data science direction. So what was maybe the inflection point there or what made you going into that direction? Yeah, and it is kind of a seven deep it is, you know, choice. Well, first of all, I was after my master's and some other study that I made in anything and publishing I worked for two years in publishing houses editor. The issue with that I was just working on textbooks and it's always felt like I was just missing something about, you know, text was very important, you know, people just students learn from those. But at the same time, how just could I influence, you know, like the future of people, you know, what is it that actually is being built now is happening now and you know, and where can I make a difference. And at the time, didn't yet, you know, materialize into all that that's moved to the term data science wasn't really used when I moved to that field. The actual thing that made me move was actually as always, possibly love because my partner was such I moved to the Netherlands and as you can imagine, you know, there wasn't much to do for an editor native Italian speaker in the Netherlands. And after some times I found them kind of a scholarship, which allowed me to do a masters in an Netherlands and I found that very interesting, master in information studies, which sounded like techie, but not too much. I say, well, information, you know, not too far away from, you know, what I was doing and actually it was indeed. And I had to learn programming and just months, et cetera, but it was interesting. It was exciting. And so, you know, then it came a PC that I did actually have after having said, OK, starting that's enough. I'm not going to, you know, to continue studying anymore and then you know, end up doing a PC. And that's how I got here. And why was it to BBC you chose as your employer? Oh, well, that's not good. Another kind of certain deep just moment, you know, sliding door moment. Actually, I mean, it is just as I'm saying that because I am at the BBC now. I always to one person who used to be a data scientist in our team, Maria Panthelli. And if you're listening to us, Maria, thank you again, because I think she was already working there. She was a friend of a friend I was talking to her. She was the second time we met each other. I was saying, Oh, I really would like to just work in data science and work as a data scientist. But I would like to work for any company. I would like just to have purpose and, you know, an organization like the BBC will be there. Like a good trade off between maybe a large organization and at the same time having, you know, values. And she said, Oh, I'm actually the scientist at BBC. And say, Oh, wow, that's great. And after a while, she contacted me and said, Look, there's a new role that has just been open in my team. Why don't you apply? And that's how I got where I am now. Cool. And now five years later, you're still with the BBC. And I guess I've collected many experiences that we might then talk about throughout the episode. And I already like to things or, yeah, two words that you brought up there, which are serendipity and purpose. So I guess two things that somehow reflect about or reflect the moments or your personal career, but also something that we can. Yeah, relate to recommender systems in many ways. So seeing what is a BBC is successful and provide. also its users with serendipitous moments and then actually how you translate the purpose or the values of the BBC into something that recommendations or recommendation systems are going to support. So yeah really excited about that discussion. Hey Duncan I mean I've already also introduced a bit and provided some background about you but I guess you are the better person to talk about yourself. So can you also quickly introduce yourself to our listeners and maybe share with the audience what is your excitement in recommendations and how you yeah ended up in the recommendations team. Yeah absolutely. I think it probably made sense for me to start sort of coming to the end of my PhD which was in effectively large scale simulations of what goes on at Hadron Colliders. So you're running huge amounts of 4Tran code to reduce your arabas by very very small amounts that is taking 10 years and at some level it's a really really interesting engaging thing to do but it was getting to a point I think towards the end where the whole idea of generating these theoretical predictions is that you compare them to data and if they match fantastic you know what you're doing if you don't match if found another particle or discovered gravity effect or something like that but then on a personal level I looked at the timelines for the new experiments that might come out and it would be 20-30 years down the road to build a new tunnel in Switzerland for size of however large and to tell you what new experiences and your experiments to compare against. So at that point I was sort of ready to move elsewhere but also having enjoyed the simulation modeling component of what I was doing and at that point I found effectively the BBC's research and development grad scheme with sort of a vague purpose of knowing that I wanted to do research or modeling work but not I tend to find a kind of person who tends to find very interesting things no matter what I do if I do it for long enough you'll find that the nuggets are really really engaging and the nice thing about that scheme is essentially you do certain amount of time in a number of different teams and you sort of pick up those experiences so it seems like object detection for cameras going on wildlife ships was one of the projects I did but the final one was on recommendations on personalization sort of looking a bit more to the longer term it was a project that involved some level work with University College London in a joint project through a data science research partnership that was running at the time and yeah I found that nugget of things that I really enjoy and just kept at it really and then transition to working in a bit more full time a bit more permanently as you do as part of these schemes but I think there was always at the back of my mind the kind of the itch where you know recommendations it's a whole field it's an interactive field and you can only do so much by offline evaluation without seeing how your what you want to do works in a real world it's that difference really between the academic side almost and the and the industry application of it how it survives contact with reality almost and yeah at that point I saw those opportunities going in in the product side to work on recommendations and I took it really and I guess well that was about 18 months ago 15 months ago and here I am. So it's basically that you're within that R&D Graduation scheme with which you joined the BBC that you basically got in contact with different use cases of data science or different let's say domains and then you basically found that recommendations and personalization was that feel that drew your most attention or how can I understand it? Yeah absolutely I was kind of aware of data science generally machine learning also it was in use in certain areas of visits but I'd never really sort of coming to contact have that personal experience with it and so yeah it was that real industrial application of it so we have this need for it we needed to do this one thing yeah and I don't know like in some ways I'm a very simple person I enjoy something I want to do more of it that's usually my motivation for most things in life. Just with the difference it nowadays I guess Fortran code plays less of a role as it did a couple of years ago. I have to stay quiet about some of the languages I know sometimes. That's great yeah cool so first of all thanks both of you that you joined this interview and that we now today get the chance to share with our listeners a bit more about the BBC and recommendation systems at the BBC so and with that I would say that's dive into it and first maybe I would like to get a better understanding of what is actually the offering of the BBC because I've seen so far that you folks you have something like that is called the eye player which for me seemed like oh this is some kind of Netflix by the BBC and you're so of course to have on BBC.com news recommendations sports recommendations all of stuff you have some audio offering but can you help me and I guess you can to make structure of all of these things so what are the offerings of the BBC and where in those offerings does personalization play a role and where do you use recommendations. That's another very big question just one thing I would say Netflix is a type of I play rather than just look so as I said it's a tough question because also just being at the public service and we can come to what that means later I think the BBC is very large much larger than some people expect especially outside the UK so just while everyone knows just we have news you know where it can be usually just one associated to the news website or the app so just that's possibly just the flagship service we have there is very cheap sports so just we have textual and just possibly also video content in that's of multiple formats already. We also have in terms of digital products because let's say possibly we could start talking about those we have a video in the month platform which is I player we have an audio in the month platform which is BBC sounds which includes podcasts it includes music but not as kind of single kind of songs it's in the form of music mixes you know created by DJs and and also live radio just you can access live radio from that so beside those that are possibly the most well-known there are lots of other websites I definitely may some but just there's the lots about education like this BBC bite size which I think possibly don't know more about it because it's grown in this country but it's a great resource for students from just early age to you know to learn prepare their exams at school there's again the whole children services just we have CBBC so there's also the whole site for children's programs and and then services what has in there this BBC food well there's a word service which is very important more services I don't know the terminology here it might be another division but it's just a whole part of the organization where BBC produces news and little several languages across the world and I think it has a very large audience which might get to half a billion people just accessing sites every week no 400 videos something like it's very large the BBC is many BBC again but we can come to that later but there's just BBC produces content and distributes produce and distributes content in many formats the many tabs you many topics etc and so far we we have developed recommendations for also be the the most well-known of those products and which are like a player for instance which are on sport which are BBC sounds we have content to content recommenders in place live and many more service sites we used to have some or content to recommend they're on the UK news app but it's been taken offline for the time being so we do recommend this is for we personalize lots of you know content and lots of those services but did the there could be many more we could expand to one day and yeah yeah okay so if we take the very widespread offering of the BBC then what I'm getting from this the main surfaces or offerings where recommendations play a role are basically the iplayer sounds news where it's currently not running but when you say sports are we then also talking about sports news or about something different or what do you mean when you say sports I think it's BBC sport is news about sports yeah updates about any sports there are live kind of livestreams for instance when there is a football match or rugby match or any sport I mean cricket I don't know enough about cricket to know if you can have a live stream because as far as I know it might be very long you know a cricket match but Oh, okay, Duncan is nodding. Yeah, I think it's these things are these sites are definitely quite integrated. So, you know, the audio livestream, your eight hours day of cricket from a match will be part of the sport offering. So it's not just news necessarily. It's a bit more integrated, if that means. I saw that there was one of the papers that you folks published was actually in, I guess it was 2021 and I have it open here, even though I have to confess I haven't read it. Building public service recommenders lockbook of a journey. Can you take us and the listeners maybe on that journey and guide us a bit through where did you start with the recommenders and how have you maybe expanded into different domains and where you know are days? I don't know if mentioned already, our team was created in 2018. Initially, I think their remit team was a bit broader because the aim was to have a team very much focused on machine learning that would develop machine learning in house but not as an R&D team, but you know a team that will be engines to be deployed various services of products. I'm using this term interchangeably but by service of product I mean something as news or sport or I player or sounds. So something like customer facing? Yes, yes, not necessarily because at the beginning it was thought to be you know machine learning, I thought just kind of posturing the use and the adoption of machine learning. I think very quickly around 2018, the team worked on its first kind of use case which was a recommended, the first recommender that was deployed and built in house. So the first kind of example of algorithmic curation within the BBC and it was a recommender. I joined after that but the first I'm aware was recommending clips. So short videos to users based on their previous interactions and it was deployed in an app called BBC Plus which has been discontinued. I don't think it has just a very large number of users but it was a way to kind of you know made the foundations for the team. You know from the very beginning it was a cross-functional team so there was this idea that we would have not just data scientists and software or data engineers but there were also first product managers but delivery managers editorial which has been I think super helpful for us. And having all these people just working together and just working on a concrete use case actually enabled the team to identify all the problems, all the challenges that were connected to building recommenders for a public service organization that's which kind of basically just comply with and follow the editorial guidelines and possibly support a public service values and even in terms of the practical approaches to do that sometimes you know just thinking about oh let's come up with a simple approach to work so you know what is the data that we've got? How good is the data? You know just problems that's hadn't been addressed previously to build a recommender system in house. How has it continued from there so was it then that you just had your first recommender and try to extend it into different domains or you tried to improve it beyond what it was like at that time or how was the journey from that point continuing? Before I'd turn the question I'd like to add something about why that first experience was important. One of the key things was that we could learn how to work alongside non-tech people. Just you have just the tutorial people they are used to work you know to do manual curation and sometimes just you have to find common ground you have to find the language you know that both sides can talk and speak and you know can use and so you know how do you just phrase things as a data scientist that's you know are clear and can be acted on by someone from a tutorial and how can you just express the fact that well you are used to maybe arrange like a base of content you know like few dozen piece of content maximum at the same time just how do we start thinking in terms of thousands of possible piece of content and then on the data side is understanding what are the problems that might arise because we are lucky everything has been kind of checked just produced you know a tutorial just guarantees that calls you to content is good but then you can have issues like you know two basis because content recommended together and having those piece of content together might suggest kind of inappropriate I don't want to get into examples because then but you might have you know something just associated like some some sensitive topics you can have an article about a murder and then something that is funny it's a matter of all sort of taste and the organization it's just there's a reputation of risk that's what we learn about where just you know the the the journey from that first recommender to the subsequent projects I think that recommender was almost abandoned very early on we didn't iterate on that also because as I said you know the number of users you know all that app were very 100 it was very small but the team I think it was you know when I more or less when I was high you know expanded we started you know being getting into a squad structure so we had multiple squads we started being a lot more just more aligned with single products so there was a squads that from very beginning was working alongside BBC sounds to develop a recommender for them another squad was aligned to new sports and the board service and I think initially was these two squads plus an engineering squad that was working on the infrastructure I'm getting also 2019 is a long time ago for me but yeah I think at the beginning we're like three squads and and that's where we started and I player got into the folds much later just you know the join our team was 2023 I think it was around February 2023 so that means that the I player team or the people working on the I player like Duncan were then becoming part of the sounds or also a dedicated team that was working on the I player later on well I guess I joined sort of as part of this as this I play team was sort of being developed there had been personalisation on I player prior to this but it was not necessarily built so much as a data science environment and so it had kind of been something that worked technically but maybe hadn't been iterated on and so on as much as it potentially could have so I guess the last year it's been a case of folding that in trying to apply the knowledge that has been obtained in some of the other teams to the I player experience in trying to to start iterating on version two some version three okay before I would like to dive into more of the specifics of the recommender said you have nowadays in place and that support your users across the different domains we already laid out I would be a bit more interested in that topic that arises in kind of every paper that you folks have published which is around public service values and that providing recommendations in the context of a public service media organisation likes a BBC is a different challenge than if I would do so as let's say a private streaming service or something like that because there is that additional dominating public mandate that you folks have to follow or to fulfill can you shed some light on what this means and what this is all about and show given those those values how this actually translates into your work on recommenders okay I could start a little bit about what public service means so possibly this is something which UK audiences are more familiar with but you know the constitutional basis for BBC is the royal charter which is something where that walks seal actual and I think it's signs by just be the queen even the next one will be by the king and it as I said it's a constitutional basis it's as hard like the BBC object mission and public purposes sets out also other things but I think those are just the main one you're that interesting to talk about now these are these mission and public services are kind of describe what it means for us to be just a public service organisation can quickly read kind of our mission to just make it a bit clear and it is to act in the public interest serving all audiences through the provision of impartial high quality and distinctive output and services which in for education entertain I think one key bit of these is the fact that outputs should be impartial high quality and distinctive. And let's remind ourselves that recommendations just generate or kind of determine our outputs, you know, what puts in front of our audiences. It's not just the content, it's just how it is distributed, how just about users see when they open a player or sounds or sport in depending on where they are on the site. So this is already kind of suggesting that recommendations should follow like those publics or the dismissation. But also it's about, you know, that the public service purposes, the public purposes, which is, you know, providing partial news information to kind of help people in the sending engage with the world around them, support learning for people of ages and kind of show the most creative, highest quality and distinctive outputs and reflect represent and serve diverse communities, the United Kingdoms, and well reflected United Kingdom culture and values to the world. And I think here possibly just, it's good to stress the fact that we want to represent and reflect the diverse communities within the UK. So diverse is just part of our commitment. And so just giving what I said earlier, that's, you know, like recommendations determine and are part of our output, those public purposes should apply should be followed by our recommendations. Or it should be supported. That is a challenge because all those things, but our opinion to interpretation in sense is, when you get to the implementation of those, how do you do this? Just what does that mean in practice? And how do you implement those? Because again, when it comes to, you know, front page, you know, on page of BBC News, editorial just could manage, could curate the page manually. And they could say, okay, just we have this piece of content, let's balance with these other people content, don't put, you know, those two things next to each other because it's in appropriate pairing. Don't do it. Just how do you do it and control that that happens when in theory, everyone, every user could have their personalized homepage. Right. So how do you scale basically? Yeah, that's a very good question. I think just there are two sides of this approach of these kind of abiding by our public purposes. When it's kind of active, there is kind of, you know, support those and make sure that everything is reflected. And then when it's passive, which it's basically don't screw up, just ensure that there's no reputational risks. And we had ones, you know, like editorial sharing the different things they have to worry about. And there's of course accuracy and relevance of recommendations because, you know, high quality outputs, but there's also in personality, harm and offense, difficult content. You know, you could have sensitive content about self harm, suicide, graphic content, diversity, then, you know, you have protected characteristics you want to show different points of view. So, and sometimes, which is, first of all, there's the approach of limiting the risks. And so we have a whole set of business rules that reduce the risk of having inappropriate pairings, that prevents anything that's but to make as liable just for breaking any any loan, you know, for any legally relevant action or, and there's the other side that is basically we have process in place to have editorial to review our recommendations. And right now, just you ask, does it scale well? Well, we do have just a very somehow manual process in which editorial select very often and at the very high level, then it happens in different ways in different products, but editorial selects possibly different piece of content that are representative of our catalog and, you know, what we can also create personas. And then we generate recommendations and we have tools that we have developed in house just help us visualize recommendations and help us gather feedback. So, editor can, you know, say whether there are inappropriate pairings can add some, you know, text, whether it's some fields. And then we gather all the feedback and then we have regular meetings where they explain us. So, there's a lot of manual intervention in that in that, and we as a data scientist, I find it very exciting actually, we have to translate it and say, okay, this is a very weird thing that we need to avoid. How do we, our instruments, we, what we have available, we the tools we have available. How can we prevent that not happening? And sometimes, you know, like we, you know, block least some content because there's no other way to be completely safe. And at the times we do something that is forced nuanced. Duncan, do you want to add anything about that? Yeah, I think there are maybe a couple of things as well to mention here in that, a lot of these editorial priorities, they're not necessarily static things. They will change off the course of a year, they'll change off the course of the lifetime of the BBC in broader terms. So, there are things that may not be inappropriate today that might start to be inappropriate tomorrow, particularly with current events, breaking news and so on. Also, thinking about things like elections in the UK, there are particular constraints around election times, particularly regarding content that covers politics. So, these things are also in evolving, how do I put it? And we're working on an evolving system of requirements. And as a result, we also, you know, obviously we talk in solicit feedback and have this ongoing dialogue with our editorial counterparts, partly to understand what the rationale is, can we generalise and apply these, these learnings are more widely. It's not just saying that this article, next to this article, is inappropriate, is what's the wider learning that we can generalise from that. But we also have a lot of manual overrides and sort of not emergency features, but the ability to turn off certain recommendations quite quickly in the case of a problem has been observed that is sort of editorial inappropriate. And we have to kind of work in that and that might be at the rail level or a particular part of the product, it might be wider, it might be the entirety of the BBC, thinking about, certainly, through some. It's not even a static target, I think, is what I'm trying to get at. And as a result, you have to be pragmatic around that as well, so that you can act quickly, debrief and then solve the problem in a more general way, given the lecturing is a bit more time. The first time I hear this, it sounds for me like, this is a very manual cumbersome work as well, also some work where you could learn a lot because your editors can support you in better understanding what is appropriate and what is not and help you in translating those values that you just mentioned, Alessandro, into something that you can then encode into some objectives or some rules or something like that. And then you have that constant, let's say, back and forth process in which you would like, for example, to develop a first model, generate recommendations from that model, for example, for different seat users or seat items, depending on what kind of recommender it is, and then offer them or let them have revisited by your editors and then use the feedback to iterate and go back and forth and back and forth. What makes me thinking a bit is, I mean, in recommenders, of course, we have learned over the course of the past year set, relevance and accuracy are not enough and should not be the only goals. However, they are more or less easy to evaluate on a large scale, given the feedback of users. So, for example, in the domain of videos and Duncan, maybe you can, can shed some more light on that, taking into account how much users are completing a video or something like that. So, for example, there is the German CTF, also a public service media provider, and they have offered model cards, shedding some light on their models, and one of the models that is responsible for recommendations, for video recommendations there, they report that they actually take into account those videos that have a 35% completion rate. So, they are very specific about what they use as input data for their models, and to go back to that relationship between us, the, let's say, technical or algorithmic experts and the editors as, let's say, the content experts. What are the up and downsides of this process in terms of iterating quickly on models, but also in terms of maybe evaluating how effective your recommendations are because I mean, if there is a whole bunch of business rules coming on top of something that is algorithmically created, it can also be quite a challenge to actually evaluate what your recommender is doing when there are so many effects taking place afterwards because, and you never really know is what I'm evaluating actually what I'm finally also showing to the customer and how good am I actually approximating what finally happens in reality. So, maybe a bunch of questions, but can you maybe also with some example, it's deeper into that collaboration between editors and you and what this means for different things like evaluating the effectiveness of recommenders how fast and quick. quickly you can iterate and so on. - Yeah, it's a really hard thing to do, right? Because it depends on how you view it is, fundamentally what your model is. And there is a sense in which your model could just be the machine learning component, but in actuality the model that interacts with society at large is this model with business logic in it as well. But not only that, it extends a bit wider as well, particularly when we're working with curated groups of content that are editorial colleagues who know the content very well, say this content belongs together in this category and then we recommend from that. How we also understand the human and the loop component of what we do, where editorial colleagues will change their behavior, obviously dependent on what the model's doing, and therefore kind of start to build part of the complex system. That is a recommended system and to end. And it's something that we are trying to improve in terms of how we evaluate things both online and offline, in terms of working at what it is that we're evaluating. We are trying to know whether we are evaluating model or the model and human component together along with the business logic. So essentially, for example, with our AB tests, editorial colleagues might have learned something about our models and start, they're used them, they are taught essentially in order to write certain things in certain ways. And so they adapt their behavior in a way that might not necessarily be fair across different variants that they might be using. And might not correlate with the behavior that we talk during log data and trying to contract these models and evaluate them offline to give them with. And it's really tricky. We've had this a few times in terms of building out capabilities such that we can put rail and dynamic the on pages in different places and the complaints into play with the construction of those rows, with how the model behaves, what models you choose, what variants you choose. And it's something I think we are definitely still learning, but what we found is to quick iteration process. It doesn't have to be a very formal thing where you write contracts with editorial staff. That gets you nowhere fast. But it's the case about saying having those quick meetings ever so often, so right? Given what you've shown us, we've done our offline evaluation. This is kind of how it looks. This is kind of the impression distribution. Does that kind of roughly match? Is that in the ballpark? Does that fit? And what you need are tools to be doing. And if not, we all go away and have that discussion not only internally, but with them about how they might be able to construct a group's content, say, that might work better with our models, but also vice versa, how we might work our models better to work with the content that they want to promote. And it's tricky because they're often, when we're developing a tool and a new capability, it's a learning experience, not just for us, but for our editorial counterparts. They're not sure exactly what the optimal output of these capabilities might be, because they've never used it before. They've never had it available to them. And so each case of really trying to shorten that feedback loop as much as possible, whilst retaining guard rails, or things like your accuracy metrics to make sure that you're not going completely off-base and accidentally generating a random recommender. And the same with some of these public service values, like catalog coverage. And we have metrics around that as well. And if something keeps coming up, fold it into your H, just that standard workflow of if there's a problem that keeps coming up that you could have identified, fold it into your evaluation process, and just use it as a guard rail for what you do. Alessandro, you also want to share something? Yeah. Just first of all, I wanted to make something clear. We have possibly we haven't mentioned it. When we talk about the tutorial, it's not the same team or the same person. I mentioned we have someone from the tutorial, editor's embedded in our team. But then when we develop the developer recommender for a specific product, we liaise which editorial from that product. So there are kind of these three polls. Well, there might be more. But let's take only those in isolation. There are just data scientists. We are the experts about the algorithm and how you can tweak. And pull some levers to make recommendations, go maybe more towards diversity, maybe just more for its relevancy. There is editorial from our side. They are experts of content. And then the editorial guidelines and public service purchases. But they are also just in their way, experts about recommender systems. Because they've been working with us for a while. They've been talking with us for a while. And they are embedded in the team. So they fought about the editorial side of recommendations. And then there are the editorial people, editors from other product we developed recommender for. They are the experts in that domain, that specific product. So they know their own inces. But they didn't create the content, but they curated the content. So they know what kind of content can be problematic to just have together. And I think-- and this is-- and in that triangle, editorial, our embedded editorial actually allows us to translate needs just back and forth from our side to their side. But I think it's possibly just you're refreshing with, then kind of just said. I think it-- from our side, what if we get from editorial? It's like, yeah, we can have a metric. We can look at different metrics. We can have a kind of metrics. But my belief is product and editorial product managers. I mean, product management and editorial, those that should tell us to do this, we are OK with this kind of degree of diversity of courage. Just we would have trade-offs. We would have some trade-offs between engagement metrics and just coverage. I don't think we, as they designed this, we are not the domain experts. We are not the content decoration experts. We shouldn't decide what the good balance between diverse and engagement should be. It should be those who are the experts in the product, those who are also setting product roadmaps and say, OK, they go, so the organization are those. And just the check-offs of the editorial, what are we comfortable with? Because again, there are some just public service values we should kind of abide by. And again, possibly, we don't do yet enough in terms of actively supporting those values. And we are working, if not our team directly, there is a responsibility team that works towards it, towards supporting understanding and measuring public service values in our recommendations. And in other machine learning engines across the organization. But I don't think the technical side-- in this triangle, it's not a triangle that are many functions and many people. It's not the technical side. Data scientists should decide what's good. It's about having it told other than having people say, this looks bad. Or you couldn't imagine just how many inappropriate pairing just could happen, even in terms of titles that cost together, you say, oh, no, no, that's not good. OK, OK. When you say this, there are a couple of things popping up in my head that I would like to better understand it, also used to-- so first, challenge them to maybe better understand. One of them is you say, so yes, you are the algorithm experts, and you can help encode the demands of your stakeholders and something that algorithms can optimize for. However, being the algorithm experts, we are also confident with statistics. And for me, it's always hard to understand how people can manually evaluate something that is the output of quite a complex mechanism. And what I mean by this is, especially is a difference in manual evaluation of, let's say, user-based versus item-based recommendations. That's my personal opinion. Using some kind or notion of content similarity, and using algorithm that performs that, so maybe based on topical similarity that is embedded and say, this is my source item, be it a news article or be it some video or whatever, and now show me something that is similar with regards to content. I would agree that this is easy to evaluate for humans because you kind of have that human notion of what is content similarity, and then there is an algorithm basically formalizes this and, as an output, for one or many seed items, I get those other item recommendations that somehow are exploiting that content similarity. However, when I look at the other side, so for example, at users, and I use user-based recommendations, so for example, with collaborative filtering, then I have a user embedding, and I want to use that user embedding to come up with item recommendations. And there I already think it could be already quite hard because this doesn't have to be similar with regards to the content of what the user has consumed because it embeds patterns of co, let's say, viewing or co-consumption behavior, which is, I would say, much harder to understand at first hand when looking at it manually. So. How do you deal, especially with those letter scenarios or how do your editors deal with that stuff, even though it might be dissimilar from what the user has already consumed, can still be relevant for the user? Yeah, so I think that, at least in the art play domain, it is a challenge, right? And this one we're grappling with. I think it's also useful to understand our recommendations in terms of the context in which they occur on the page. And so we will have a number of different spaces on the page. Only one of those is around full catalogue recommendation, where we don't have any kind of, I guess, item similarity constraints on what we're doing. Even if that similarity is at a level of new and trending content, for example, or most popular content on the BBC over the last week or so. And as a result, we don't tend to have to look purely at the user-based recommendations in absence of any kind of item meaning. A huge amount. It changes as well, dependent on the product we're talking about. I think it's also worth saying that a lot of these spaces are also quite heavily curated already. So they're there for a purpose that a tutorial want to achieve. And so there might be a relatively limited compared to the entire catalogue with the BBC's content, pull from which they are able to pull from and recommend. Okay. So basically your your your candidate space from which you finally recommend is already, let's say, constrained by editorial mechanics. Precisely precisely. And that also gives them a lever with which to use our algorithms to promote what they need and so on as well. That provides at some level a reasonable amount of safety because by the existence of the item as a candidate in that group, generally speaking at least from the other perspective, it's very unlikely that they'll have selected two items in that group that are really very inappropriate when put next to one of them. I don't say that happens. Again, this is very much connected to the language I just said. And the total actually are those that help us understanding the kind of experience we want to provide. With any railway that is powered by recommendations. And so they could, I mean, depending at the level, which was we use for recommendations episodes or programs or they could help define, you know, even simpler business rule, but saying, okay, this is the experience we want. We want to just use it, which is not just uniquely on the fact that we get metrics that score higher, but it's also the fact that just we know that this is not a good experience. Or this is the experience we don't want to be associated with this rail, with this place within the page. They know, for instance, that's each rail goes within a page. And so other rail that are possibly only just manually curated will provide that type of content. And so we say we want to avoid the type of content for appearing multiple places on a page. And of course, some of those things could be then addressed with automated approaches that we will need to refine. But others, you know, those are, we need that kind of domain knowledge to be able to provide that kind of experience. I think it's also because you want the algorithmic driven experience within the BBC to have the same kind of style to be informed by, you know, our tutorial guidelines. And so you want someone with that domain knowledge. I mean, I wouldn't let a bunch of data scientists for all of this. I wouldn't trust myself. This is a reason why we are all trying not to trust ourselves, but rather to trust the data in order to see whether the models that we have come up with are somewhat useful for the consumers. Of course, that feedback can sometimes be misleading and also lead to the amplification of biases. So there's the same content is narrowing down and consumed more and more. It might be a sign of relevance, but definitely not something that you folks want to achieve. Since you want to go in the opposite direction and rather want to see that users are exposed to, let's say, a diverse set of experiences and in order to be consistent with your mandate. Alessandro, you mentioned these values like being impartial, being distinctive, being high quality. I guess high quality for most of the listeners might be something they could directly relate with. I guess it's not only about accuracy, but I would assume it's more related to accuracy. However, when saying impartial, distinctive, high quality, what are kind of the algorithm objectives you translate that into or how do you translate it? And might this look different from the domain we are looking at? So what forms can these values take when we say we are in news or we are in sounds or we are, for example, in the eye player? I should say that we are very, at the moment, the wages, as I mentioned before, these active supporting, you know, our bodies. And it's very naive. It's done in very, using guard-real metrics. We do have engagement metric and this is possibly most of the interpretation we currently have about high quality recommendations. Here's the combination that they are engaging and rather one for just the user. Because after all, the possible way haven't said that at the beginning, the BUC's license fee funded, which means that audiences, they pay the license fee, it's not state-funded. It's license fee funded. People just pay the license fee and they get all the different services we mentioned earlier. And so, personalization is expected by user. Personalization allows them to get what they want in the small amount of time they have, because in terms of time, we compete with lots of different, not just other media provider, but social media. And people just have other things to do possibly that don't involve being online. And so, that's being high quality. Now, in terms of being distinctive, I could associate that to what we were saying, or just having this editorial saying, "Okay, this is where the ray and what we want to convey with that." In each page, within our online platforms, you know, editorial selected content is present, as well as algorithmically selected content. So, we have a combination of those types of content in every page. However, in terms of actual metrics, actual optimization for metrics, it's represented measures to some public service values. With CR at the very early stage, there's work being carried out within the response to play I team, because this is topic that's interesting, not just for recommendations, but also for other data science teams within the organization. We have a data science team, for instance, that took more content publishing and, you know, the study of supporting editors and journalists to produce content. So, I always see this as, you know, this kind of continuous line that goes from the creation of content and the publication to us that are, you know, the end stage of this audience-facing and distribution. And the responsibility I team is actively working, just on developing public service metrics that help us, first of all, understanding whether we are actually doing what we aim to do. That is supporting our public service metrics, but also then, you know, at the later stage, to optimize our content for those values. Duncan, maybe more looking into the eye player. Can you compliment this with some examples how you translate these values into corresponding metrics for the recommenders that you have in place for the eye player? Yeah, so, I think at the base level, I say we're at an early stage on this journey, even defining what these things are in numerically, can be quite tricky, but what we'd started by doing is a while back looking at the kind of standard metrics that early speak to the non-accuracy type, understanding of what we recommend to use them. Be it serendipity. I think there are about 20 different definitions of serendipity that are found in various papers in different places. I can't remember which one we picked. But we measure that and then same with even interesting type catalog coverage. Actually quite basic things. When you look at them holistically, you really do start to get a better understanding of what it is or achieving a scale. The other very valuable thing that we found is in terms of impressions, distributions, obviously. But even just eyeballing that can very quickly tell you, you can put a lot of warnings around it as well, but if you just have a massive skew in one direction, you know, we are quite obviously not serving up public service values. That alone in some of the actually stuff as well. And I think there we often have that quite direct conversation going in real time with editorial curation counterparts who have this sort of quite fundamental understanding of how the item itself, content itself relates to the public service and certain content will address certain parts of our public service remet and some will think of a documentary compared to a comedy show, we have these guiding Ruthie and Prince Pauls which are in form, etiquette, entertain, obviously a documentary fits into one part, comedy fits into another. And so understanding that breakdown and when at scale we look at the aggregated impressions of the way the recommendations are looking on maps, doesn't fit with the kind of expectations as our editorial curation colleagues might have for what a public service, you know, if we recommend documentaries to five people but comedies do however many millions obviously we're doing something wrong. So as I say it's early days but we're having those kind of continuous conversations and the goal will be and is to constantly update those and every time we find an issue or talk to our editorial colleagues and say and they say actually struggle to give examples off the top of my head but we're observing these kind of issues and the diversity of genres that you're expect on the top of the homepage so pulling something out of dinner. That is absolutely something you can operationalize, you can put metrics around that and then start to measure and you don't want to end up in a position where you have thousands of metrics at the point where you have effectively no metrics because you can't look at them and aggregate them and understand them but it is about using a number of touchstones to tell a story about what's going on and to really get a better list to understand. It's tricky right it's that trade off between trusting in the statistics but also this kind of softer understanding of what your algorithm is doing. Okay, gotcha. So far I've been yeah looking a lot on that editorial where the algorithmic or combined a agotorial topic and how we can translate those values into something that we could also then measure and hopefully optimize for properly to reach those values with support of algorithms. However, what are other challenges that you face in providing that personalized experience and delivering recommendations across the different products that you have so what are those challenges can you can you share a couple of them? So yeah I can maybe give an example of one that we often find ourselves grappling with. Indirectly and not necessarily realizing it quite so often. Well we've got a huge reach right the BBC particularly for iPlay in UK at least. From the 90% of the adult UK population that we reach obviously there are large skews in our data towards certain contents and touchstone things that might have happened during the week but that engagement is also driven by things that we don't really have access to personalized information on. You know we are we exist in a world where there is both video on demand and live streaming services but also a huge amount of BBC consumption happens. Not necessarily off-product but on actual TVs and the same with with radios and so what we find is that these linear schedules would say the news happens at 6 o'clock every night and people tuning for the news will really really skew a lot of the training days that we have in ways that is often first order so it's often this is on a prime time television by virtue of that it will get a huge number of people going and trying to find it but disentangling our understanding of that content independent of the promotion that it's been given not just on product but outside in in UK society can be quite challenging in terms of you know over-recommending things that are just very very popular. But it can also be found that those things kind of induce really powerful second order effects that might be a little bit less visible until something goes start behaving in an unexpected way elsewhere on product. So often we have things like onward journey recommendations of the back of one piece of content that you've watched will recommend you something else and what we can find is that if you know certain content gets really very high traffic we've observed it over the course of days we can see an oscillatory behavior in other content on product where the kit is drive other impressions of other content and clicks on other content which then gets promoted and becomes very popular so it's quite a complex system that popularity will drive yeah if that means sense. There's something we're struggling to grapple with generally. Yeah so and what is it that you that you do about these problems so one thing you mentioned is popularity and another one that you mentioned is that release cycle of certain recurring formats like for example is new so maybe there might be a tv show which releases once a week a new episode. I guess we haven't fully cracked it yet something where we are definitely working on this is actually sometimes where editorial curations will come in quite valuably right if we notice off the back of one piece of content is driving these second-door effects or even the first order. We can the editorial counterparts will view this as sort. Not being appropriate we can remove that from the set of content available to be recommended but the other thing that this does you can go down the route of popularity, de-buysing techniques and so on which we can and have explored but it becomes quite tricky when we don't have visibility of a lot of interactions with content a lot of the drivers that go on. And so it does lead us also to not just fully depend on collaborative filtering that you can't just throw it into matrix factorization and hope. We also need to lean on metadata similarities quite heavily to try and overcome this. So I think that's certainly one thing that we found does help is sort of reducing the size of the collaborative filtering effect on the final recommendation either directly or indirectly compared to metadata can help a lot with the with the so. Popularity bias is even at first order and then subsequently the second order and the more you personalise the more you can achieve these things accidentally and so it's a case of being really really careful and having adequate monitoring set up so the one things go when distributions start going how you are you're able to pick it up and address this as fast as it can before it starts affecting the training data of other upstream models and so. That's actually a good point where we could have some within experts advertising so for the listeners you also actually want to know more about popularity bias in recommender systems highly recommend going back to episode 19 where I talked to Heman Abdul-Appurri and talked about his work in popularity biases and also in techniques to debi so definitely give that a try. Alessandro in terms of the challenges what in addition do you have a mind or come across when I'm thinking about challenges today but also for the future that you want to address and haven't addressed yet. I think something that's it's very exciting you know it's one of the it's challenging but in the perhaps in the most part of some sense in the sense that we have hard looks to get to where we want to get and we know about those but it's got to be fun. We've talked so far about you know squads and you know recommender systems aligned to different products I play your sounds and as you can imagine you know this kind of suggest there are mainly BBC's and this is partially true because currently very often even data is siloed by product and and as well as our team is currently aligned to single product. But what we've gained a lot in terms of building relationship with editor or rather product managers from those single products but there's an intention within the organization to kind of connect up the BBC having a just all-around BBC experience rather than just simply offering it as a set of different products and that means that we will move we need to move towards you know thinking the BBC as a whole you know thinking about it just recommending content for across different products and in processing you're looking at user behavior not just within each product but across multiple products and there are kind of infrastructural and algorithmical challenges to that just we are in the process of building a new platform you know that that will be a common platform for all our recommender systems to their builds by our team I have one more team currently working on cross-product recommendations so mainly you know looking at interactions in one product and and providing recommendations for different products of looking at recommendations and BBC sounds and recommending relevant content on a player and this is the first this was this is the first time I could look at those cross-product interactions. I think it's interesting because in large organization you see those things that you have to match you know that us building the algorithm but also having the space for that there's a new space for you on the BBC homepage where users can find all the content in one place all recommendations are in one place and and that means at some point you know the collaboration with the total that we had you know that was kind of limited to oh just we don't want to see the same appropriate pairings within one product we will have to face these kind of things across different products so the program will scale up to the whole organization to just all formats and types of content. And I think just going down the truth, there is so much we can do in terms of thinking about the context where we haven't done context for the recommendations for instance. Just if you want to think in terms of like audio content, video content, news content, you want to give people the right content at the right time. And possibly just recommending a documentary, a very long documentary, just while people are commuting to work and have just lived at time to watch it, might have been the best thing. Or we might find also a way to say, OK, just add it to your content you want to listen to the future. So there is, I think, one of the things we are going to address next, as I said, we are just at the beginning. We are just working like part of the team as being working on this cross-product recommendations. But I think this is the start of something that at some point, the whole team will be working with this perspective of not just single products in mind, but the whole of the BBC and more integrated experience. Yeah, I like that. You are bringing that up because it was something that I had in my mind from the very beginning when I was preparing for this interview and checking on the different offerings that you have, those so many different domains. And then even within a single domain, so let's, for example, say, for sounds, you have, again, subdomains, like music, podcasts, and so on. So for me, that looks like you folks have already been working on cross-domain recommendations within the domains. But there's, I guess, even more potential to do that really across the whole BBC, so to have a user representations that tells me about something about their news taste. But then I might find something corresponding within the eye player, or maybe also a podcast. Those people are interested in that tells them more about politics because based on their news consumption, I found that they are interested more in politics as they are, let's say, and something different. And this is something that you are already addressing with its own team, but sorry, I didn't want to interrupt. I just wanted to highlight that team's role. But yeah, go ahead, Alexandra. What I'm about to say is just kind of a personal view, and not just any-- there are no plans within the BBC. But if we go back to what we were saying at the beginning, it is not just a news for type of players, sound. We have food. We have white sides for education. We have two, three, four. You could think about at some point just giving this integrated experience where you know that it's about dinner time. Just you could recommend some BBC food recipe and some music that is, you know, very well with that, or some podcast, you know, some drama that is set in Italy or in Turkey. And you recommend the recipe for all those places. And you have those all-around experiences because we have so much content of so many different types that you can think not just about an experience of a single piece of content, but a multiple experience that goes beyond. As I said, we have recipes. Why not thinking, oh, just why don't you just try this recipe? And at the same time, there's this nice drama and it brings you for an evening in Southern Europe. And you can have good time. It's not just about enjoying a piece of content, but having good, just valuable time where you do something and at the same time you listen and you act or just imagine for teenagers preparing their exams. Just we have bicep to prepare their exams. And then you can say, well, we have this documentary about history. Why don't you watch that? Because we know that you are interested in preparing this. You can add more depth. You can find it even funny. We have just even very memory if it's a podcast or you get this horrible histories, there's four kids and it's about history. And it's fun. So you can really have those kind of integrated experiences that's only an organization with a breadth that the BBC has can provide. We didn't know that there's far away in the future. I wish we could. We are not there, but hopefully we will get there at some. I guess definitely in some of the things that you provide or you are definitely doing a better job than we are right now in Germany. What's funny though is that the German public service media organization has been founded, taking the BBC as an example of the World War II, which I just found. But when I look nowadays, then I feel like you are already a bit more. The amateur was a unified is a wrong word and there were also good reasons in Germany to have that more of decentralization in terms of the democratization. But nowadays, when I kind of want to consume some political talk shows, I do always have to choose between two different apps because there is the app that ARD has and there is the app that ZDF has. And on your side, it already seems, oh, you got it covered with IPlayer and that's it. And even IPlayer offers different things from different, let's say, local providers, but it's all at least in one app. And in Germany, it's kind of at least it distributes across two apps, which I don't see any value. (laughs) I guess it seems like there is quite a lot of work that is still to be done and also great plans and interesting challenges for the future. Looking a bit more to the Rexus community and your involvement there, Alessandro, I also found that you have been a organizer for the queer workshop that has been taking place for the second time as of last year's Rexus. And I'm not sure, but maybe you can tell us if it's also going to take place for a third time this year. So what is that workshop about? And are you going to do it for a third time this year? - So I guess that was the last question. And unfortunately, we're not going to do it for the third time because the proposal wasn't accepted this time. - Unfortunately. So the idea behind the workshop was that kind of, you know, there might be different goals for explanations for recommender systems. But so far, kind of evaluating, you know, just like explanations, just along with different goals. There have been kind of spoke matters that have been different approaches. But the six approaches, like, first of all, just have been, there are lots of different approaches, not clear, which one can be used in different use cases. But also, and I think, and say this on the kind of industry side, it's been hard so far to find something that could be easily reused in an industry context. So I personally think that explanation will be passing or something we should explore at the BBC, just because, first of all, they don't just provide transparency, but it also increases the effectiveness of recommendation, the stressiveness, or, you know, your recommender system. But the amount of work, and you need to put into developing explanations, and the amount of work you need to just develop, just to develop, and approach to evaluate the, with all those aspects of explanations, is, doesn't much, you know, how much this can be prioritized, and the much does, how much, like, an organization like the BBC, which is, you know, large in terms of broadcaster, but not large in terms of big company. And what we wanted to do was, with the workshop, was just creating a place to bring together practitioners and academics to kind of discuss those things, discuss around goals, around how those had been evaluated in the past, and about how we can come up with approaches to the enable, organizations that the BBC, but even just our type of organization to evaluate those goals. - Nice, okay. Yeah, and anything else from the work that you folks are doing, or from the challenges at the BBC, that you might want to share with the community, or something that excites you also about the future, which we might have missed so far. - Yeah, something that possibly might be worth mentioning is that, and this, just by the same obvious too, people, but just over the year, we come to a stage where we build lots of recommender, and we'd be able to say, just have recommenders live, and most of them just may be the product. But at the moment, we are kind of, I wouldn't say, it's a step back, but kind of bring focusing on, okay, just we've built this so far, how do we get common approaches? Because very often, when you are just focused on delivering, or the determined product, you start just running away, and just you build your thing. And in any, I think, in an environment where you have, type that line, and you want to deliver on some goals, you then maybe just you need a realization tool, and you have a very skilled developer, and say, okay, just we can use this, this is very easy to make, and we have a visualization tool generated by a squad. Then another squad needs another visualization that has slightly different requirements than the other one. And you end up with another visualization tool for a tutorial. And then so what we're doing at the moment is just trying to kind of, for different experts and so far we have identified for, which is like just modern relation approaches, modern definition approaches, dataset construction, editor, build-as-reization and feedback tool. Just we want to say, okay, this is everything we've got, you know, we've developed so far. Let's have a look and let's see how just, you know, we can come up with just a single approach which shouldn't be, we should look because strain anyone, you know, we should easily adaptable but not to complicate it. It could be complex but not complicated, you know what I mean. And, you know, and having tools that, you know, we are a team just only counting the scientists of 14, 15 data scientists within recommendations. I think just you want to start, you know, like optimizing rationalizing, you know, our approaches to also just enable everyone to say, you know, what the people are working also be on a product and we're in a type of recommended person time but we want to enable also people to move to another topic and we want also to optimize efforts. So that's, you know, again, we have common tools. We have a common way also to define, you know, a model interface. I think this will be so important, you know, and because at the moment we don't have a common model interface but that will be enabled us. Okay, just we built this model for our player. Can we try it in sounds? Oh, that would be much faster. Possibly we will need to adapt it because it's different domain. And again, as I said, we want to connect up the whole of the BBC. So just we will look at this cross-products, interactions and behavior, but having first we need the tools to do this. And just we've recently focused on that. And I think it's been very valuable. Yeah, maybe if I could piggyback on that as well. There's some really interesting things about the fundamental understanding of what people do, how people behave with BBC content. So I think by understanding these things is rather than just being solo product things, if we can start to really tease out the question, not to get a fundamental theory of all recommendations in use of behavior ever. Lovely. Coming from a physicist, that would be nice but maybe not. But it will be really valuable for us to know what learnings and what understandings are consistent between our products. And that is an incredibly valuable thing to really get. But also what is, by disentangling that effect, you also, for free, effectively get the product specifics. I'm being able to clarify that and understand, right, we've solved this problem. It's only for this product. However, we've solved this problem and we can immediately deploy that out and repeat that learning in multiple places. Not only do we get better recommendations, which is what we're all here for, we develop that better understanding or the generalizability of our learnings, does that make sense? Yeah. By doing that, you learn faster, you move quicker and repeat that over years and years, hopefully you get further and further faster and faster. So it's learning those kind of fundamental similarities and differences, I think will be really, really valuable for us. Yeah, definitely I wish you all the best for that learning journey. Sounds like there are still a lot of things to be done. Okay, towards the end of every interview, I'm asking my guests also what they would like to hear more about, especially thinking about people, but it could also be topics. Is there anyone that you are having in mind or that you would like me to interview as part of rexperts or any topics that you might want to know more about that we haven't covered so far? I'm always thinking and you just remembered me of cross-domain recommendations than having already a couple of people in mind who I would like to talk about this, but yeah, if there is any person that you want to hear more about or any topic, what would it be? I don't know if it's both someone in particular, but I mean, will we be interesting to hear? I know that many of your guests that have heard in the past have this academic and this re-experience, but just as a topic, it will be interesting to hear, very often you see models and you see, there have never been papers published in the topic. You see, those evaluations in academic contexts, and then you try to use them in an industry context, and they are either useless because they don't get the performance you get, or they need huge machines to be trained or real world data. So, this kind of relationship in terms of, what is development in academia and having insights from someone who's going through the process and advice, just we could share anecdotes as well, but I remember this was when I moved from my PhD to work for the role of the scientist, you really see that, oh, what I learned so far is something to do with what I was doing in my PhD, but not quite because there's scale, and there are lots of problems, doing things in practice in a real world setting, very different. - Yes, I guess related to that is the group Milan who do a lot of work on the offline evaluation of models in academia, but also how we translate that to better understand how these academic models work in practice and production. I try and follow what they do, being very influential at least, I might think. - That sounds good. We'll put it onto my bucket list. If it's not there already, I have to check. - Great, thanks for sharing that, I guess, definitely to quite valuable points that we could address in some of the future episodes of our experts. - Yeah, and with that, we are at the end of the episode, and I thank you very, very much for taking part in this and sharing the experience that you have with your ex-us community. So thank you very, very much. - Thank you for having us. - Thank you. - And then I would say greetings go out to the UK and have a wonderful rest of the day. Bye. - Thanks you too, bye. - Thank you so much for listening to this episode of "Rex Birds", recommender systems expert, the podcast that brings you the experts in recommender systems. If you enjoy this podcast, please subscribe to it on your favorite podcast player, and please share it with anybody you think might benefit from it. If you have questions or recommendations for an interesting expert you wanna have a my show, or any other suggestions, drop me a message on Twitter or send me an email. Thank you again for listening and sharing, and make sure not to miss the next episode, because people who listen to this also listen to the next episode. Goodbye. (upbeat music)

Podcast Summary

Key Points:

  1. Personalization helps users efficiently find desired content, balancing high-quality outputs with manageable metrics.
  2. The BBC, as a public service media organization, faces unique challenges in algorithmic recommendations, requiring trade-offs between statistics and editorial judgment.
  3. Experts stress that scientists should not dictate the balance between diversity and engagement; domain experts and product teams should set priorities, especially as content appropriateness can change with current events.
  4. Collaboration between tech and non-tech teams is essential, involving shared language and understanding, as seen with manual curation at the BBC.
  5. The BBC offers diverse digital products—iPlayer, BBC Sounds, news, sports, education, children's services—where personalization is applied in iPlayer, Sounds, and sports, with news recommendations temporarily offline.
  6. The BBC's personalization team, founded in 2018, developed its first in-house recommender for short video clips, establishing a cross-functional approach with editorial and product roles.
  7. Guests Alessandro Piscopo and Duncan Walker bring varied backgrounds (arts to data science, particle physics to recommendations) and highlight serendipity and purpose in their work.

Summary:

The transcription discusses personalization and recommendation systems at the BBC, a public service media organization. Key themes include the need for personalization to help users efficiently access content, while maintaining impartial, high-quality outputs. The BBC faces unique challenges balancing algorithmic recommendations with editorial values, as content appropriateness can shift with current events.

Experts argue that scientists should not decide the trade-off between diversity and engagement; instead, product domain experts should guide these decisions. Collaboration between tech and non-tech teams is vital, requiring common language and understanding. The BBC offers numerous digital products—iPlayer, BBC Sounds, news, sports, education, and children's services—with recommendations active in iPlayer, Sounds, and sports, while news recommendations are paused.

The personalization team, founded in 2018, built its first in-house recommender for short video clips, emphasizing a cross-functional approach with editorial and product roles. Guests Alessandro Piscopo (lead data scientist, background in arts and classics) and Duncan Walker (principal data scientist, background in particle physics) share their career journeys, noting serendipity and purpose as drivers. They stress that personalization connects users with relevant content, especially given the BBC's vast daily output of over 1,000 items, while aligning with public service values.

FAQs

The BBC uses personalization to help users find content quickly, focusing on impartial, high-quality, and distinctive outputs. They balance statistical metrics with a softer understanding of algorithms to tell a story about user behavior.

Recommendations are used in BBC iPlayer for video, BBC Sounds for audio, and BBC Sport for sports content. The BBC also had recommendations on the UK news app, but it is currently offline.

The BBC involves editorial experts and product managers to decide the balance between diversity and engagement, rather than relying solely on data scientists. This ensures recommendations align with public service values.

The first in-house recommender was deployed around 2018, recommending short video clips to users based on their previous interactions, initially in the BBC Plus app.

Domain experts, such as product managers and editorial teams, set the product roadmap and decide the balance, not data scientists, as they are not the domain experts.

The BBC acknowledges that content appropriate today may become inappropriate tomorrow, especially with breaking news, and they work alongside non-tech teams to find common ground and adapt.

Chat with AI

Loading...

Pro features

Go deeper with this episode

Unlock creator-grade tools that turn any transcript into show notes and subtitle files.