Jean Carlo: Data science, Machine learning, AI, Mentorship
38m 48s
John Mashadu, Data Science Manager at Getty Guide, shares insights into the evolution of AI and machine learning in real-world applications. He highlights the rise of multimodality—where AI now processes images, text, and context—transforming how systems understand real-world complexity. Despite rapid progress, deploying AI at scale remains challenging due to unexpected data, system dynamics, and reliability concerns. John emphasizes the importance of MLOPS in ensuring machine learning models are reliable, observable, and maintain performance over time. His journey—from early computer games and hardware projects in Brazil to building data products in Berlin—reflects a deep appreciation for systems thinking, simplicity, and iterative problem-solving. He underscores the value of mentorship, open-source communities, and knowledge sharing in advancing technical capabilities. Key AI practices like fine-tuning and retrieval augmentation are discussed as tools to tailor models for specific use cases. John also notes the growing accessibility of AI tools, such as code assistants and agents, which enhance productivity but still face reliability gaps. Ultimately, he stresses that staying current requires curiosity, active engagement with communities, and a mindset of continuous learning and value creation—especially in high-impact, scalable environments like travel tech.
how to transition from this kind of thing that this model that we have and it works here into
these experiments, but I move that to a real-world scenario where there are data that you didn't expect,
there is a scale that you didn't expect, there is all sorts of real-world kind of changes,
and this remains a challenge like how to get all of this technology and ship it as a product
that is reliable and you can really trust all. How is AI changing the way we solve and approach
problems? What's multimodality and why is it so exciting? What is the impact of machine learning?
And what's the journey like in data science at one of the most important travel tech companies in the
world? Hi, I'm your host Pratik and this is behind the journey by Getty Guide, a show where we
explore and uncover what is really like to build the experience economy by diving deeper into the
journeys of people making it happen. Today I'm excited to have John Mashadu here. John is a
data science manager at Getty Guide and he shares a fascinating perspective on AI, machine learning,
mentorship, the importance of community and more. He draws from his journey dealing with the
intricacies of computer games back in the day to not tackling some of the most important problems
in tech. I learned a lot from this conversation and I hope you do too.
This podcast is by Getty Guide, a company on a mission to connect millions of travelers with some
of the most unforgettable experiences around the world. Imagine local experts giving guided
tours, skip the line tickets to your favorite attractions or exclusive bucket list experiences.
Getty Guide is headquartered in Berlin and since its launch in 2009,
travels a book more than 80 million activities through the platform and all of it is made possible
by a global team. So if you're looking for your next role and the thought of making an impact on
how people experience the world sounds exciting to you, head to Getty Guide dot careers. Let's get you
guide dot careers.
Hi John, welcome. Welcome to behind the journey. How are you doing?
Hello. Thanks for having me and doing great. It's Friday. I'm excited for the weekend.
That's good to hear. Coming to excitement. Actually, I have a long list of questions for you,
but the first one I wanted to ask is what is the most exciting thought that's living in your head
at the moment? Yeah, that's a tough question, but one thing thinking on the spot, it's easy to pull
out, is thinking around this area of multimodality in AI. Sounds very interesting for me. I'm
pretty much following the news and trying to stay engaged on it. Can you expand into a data team
general for everyone? Sure, it's this idea that with AI machine learning and so on,
there is usually data using forms of numbers or text and so on, but how to bring the other
senses to the picture. So a picture indeed is one example of bringing images back to machine
learning to solve all sorts of problems now with a little bit more of senses of the real world
beyond just bare data. Where do you see it going? There is applications everywhere. I see that
open AI for instance is doing huge progress on this and we are seeing all of this kind of
and with a lot of speed. A lot of speed. Yes, exactly. This is a very research topic at the moment,
but there are shipping products at the same time. So they are speeding up a lot of stuff in this area.
And you've been very close to all the AI action, especially we get you right. Can you reflect on
the last one year? How has the landscape changed? Yeah, there is a big kind of positive surprise,
I have to say, everybody is engaged about this, everybody is talking about it,
society, governments, everybody really, and speeding up indeed huge progress every week.
It's a new development and it's very hard actually to keep up, but at the same time,
it's a much more inclusive space at the moment, like so much more folks that can deliver value.
So pretty exciting. I have to get together. So we had been using AI for a very long time.
How has that changed in the company? I think we've been very much fast movers. There is a lot of
momentum within a company. TensorFlow folks engaged, TensorFlow presentations, knowledge sharing,
use cases, going to production, or more exploratory topics as well. So yeah, we are following the
kind of the big movement that's going on and learning a lot. Some of our practices improve,
but a lot is also much the same as before. And given that this space is so exploratory,
there's a lot that are open questions that need to be figured out and we can use this learnings
from the way we've been doing data products before and bring it back best practices and so on.
What are some of those open questions can you put on that? There is so much about trust and safety,
like what you're shipping to production can you really control and how to observe what's going
on, how to assert that it meets certain criteria or quality and so on. All these questions
tie back to best practices, how to go to production, which is a big topic dear to me as well.
Yeah, can you expand on just the challenges of in general going to production and yet it has a
huge scale and as an industry as well, what are some of the challenges that you've been
hitting from other people and different industries? That's a great question. My she learning has
as hard as much of academic pursuit and there was at some point before MLPs become a big thing
and still a problem that how to transition from this kind of thing that this model that we have
and it works here into these experiments and move that to a real world scenario where there are
data that you didn't expect. There is a scale that you didn't expect. There is all sorts of real
world kind of changes and this remains a challenge like how to get all of this technology and
ship it as a product that is reliable when you can really trust on. It's a big topic and I don't
think we will run a way of problems to solve in any point. Yeah, that's exciting. But what has
been your usage like of in general like personally from productivity standpoint? AI tools? AI tools.
Yes, I mean I mean touch to be so powerful, right? I think it was that big surprise on how much
even like just navigating German bureaucracy becomes so much easier now using these tools.
Going beyond like more to development tools, big user of co-pilot. I'm also a big believer
in building your own tools so use the SDK here and there and so specific problems using this kind
of technology. So yeah, I think I don't have any kind of super nitty-gritty specific to beyond
the ones that I built for recommending here just the classical ones stick with some sort of
complexion to charge your pizza. Yeah, I think the developed community is also shipping things
so fast and with respect to AI tools. I remember there was a few months back there was a talk about
this tool called cursor. It's a basically a folk of VS code and cursor is amazing. One of the things
that it does really well is like you can use your own open AI API key you can put in you can use
3.5 for GPT-4. But I think one of the best things that I like about it is that it has indexed
documentation. So you can chat with your code, but you can also reference documentation which
makes it much more reliable in terms of output and I think they are going to be more tools like that
in general. I think the great pros are still being built, but indeed there's all kind of obvious
problems that need to be solved and one needs to really make it very well polished to really solve
the problem. It's just one topic for instance on like you get an error on your system how to move
from that very descriptive error message to solving that problem. A tool that is language
you keep it with the large language models can definitely have a very good chance of solving such
type of problems, but the product is still not there. But it will come for sure. Any tips for
people who are excluding AI and maybe non-engineers in general? How can they keep up? How can they,
for example, look at the seismic shift that's happening and make use of it in their interest?
I like a lot of this area of agents also that solving very narrow problems deeply, but at the same time
trying a little bit more to your question. I think the ecosystem we evolve and new tools will be
available. We would stick a lot with OpenAI at the moment as they are clearly on the lead and try to
use strategy for your problems as much as you can. Yeah, following some new sources and so on to
stay on top, there is all this assistance on like other tools that I think will be exciting if
you're using some cloud Google Docs and so on. There is already some prototypes on Google. Yeah,
these things will get much better over time and so whatever you are doing in the tools that you use
are there tools that are already solving it with AI enabled to make it even easier or in the same
class of two or maybe your two already has a new just need to learn it. Can you expand your
bit if you mentioned about it? Can you expand on agents in the room? Yeah, so the prompt format is
open text, no context or you can add a little bit of context in the text format, but in the applications
we are using there is a lot of context. We are using an application for a specific purpose and it
has connections to other tools. Can you give me an example? Google Docs for instance has connections
to presentations and so on or to your drive with a lot of your data. Agents will play a big
growing neck thing of this dot and solve very well a problem, not so much into kind of stars from
first principles where you have to outline the whole text to a problem, but they will use a lot
of the context and solve problems for you. I'm pretty excited about that. Yeah.
That's very cool.
And on context specifically, there's also a lot of industry
talk about fine tuning versus drag retrieval augmentation.
Can you walk us through what does that mean?
And what's your view on it?
When to use which?
Yes, drag specifically is not a part that I focused a lot.
So it's my personal opinion.
But my view is that drag is for retrieving documents
and where you have an effect that you are trying to retrieve.
And so you can basically use drag to get you to the right place.
Why fine tuning is more generic, right?
It's like changing your model to become better at certain tasks.
You could use fine tuning for improving
the performance of a drag system by adding more context
and so on to make it better.
So I think they have a few different use cases
and they can be even used in conjunction.
Can you give an example?
Yes, let's say we are building a shut bot, which
has frequent ask-it-questions from a support system.
One can basically use ChatGDP and use drag over the documents
to basically have a vanilla version of drag on ChatGDP.
Or you can fine tune it.
Basically your model, let's say, ChatGDP
for having your tone or voice or something like this.
And then using these two things in conjunction,
you're going to retrieve documents.
And the answer to your questions
will be even more tuned to basically.
Similar to what Openities, to these custom GPDs,
but you can upload your knowledge base
and then have it add in a certain way.
Basically a much more tuned version of a chat, for example.
I think there are no fine tuning behind the scenes.
I'm not so sure because I'm not totally on top of this.
But indeed, you can improve performance
or factors that you also care about tone or voice
or structure if you need some sort of structure in your output.
And fine tuning?
What about--
For example, fine tuning.
There is many reasons one can use it.
I think from the AI perspective, the most common
is the tone of voice is why I think
the core one that OpenAI also advertises a lot.
But you can remove the head of the model for any kind of tasks
like a text processing task.
Like you can start doing classification
with a large language model.
So you would try and tune that large language model
in a classification task and it would suddenly start
not being a generic model anymore,
but like a classification system.
There is many reasons to do it.
So if I understand correctly, fine tuning
is more about teaching the model to do something truly novel
to something new that it does not know in general, right?
Yeah, exactly.
One can teach you something novel or influence it
to do something in a slightly more refined way,
specific one indeed changes the behavior of the model
in a more fundamental way.
What's been your experience?
And I know you fine tune a lot.
So how has your experience been just fine tuning models
for your use cases?
It's early days.
And if you evokes a lot, there is different kind of ways
one can do it, one can use OpenAI,
and then you have this API experience,
but you lose all the control.
Although it's probably a great way
to start for many problems that you want to solve fast.
If you want to go more deep, fine tune open source model,
like Liyama, then yeah, that it's also getting better
and better easier and easier.
There is a couple of tools that are becoming very standard
like Hagen phase.
Of course, you need a GPU.
You could also do without it,
but it's so much harder that I don't count that as an option.
And then there is this standard process as well.
It's very much like machine learning,
more classical machine learning,
but it involves a little bit more hardware
and some steps of preparation.
- Got it.
- So when we were talking a while back,
just before the podcast, you mentioned games
and complex systems and how this sparked
your interest in computers.
How have these early experiences reflected
in your current approach to problem solving in general?
- That's a very deep question, let me try to unpack it.
I think the problem solving component on it,
games, they are very kind of logical things,
and you can see with simple rules,
you can have very complex behavior.
And back in the day, also, when I played games
in my first experience, I had a very bad computer
and I had to hack it around to make things work.
- How was it like, where were you?
What stage of a live video?
- It was maybe 13 years old and got this old computer
and wanted to play some games, huh?
- Did you remember which one?
- I mean, back in the time,
I guess there was the DW2, Age of Empires,
these were the things that kept me hooked.
Yep, and basically, one had to hack out a lot of stuff
to get it to work in my Pintune 2.
Yeah, it was a great learning experience,
even though I didn't know it conscientiously,
like I had to go into the OS and figure out
a few other process that could be cured
and make things to work.
Yeah, I love that time
and there was a lot of energy for me
into this kind of real time strategy games,
simple rules, complex systems, which kind of kept me hooked.
Somehow, I managed to then invest this energy
that I found there more into kind of programming
once I got those skills.
And yeah, that gives me also a lot of confidence
when programming after doing so much
on the side exploring my own ideas,
which I think was very helpful
and it's the same energy that comes back then,
like looking at this system's working.
- Where were you at that point in time?
- We on Brazil. - In Brazil.
And you mentioned programming in general,
I'm gonna ask you a controversial question.
What's the most beautiful programming language?
- Yeah, it depends on which criteria, I would say.
I have done my chair of exploring programming languages.
I had a brilliant friend at some point
that was really genius level into programming languages
and I got super inspired for him
and I did my own research
and explored really tons of the programming languages
and my conclusion after doing that for a long time
was simple is the most beautiful,
like solving problems faster and reliably,
that's super power.
To come back a big fan of Python,
I also helped the community in ways
that I can organize by data in conference
that I'm happy that data science and Python also found
like a perfect match there.
And that's from overall programming language question,
but if you would ask from purity perspective,
like of mathematical purity,
for sure at the end we would learn more kind of Haskell
or Elm or languages like this.
- As a non dev, I think I also find Python
to be very inclusive and accessible as a language.
And the community is strong.
- There's a lot of purity on that also that,
it's an open community with open values
and yeah, that it's a big plot.
- Yeah, I agree with you, totally.
And you transitioned from building really complex things,
also from a research,
you were doing research back in Brazil.
And not exactly was more an academic context,
building hardware or prototyping hardware
for like accessibility use cases.
- Okay, can you expand a bit more on that?
- The biggest project that I was involved
was around like a brightly reader,
some pins, hardware pins that goes up
as you are reading something in the computer.
So the blind person can basically reach the computer
what's going on.
- Oh nice.
- Yeah.
- The project was not really successful
by the time I left, I spent so long.
- Yeah, but I designed the kind of electronic prototype
and that worked well and they were happy with the results.
But the toughest part to escape at that point
was the hardware and I had no clue about that.
It was developed on the side by other folks.
- And then you transitioned into a depth job.
- Yes.
- How did that transition go?
- I think the context is rather different, right?
You are into this kind of more research kind of exploration,
which is very low level, more than going to electronic side.
I studied electronics for some time
and then moving to programming as main task.
And then it's about building things reliably fast.
I think knowing the low level really gives you
an extra hint on how to build things in the right way.
We have, that's for certainly helped.
I really liked to go higher level over time.
I first I was really passionate about,
oh the bike's here going through and there.
But over time we came all about
how can you solve this problem in the right way.
- And how long have you been doing this now?
- Yeah, more than 10 years.
This was 2012.
- Yeah, and when did you move to Berlin?
- In 2018.
- Okay.
Did you move or get you guys by the way?
- Yes.
- Okay.
- Yeah, six years now.
- That's amazing.
Half a decade.
More than half a decade.
That's great.
And we all learned from other people mentors
and whatnot have you had those guiding figures in your life
who you have learned from?
- Yes.
One very instrumental person in my career
was El Tominato, his big relatively to the kind of space
that we are like in the software kind of spaces,
very well-known in Brazil, particularly there in the,
he was back in the day in the PHP community
and he was a superstar there.
- Okay.
- And nowadays he transitioned more into Go
and more into tech leadership conversations.
- Yeah, he was really instrumental very early in my career.
He gave that push on gorgeous conferences everywhere
and presenting and connecting with the leaders
on that field and getting to know the kind of people
that were moving the needle for that community
and yeah, was really very good for me.
- In general, with events and conferences,
what we are biggest learnings has a mentee.
- He teached me a lot about learning through teaching.
So, if you want to get good at something, it's a very effective way.
have to teach us to do somebody. Also, that's most of the people around you. I struggle
with the same stuff, but they usually don't talk about it so much. So if you give a presentation,
you're going to find tons of folks that can relate with the problems you're facing and
then you can discuss things and move forward. This very open community mindset of open source
also was very much tied to this culture. And do you have current mentors? Do you actually
mentor at the moment to anyone on my current role as a manager? I see there is a relationship
here that I'm very much committed to their growth. How can I move the needle for them further?
Yeah. When I was more in an individual contributor role, then I had more kind of a formal mentorship
conversations. Pretty busy with parenthood right now. So I will go back to it, but it's good to have
a tidy break at the moment. Yeah, for sure. I really understand that there are key lessons that you
learn from people when they mentor you and both as a mentee as well. It was also part where you
were leading the technology part of a startup under his leadership and mentorship. What if you go
back to that time, what are some of the things that come to your mind that really paved paths for you
for the rest of your career? He teach me a lot about how cloud solutions pace, like how to build
systems for the cloud with this care that you can solve problems for the entire world. Yet being
very simple, the problems, ideal, they are complex and the solution is simple. Trying to really keep
the simplicity. There is a lot of beauty there and I learned a lot about it as well from him.
Also about how to deliver value iterative and going deep into systems and understanding
all the components that that moves our cloud environment. Yeah. Tons of learning. Cool.
All right, you mentioned about moving to Berlin forget a guide and you also mentioned about PHP.
There's a connection there. Take us back to the time when you joined get a guide as a PHP engineer.
Yes, so other his mentorship, I could very much speed up my career and get very deep into the core
communities in PHP and know everybody. So it was relatively easy then to learn the job in Europe.
If people can find you everywhere online, like in the right communities, then as a developer,
it's a great way to get a great job. I was pretty happy on the startup there. There was a lot of
progress. It was very hard to fail because there was a big company behind it, but I wanted a
little bit of adventure, quality of life of Europe, and the kudo may have trouble. All these things
play the role to make it into the season. You constantly stress on the topic of
the depth being a bit more out there, going to conferences, doing meetups, writing blogs,
things like that. Knowledge sharing is a big component. I understand that you are stressing on
both to gain new perspectives when you share and also to share your perspectives on things.
For people who were just starting out in these cases or are not that out there specifically,
what are some of the things that help you be a little bit more out there, support communities,
also do so much at your job and give back. Yeah, great question. I think you have to figure out why
you were doing it, helping to improve the status quo. So many systems out there being viewed.
It was very much into Linux at the beginning of my career. I wanted to become a kernel developer.
But yeah, what impact, because some folks have brought the core kernel kind of parts, their code is
the impact of that is like huge. If you make a contribution to one of these core projects,
billions of people all day are using your code. So I think with computer science, if you are
really being driven by this kind of impact that you can have in this field, then I would certainly
recommend people to meet other people, to start establishing a connection, build some hobbies,
kind of stuff that interests you naturally and just try to play with that idea a little bit.
Make sure that this work is somewhere visible that other people also can reach it out.
Like GitHub, of course, is one building this portfolio of things that you are interested about.
It's over time that that basically becomes a lot of evidence that you are very much interested
in stuff and people will reach out and want to stop together. The stop about machine learning.
And let's do a machine learning one for everyone. Can you explain what it is?
Why it matters? And why should teams care? Sure, what it is. It's yet another way to solve
problems that otherwise could not be solved. The programming you can do a lot of kind of specific
things, but like following recipes, but when machine learning, if you put some data together,
you can start looking into predictions. You can start looking into how things would look
like if a certain scenario is true or false. And that's extremely powerful in all sorts of
scenarios. It has such a big potential to use data in healthcare or education, like building
predictive systems that can tell you what is your tendency for certain diseases or how to prevent
things even ahead of time. Education as well, how to tailor education to your specific knowledge
gaps or your culture and make the most of your specific needs. There's a lot of statistics
there looking to historical data on how these things look for other people and you can really make
a dent into these problems. It's chaos. So everybody benefits basic. So that's more kind of
society part, but of course, for business, it's extremely useful as well. We can all sorts of
business problems at scale with these tools. It's also can be seen as software tools. Oh, right,
that Andrew Kaparje talks about in the famous web post that the space of things you can solve
with software in a Venn diagram is X, but there is a much bigger space that you can solve with
data and the machine learning and AI too. For those who do not know about him, can you?
I think he was ahead of AI at that point. Yeah. And he was involved with NOPE AI at much
her home. He has a very good tutorial on how to build judgment. Yeah, I think he published that
in, I think it's one or two hours video. How to build your custom GPT. How to build an LLM
or something like that. I find that insane. Like he put all that knowledge out there for people
to see in a YouTube video. And it's also pretty popular. I see that as millions of views.
I did that tutorial. Yeah, it's brilliant. I think it speaks a lot to the moment of this
openness of knowledge and like that makes things much more inclusive. There's other perspectives.
People think also, okay, there's this big company that own all this hardware and they control
everything that's certainly something to watch out for. But actually, I feel that there was
so much more inclusivity going on after the introduction of this. How have the last 12 months
changed for machine learning in the room? Having been in different ways as I mentioned in the
beginning, there's an overlap of forces. Standard machine learning can bring a lot of good practices
to the LLM space because there we already figure out so much. But the other around as well,
right? Like the power of the language can unlock so many new opportunities and possibilities.
That these models are so good with language. You can solve some problems much
faster, although reliably is still open question, but we are okay on this and making progress.
So there is a lot on this aspect of just unlocking possibilities that before would be too costly
to build or you would take so much more time to build it or much more people. And now it's
reduces a lot of the complexity. It allows people to focus more for solving the problem rather than
all this kind of technical sub-problems that arises when you have to build this stuff from scratch.
You've been pretty active in the MLOPS community as well. How did you get involved?
Yeah, that ties very much into my transition from a pitch. My engineer took more into the machine learning
field. Three, four years ago, people realized that actually there is a new field here. There was a lot
of talks before about Google has been doing this for much longer than four years and the other
tech giant. But I think that there was like the realization that this was important for business
because a lot of people wanted to put machine learning into production and didn't know how
and there was a long time where there was a model of throwover defense problem that you assume
that there is a data scientist with the particular skill set. They solve the problem that they send
to the engineering team and they solve the problem. But actually there is a better way to do it,
which is you keep the ownership within a data science team of going into production and you
try to understand what are the needs there and build the tools and systems around that fulfill
those needs. And there's quite some particular needs to go to production with machine learning.
Yeah, so get your guide was in that point as well. We were struggling to put ML to production
and MLOPS arises this topic that was the materials kind of leading this very famous
community podcast back then was mostly a podcast guy super funny and I just I messaged him at some
point Hey, do you want to come to talk to us and get your guide about the stuff that you are doing?
I think there is overlap. We met the he come in our online conference and we had a great time
and eventually we decided to create a MLOPS position get your guide. I transitioned to there
and I've been connected to the MLOPS community so far. We created a Berlin chapter of the
meetup. Yeah, tons of learnings. We run a lot of meetups. Well, the first one get your guide.
The second one was in vote. The third one was in Google. But the fourth is going to happen soon.
So looking forward to that. Can you expand on the role of an MLOPS engineer? What does the
Detroit look like? What are the challenges? I think it has to go back to what is MLOPS, right?
There's everything around trying to go to production with machine learning. So it's not about modeling.
That's about what are the tools and the ecosystems that one need around it,
like which systems needs to be in place to do it reliably.
So the ML ops engineer will basically support data products team,
so we're building machine learning production to serve the customer in data products.
So the data scientist has a modeling problem how to fulfill that use case with the systems,
which systems need to be in place to get the right data at the same time to make sure that
the data remains you need to observe it. So to make sure that the data remains constant,
you need to basically figure out how you do machine learning. There's a lot of
model features, which you get the data you're transforming it in a right way.
And that has to be retrieved in the right way also in production. So one needs to figure out
how to do that for production use case where you have maybe thousands of requests per second,
or you have a huge scale on the data load like terabytes of data to process.
So how to retrieve data correctly, how to make sure that it's observable,
how to make sure that your system that uses machine learning is optimized to be fast enough
because that's a compute bound problem usually. So you have to do some computes and return the
results. And there's tons in the particular models that needs to be optimized. Yeah,
looking to how to run it reliably in production, keeping costs at bay, keeping
scale ready, which tools needs to be in place. Can you take a real life example for get your guide
and explain what does a day-to-day look like? One of my cover problems is the real-time
rewrunker. For each activity, you're seeing get your guide's website. There is a system that
after these activities that can fit somehow that location were defined. There is a system to
define which order those activities should appear. So what comes first? And what comes first
is usually it has to be what you want. And so there's a system that takes input about your
context where you are, the kinds of things that we've been interested in get your guide so far,
and basically figure out which activities to show you next. But that's a very high scale right,
like thousands of requests per second and in the high peak. And there is my machine learning
model basically deciding that question. What fascinated you about data products? You mentioned
briefly about it and get your guide. And also what can you tell us a bit more about the team,
you know? Yes, what is data products? We are basically building machine learning solutions
for solving customer problems and yeah, as short as that. So looking into basically the customer
journey and all its experience and figure out how data can make that better and optimize it to
make it more relevant or get you the right content in place at the right time. That's basically
what this team is about. I think what fascinated me about this area in general is
indeed the natural kind of your complexity of the problem and the real impact you can have
with it at scale. So you can really change this experience of so much people with, yeah,
looking at the right problems. So there's a lot of emphasis in problem definition figuring out
how to crack that in a sort of a mathematical way or so. So I like the inherent complexity of
solving problems like this and having them at scale. So I think these are kind of ways that I got
attracted to it. In complex systems like you work on, what's a decision making framework like?
Well, how do you make decisions? How do you choose not just what to work on, but how do you decide
what's what it in general? Yeah, we tie a lot back to get your guys values and these emphasis
in impact. Yeah, that is a big lever. Also strategic considerations is important for the business
not only now, but in five years or longer. And how we decide on if we are successful or not as
part of the rest of the tech organization, we are very much experiment oriented and data oriented.
So there's a lot of maybe testing going on. That's the way we operate and in data products,
we are saying now more often than not, we have superpowers of like knowledge sharing, a lot of trust,
a lot of knowledge sharing. Yeah, that's really helps a lot and makes it very engaging to keep
trying to solve problems there every more kind of impactful and so on. By knowledge sharing,
you mean you have sessions weekly and things like that? Yes, we have a great collaborative culture,
the huge overlapping values and the same kind of passion for core kind of topics and ceremonies
and we try to look into the way we work together continuously to make it as streamlined as possible
and a lot of knowledge sharing presentation, meetups and so on. What are some of the skills that
and quantities you look for in your team members? The get your guide framework helps a lot here,
having this cultural values, it's important having a big overlap in them, having an impact
oriented mindset. So how can you make the biggest change for the customer? Ownership is also important,
what you build you are responsible for and if it needs to change, you are also responsible for that.
Can you expand on that a bit more? Yes, this goes back also to the topic I mentioned with
Emma Lobs that things has been done in the longest time to go to production with machine learning
through throwing over the fence the problem and suddenly becomes, you build something the other
person has to replicate that in another system to basically make it work and we manage to change that
to you build it you want. So as a data scientist, you have a lot of ownership on the model that you
build there, we put tools in place that you can also observe how that changes over time and a lot
of automation. So for instance, there is this concept of drift, let's say the road changed it,
there is now, I don't know, one market becomes more popular than the other or there's a
zonality. All these things influence data product, right? So one needs to be on top of them,
performance can be great over time. So ownership here means we build something and we keep
iterating on it and make sure that it keeps working the right way.
How much of that is also about the landscape in your field changes a lot.
What advice would it give to people to be on top of things in general? I think being really driven
by your corporations looking for what is that and how that can overlap with what is going on around
you and that naturally creates energy to basically go out and meet people or learn a new topic.
Also having this mindset of sharing a knowledge, right? So if you're constantly looking how can I add
value here, which topics will be important in some time? Yeah, this continues interest about
the field that leads to a lot of drive on itself and a lot of energy to continue doing it.
And things get easier to do over time. So it's rewarding also to see, we're making progress, right?
It's not only about, yeah, collecting more information, solving problems better, you're solving
more problems or at a different scale. What do you do when you hit a roadblock? Continue talking,
I would say there is always so many other options and continue brainstorming. There is always
alternatives. We sometimes feel we are stuck, but it's just give us a back and really the stuck
situation is often the case, not really there. Yeah, cool. Last question, what has been your
recent get-to-guided activity? What did you do? Ah, I had the last year a very special,
no, it was this year. A special vacation when Granny came from Brazil. First time in a commercial,
I played, actually. So changing confidence. Yeah, she stayed three months here and we went
within this travel also to Italy. As she wanted to see the Vatican and the Pope and so on. So yes,
I took some, yeah, quite some stuff there with get your guy, the Vatican museum was quite memorable.
The call of us back to that moment, how was it? It's breathtaking to see so much history.
Yeah, like quite memorable trip, a lot of beautiful moments that also resonates a lot with the
get-to-guide way that so much about memories and this was really one of the best highlights
of quite some recent years. What did you get anything? As she loves it and she's coming again next year.
Now she wants to go to Paris. It's good, but it's a beautiful city.
Yes, it's a good thing. Cool, cool. Thanks so much for coming in today. It was a pleasure talking to
you, but so many friends. Thank you very much. Thanks so much and have a good
stuff to the week for you. Like quite a great fun here. Thanks a lot for having me. Thank you, John.
Podcast Summary
Key Points:
Multimodality in AI is transforming how machines process real-world data by integrating senses beyond text and numbers, such as images and context.
The AI and machine learning landscape has seen rapid growth in the past year, with increased accessibility, production use cases, and broader industry involvement across sectors.
Transitioning AI models from controlled experiments to real-world applications remains a key challenge due to unexpected data, scale, and dynamic environments.
Fine-tuning and retrieval augmentation are critical techniques for customizing AI models—fine-tuning adapts models to specific tasks or tones, while retrieval augments answers with relevant documents.
MLOPS (Machine Learning Operations) has emerged as essential for reliable, scalable, and observable AI deployment, with roles focused on system design, monitoring, and automation.
John’s early experiences with games and hardware prototyping shaped his approach to complex systems, emphasizing simplicity, logic, and hands-on problem-solving.
Open-source collaboration, knowledge sharing, and mentorship are vital for personal and professional growth in AI and data science.
AI tools like code assistants (e.g., Cursor) and agents are accelerating productivity, but challenges remain in reliability, error resolution, and trust in outputs.
Summary:
John Mashadu, Data Science Manager at Getty Guide, shares insights into the evolution of AI and machine learning in real-world applications. He highlights the rise of multimodality—where AI now processes images, text, and context—transforming how systems understand real-world complexity. Despite rapid progress, deploying AI at scale remains challenging due to unexpected data, system dynamics, and reliability concerns.
John emphasizes the importance of MLOPS in ensuring machine learning models are reliable, observable, and maintain performance over time. His journey—from early computer games and hardware projects in Brazil to building data products in Berlin—reflects a deep appreciation for systems thinking, simplicity, and iterative problem-solving. He underscores the value of mentorship, open-source communities, and knowledge sharing in advancing technical capabilities.
Key AI practices like fine-tuning and retrieval augmentation are discussed as tools to tailor models for specific use cases. John also notes the growing accessibility of AI tools, such as code assistants and agents, which enhance productivity but still face reliability gaps. Ultimately, he stresses that staying current requires curiosity, active engagement with communities, and a mindset of continuous learning and value creation—especially in high-impact, scalable environments like travel tech.
FAQs
Multimodality in AI refers to the ability of models to process and understand multiple types of data, such as images, text, and audio. It's exciting because it allows AI systems to better mimic human perception and interact with the real world in more comprehensive ways.
AI is enabling new approaches by allowing systems to learn from data and make predictions, shifting from rigid rule-based logic to adaptive, data-driven problem-solving that can handle complexity and uncertainty more effectively.
Key challenges include handling unexpected data scale, real-world variability, ensuring model reliability, maintaining performance over time, and establishing trust through observability and control mechanisms.
Fine-tuning adapts a model's behavior to perform specific tasks, such as changing tone or improving classification accuracy. RAG retrieves relevant documents or context before answering, enhancing accuracy by grounding responses in external knowledge.
MLOPS (Machine Learning Operations) ensures models are deployed reliably, scalable, and maintainable by managing data pipelines, model monitoring, observability, and automation to handle real-world data and performance changes.
Non-engineers can follow open-source tools, AI tutorials (like Andrew Kaparje’s), attend webinars, and use AI-powered assistants in tools like Google Docs to stay informed and apply AI in their work efficiently.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.