Episode 27: Machine Learning & AI Services on AWS | SAA-C03
20m 55s
This episode covers AWS AI and machine learning services, emphasizing that the exam focuses on pattern matching rather than technical implementation. The majority of questions ask whether a pre-built AI service matches a given use case—such as speech-to-text, translation, or sentiment analysis—without requiring knowledge of machine learning models or coding. Services like Amazon Transcribe (speech-to-text), Polly (text-to-speech), Translate (language translation), and Comprehend (natural language understanding) are tested through clear signals. For conversational AI, Lex is used to build chatbots and voice assistants. Textract extracts data from scanned documents, while Kendra enables intelligent, natural language searching across documents. Amazon Personalize delivers real-time product recommendations, and Forecast predicts future demand from historical trends. The key decision rule is: use a pre-built service when the task is standard and well-defined; only use SageMaker when custom model development is required. This topic is one of the most accessible on the exam, making it a high-confidence area for candidates who learn the core signals and mental models. With this foundation, learners can confidently tackle AI-related questions and complete the full SAAC03 curriculum, paving the way for exam-specific strategy in future episodes.
Hey everyone, Balu here. Welcome back to TechTalk with Balu, your complete guide to Async the AWS Solutions Architect Associate Exam.
Today we are tackling episode 27 and this is a milestone because it's the final code topic in our SAAC03 journey.
We are covering the machine learning and AI services on AWS and I've got some genuinely good news about this one.
Of all the topics on the exam, this is one of the least demanding because the exam tests these services shallily.
You do not need to know how to train a machine learning model itself, you don't need to understand the math and you don't need to write any code.
For the vast majority of the exam questions, there's really only one skill being tested here and it's this.
Given a description of what a company wants to do, can you name the right AI services for the job? That's it.
The question will describe a use case like a company wants to automatically add subtitles to its videos and you just need to recognize that speech to text is the Amazon Transcribe.
So this episode is fundamentally a pattern matching exercise. I'm going to introduce each service, tell you in one line what it does and give you the keywords that signal it.
Learn those signals and you will pick up these questions quickly and confidently on the exam day.
Here's the mental model to hold on to.
AWS AI services come into groups. The first group is the pre-built ready to use AI services. These are like specialist appliances. Each one does a specific job out of the box and you just call an API.
These are services like recognition, transcribe, poly, translate, legs, comprehend, text track, Kendra, personalized and forecast.
The second group is the build your own tool, which is the Amazon SageMaker. That's for when the pre-built services don't fit your specific need and you want to build, train and deploy your own custom model.
We will first cover the pre-built specialist services since they're most commonly tested, then we'll finish with the SageMaker.
By the way, we're keeping the same intractive format with pulse checks, trap spotlights and memory hooks and we're keeping it punchy. So let's get started.
Let's start with the AI services that work with images and video and the main one is Amazon recognition.
Recognition finds objects, people, text and scenes in images and videos using machine learning.
If a use case involves analyzing a picture or a video, recognition is almost certainly your answer.
Its capabilities include labeling, which is identifying objects and scenes, face detection and analysis, which can estimate things like gender, age, range of emotions, face search and verification, that all-power user verification and people counting and even celebrity recognition and parting, which is tracking the part of people, for example, in sports game analysis.
One specially examined relevant use case is content moderation. Recognition can detect content that's in appropriate, unwanted or offensive in images and videos.
This is used across social media, broadcast, advertising and e-commerce to keep these platforms safe.
You set a minimum confidence threshold for what gets flagged and you can send flagged content for manual human review using a service called Amazon Augmented AI or in short A2I,
which is worth recognizing as the human review of AI prediction service.
So the signals for recognition are anything about images or videos, detecting objects, faces or texts within them, facial recognition or verification and content moderation.
For your memory hook, think of recognition as the eyes of the AI family. It's the specialist that sees and understands what's in your pictures and videos.
Now let's move on from seeing to hearing and speaking, which brings us to two services that are mirror images of each other, Amazon Transcribe and Amazon Poly.
Amazon Transcribe automatically converts speech to text. It uses a deep learning process called Automatic Speech Recognition to turn audio into accurate text very quickly.
It can automatically remove PII data or personally identifiable information using reduction and it supports automatic language identification for multilingual audio.
The class use cases are transcribing customer service calls, automatically generating closed captions and subtitles and creating searchable text archives from media.
So whenever you hear speech to text, audio to text, subtitles, captions or transcribing calls, that's the transcribed service.
Amazon Poly is the exact opposite. Poly turns text to life like speech using deep learning, letting you build applications that literally talk.
It has some nice customization features. With pronunciation lexicon, you can control exactly how specific words are pronounced.
For example, making sure an acronym like AWS is read as Amazon Web Services or a stylized name is pronounced correctly.
And with speech synthesis markup languages or SSML for short, you get even more control, like emphasizing words, adding pauses or breathing sounds, whispering or using a newscaster speaking style.
So whenever you hear text to speech, turning text to audio or making an application talk, then that's the Poly service.
The way to keep these two straight is to remember the direction.
Transcribed goes from audio to text while poly goes from text to audio. There are matched pair pointing in the opposite directions.
For your memory hook, transcribed is the ears of the family, listening and writing down what it hears.
Whilst poly is the voice of the family reading text allowed in a life like way, ears take sound in, voice puts sound out.
Next, let's cover language and text understanding starting with Amazon Translate.
Amazon Translate is exactly what it sounds like. It provides a natural and accurate language translation.
It lets you localize content like websites and applications for international users and translate large volumes of text efficiently.
There's really not much to overthink here. If the use case is translating bitven languages, it's Amazon Translate.
For your memory hook, Translate is the family's multilingual interpreter, fluently converting between languages.
Now let's go deeper into understanding text with Amazon Comprehend. Comprehend is for natural language processing, often abbreviated as NLP and it's a fully managed serverless service.
Where Translate just converts languages, Comprehend actually understands and extracts meanings from the text using machine learning.
It can identify the language of the text, extract key phrases, places, people, brands and events, understand the sentiment meaning how positive or negative the text is, and automatically organize a collections of documents by topic for you.
A classic use case is analyzing customer emails to find what drives positive or negative experiences or grouping articles by topics Comprehend discovers.
So the signals for Comprehend are natural language processing, sentiment analysis and extracting insights, entities or topics from text.
There's also a specialized version called Amazon Comprehend Medical, which detects useful information in unstructured clinical text like Physicians Notes and Discharge Summaries and specifically detects protected health information.
If a question mentions extracting medical information from clinical text, then that's Comprehend Medical.
For your memory hook, if Translate simply converts languages, Comprehend is the family's analyst who reads the text and tells you what it actually means, how it feels and what it's all about.
Now, let's talk about building conversational interfaces, which is Amazon Lex.
Lex is built on the same technology that powers Alexa. It combines two capabilities.
Automatic speech recognition to convert speech to text and natural language understanding to recognize the intent behind what someone says or types.
Together, these two build chatbots and calls and the bots.
So when a use case is about building a chatbot, a conversational interface or a voice assistant that understands intent, that's Lex.
Lex is often paired with Amazon Connect, which is a cloud-based virtual contact center.
Connect lets you receive calls, create contact flows, and run a full contact center in the cloud and it can integrate with CRM systems and other AWS services with no upfront payments and significantly lower cost than traditional contacts and resolutions.
A common architecture is a customer calls in through Connect, Lex recognizes their intent and the Lambda function fulfills the request, for example scheduling an appointment by talking to a CRM.
For your memory hook, think of Lex as the family's conversation list, the one who understands what you're asking for and chats back, whether by voice or text, is the brain behind the chatbot.
Now, let's cover extracting data from documents, which is Amazon TextTrack service.
TextTrack automatically extracts texts, handwriting, and data from any scan document using AI and machine learning.
This goes beyond simple text detection, it can pull structured data out of forms and tables and process pretty much any kind of document, whether it's a PDF or an image.
The use cases are exactly the paper work heavy industries, example financial services, extracting data from invoices and financial reports,
healthcare processing medical records and insurance claims, and the public sector handling tax forms, ID documents and passports.
So whenever a use case involves pulling data or text out of scan documents, forms or images of paper work, remember that's TextTrack service.
There is a subtle distinction worth noting here by the way between TextTrackd
and Recognition's Text Detection.
Recognition finds text within a general image or scene like a street sign in a photo.
Textract on the other hand is specifically about extracting structured data from documents,
like reading all the fields on a form.
Document and forms point to Textract always.
For your memory hook,
Textract is the family's data entry clerk taking a stack of scan forms and documents
and neatly typing all of that out in text and data they contain.
Let's go to Intelligent Search next with Amazon Kendra.
Kendra is a fully managed document search service powered by machine learning.
What makes it special is that it extracts specific answers from within your documents
and it understands natural language questions.
So instead of just returning a list of documents that contain a keyword,
a user can ask a question in plain English like "Where is the IT support desk?"
and Kendra returns the actual answer "first flow" pulled from inside a document.
It can index all sorts of data sources including S3, RDS, Google Drive, SharePoint and more
and it learns from user interactions to promote the best results over time,
which is called incremental learning.
So when a use case is about an intelligent natural language search across an organization's documents
that returns real answers, remember that's Kendra.
For your memory hook,
Kendra is the family's librarian, the one you can ask a plain language question
and who can instantly find exact answers buried inside all of your documents.
Now let's cover recommendations with Amazon Personalize.
Personalize service is a fully managed machine learning service to build apps
with real-time personalized recommendations.
In fact, it's the same technology as Amazon.com that uses its own recommendations.
Think Personalize product suggestions, re-ranking of items and customized direct marketing.
For example, a user buys gardening tools and personalized recommends the next thing they're likely to want.
It integrates into existing websites, applications and even SMS and email marketing systems
and the big selling point is that you can implement it in days rather than months
because you don't have to build, train or deploy your own recommendation model here.
So whenever a use case mentions personalized recommendations, product suggestion
or customers who bought this also bought sort of scenarios, remember that's personalized service.
For your memory hook, think of personalized service as your family's personal shopper.
The one who learns your taste and always knows what to suggest what you like next.
Next, let's cover one more pre-built service called Amazon Forecast for predicting future values.
Forecast uses machine learning to deliver highly accurate forecasts of future data points
based on your historical time series data. Classic use cases are forecasting product demand,
predicting future sales, estimating inventory needs and planning resource requirements.
So when a use case is about predicting future numbers from historical trends over time,
like next quarter's demand for instance, think Amazon Forecast.
For your memory hook, Forecast is the family's fortune teller looking at your history
and predicting what the numbers will be in the future.
Now let's turn to build your own option, which is Amazon SageMaker.
Everything so far has been pre-built specialists that do one job through a simple API.
But what if none of these fit your specific needs and you want to build a completely custom
machine learning model? That's what SageMaker is for.
SageMaker is a fully managed service for developers and data scientists to build, train
and deploy their own machine learning models. Normally doing all of that is difficult. You have
to gather and prepare data, choose and build a model, train and tune it and then provision
service to deploy it. SageMaker brings all that entire machine learning workflow into one place
and manages that heavy lifting for you. Here's a simple picture of what building a model looks like.
And I will use a fun example here predicting your exam score.
You start with historical data, things like how many years of IT experience someone has,
their years of AWS experience and how much time they spend on a course, each bed with the score
they actually got, which is called the label. Now you use that data to build and train a model,
so that the model learns the patterns connecting the inputs to the score.
Then when new data comes in, someone's experience and study time, you apply the model
and it predicts their score, maybe pass with 906. That is build, train and deploy your own
model workflow in the essence of SageMaker. The key exam distinction is simple but important.
If the use case is a common well defined task like image recognition, transcription or translation,
use the relevant pre-built AI service because it's ready to go and requires no machine learning
experience or expertise. Only reach for SageMaker when the requirement is to build, train or deploy
a custom or specialized model that isn't available as a pre-built service. So when you see build a
custom model, train a model or data scientist developing their own machine learning, that relates
to SageMaker. For your memory hook, if all the other services are specialist appliances,
just plug in and use, SageMaker is the fully equipped workshop where you build your own custom
machine exactly to your specifications from the ground up. Let's do a pulse check now to test
that pattern matching skills. Here's the scenario. A media company has a huge library of video content
and wants to automatically generate subtitles in the original language for all of it,
to make the videos accessible and searchable. Which AWS AI services should they use? Let's take a
moment and think about it. The answer is of course Amazon Transcribe Service. The signal is generating
subtitles which means converting the spoken audio into the video into text. That speech text
which is exactly what Transcribe Service does. If they had wanted them to translate those subtitles
into other languages, they had add Amazon Translate on top of it. And if they had wanted to generate
a spoken voice over from the text, that would be poly. But for generating subtitles from speech,
it's Transcribe. Let's do one more pulse check here because this pattern matching is the whole
game in the exam. Here's the scenario. An online retailer wants to show each customer personalised
product recommendations in real time based on their browsing and purchase history without having
to build and train its own recommendation model from scratch. Which service do they use? The answer is Amazon Personalised Service. The signals are
personalized product recommendations in real time and explicitly not wanting to build and train
a custom model here. Personalised is the ready made recommendation service, the same technology amazon.com
uses. If they had wanted to build a fully custom model of their own, that would push towards Sage
Maker, but the requirement specifically says here that they don't want to. So Personalised Service
is the clear fit. So here's the trap to watch out for across this whole topic.
The example describe a use case and offer several AI services as options. The wrong answers
are often services that are close but not quite right. Watch the precise task.
Analyzing images is recognition, but extracting data from a scanned form is text tracked.
Converting speech to text is transcribe, but building a chatbot that understands intent is lex.
Simple translation is translate, but understanding sentiment and meaning is comprehend.
And always ask yourself if a pre-built service covers this particular task before reaching for Sage
Maker, because if a ready made service fits, that's the better answer than building your own model.
Let's do a rapid-fire summary here to lock everything in. I'll give the service and it's
one line signal because this list is your exam cheat sheet. Recognition is for images and video,
detecting objects, faces, text and content moderation. Transcribe is speech to text like subtitles
and call transcription. Polly is text to speech making application stock. Translate is language
translation, comprehend is natural language processing, sentiment analysis and extracting meaning
from text with comprehend medical for clinical text. Lex builds chatbots and conversation bots,
understanding intent and pairs with connect for cloud contact centers. Text tracked extracts text
and structure data from scan documents and forms. Kendra is intelligent natural language document
search that returns real answers, personalized gives real-time personalized recommendations,
forecast predicts future values from historical time series data, and Sage Maker is the build your
own service for developers and data centers to build, train and deploy custom machine learning models.
The one decision that underpins all of it. If a pre-built specialist service matches the task,
use it. Only use Sage Maker when you need to build a custom model that the pre-built services
don't cover. Alright, that's for episode 27 on machine learning and AI services which means
we have now completed every code topic for the exam. We covered the pre-built AI specialist,
recognition for vision, transcribe and poly for speech in and out, translate and comprehend
for language, lex for chatbots, text track for documents, Kendra for search, personalized
for recommendations and forecast for predictions. And we covered Sage Maker finally the build your
own service for custom models. So here's the big takeaway. This topic is pure pattern matching.
Learn the one line signal for each service and recognize that the pre-built services handle
comment tasks with no machine learning expertise required. Whatsoever, while Sage Maker is only
for building custom models. Match the use cases to the service and these questions are some of the quickest
points on the whole exam.
And take a moment because this is worth celebrating,
that's the entire core curriculum done.
Across 27 episodes, you've gone from the fundamentals
all the way through compute storage,
databases, networking, security, integration,
analytics, cost the well-architected framework,
and now AI and machine learning.
You've covered the full breadth of the SAAC03.
What's now left is all about getting you exam ready.
Next, we have got a dedicated exam day strategy episode
where I'll walk you through exactly how the exam works
and the tactics to maximize your score.
If this episode helped the AI services click for you,
please consider to leave a five-star rating on my podcast
and do share it with anyone studying
for the AWS exam.
It generally helps channel grow.
So until next time, keep building, keep learning,
and I will see you for the home stretch.
This is Balu signing off, bye.
Podcast Summary
Key Points:
AWS AI services are primarily tested through pattern matching—given a use case, identify the correct pre-built service based on keywords.
Pre-built AI services like Recognition, Transcribe, Poly, Translate, Comprehend, Lex, Textract, Kendra, Personalize, and Forecast handle common, well-defined tasks without requiring machine learning expertise.
Recognition detects objects, faces, text, and scenes in images or videos and is key for content moderation and visual analysis.
Transcribe converts speech to text (e.g., subtitles, call transcription), while Poly turns text into speech with customization options for pronunciation and tone.
Comprehend performs natural language processing, including sentiment analysis and entity extraction; Comprehend Medical specializes in clinical text analysis.
Lex enables chatbots and voice assistants by understanding intent, often paired with Amazon Connect for contact centers.
Textract extracts structured data from scanned documents like forms and invoices, distinct from Recognition’s general text detection.
SageMaker is only needed when building custom, trained machine learning models—otherwise, pre-built services are sufficient and preferred.
Summary:
This episode covers AWS AI and machine learning services, emphasizing that the exam focuses on pattern matching rather than technical implementation. The majority of questions ask whether a pre-built AI service matches a given use case—such as speech-to-text, translation, or sentiment analysis—without requiring knowledge of machine learning models or coding. Services like Amazon Transcribe (speech-to-text), Polly (text-to-speech), Translate (language translation), and Comprehend (natural language understanding) are tested through clear signals.
For conversational AI, Lex is used to build chatbots and voice assistants. Textract extracts data from scanned documents, while Kendra enables intelligent, natural language searching across documents. Amazon Personalize delivers real-time product recommendations, and Forecast predicts future demand from historical trends.
The key decision rule is: use a pre-built service when the task is standard and well-defined; only use SageMaker when custom model development is required. This topic is one of the most accessible on the exam, making it a high-confidence area for candidates who learn the core signals and mental models. With this foundation, learners can confidently tackle AI-related questions and complete the full SAAC03 curriculum, paving the way for exam-specific strategy in future episodes.
FAQs
The exam focuses on pattern matching: given a use case, identifying the correct pre-built AI service without needing to write code or understand machine learning math.
Amazon Transcribe converts audio to text using automatic speech recognition, ideal for subtitles, call transcription, and closed captions.
Amazon Polly converts text to speech with customization options like pronunciation and tone, making it ideal for voice-enabled applications.
Amazon Translate provides accurate, scalable language translation for localizing websites and applications for global audiences.
Amazon Comprehend uses natural language processing to detect sentiment, extract entities, and identify topics from text content.
Amazon Lex combines speech recognition and natural language understanding to create conversational interfaces, often used with Amazon Connect.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.