#15 AWS IAM Series: Part 3 - AWS Storage & Database Services
36m 24s
This episode, part of an AWS Identity Management series, explores additional AWS storage services. Amazon CloudFront is described as a global content delivery network that caches content at edge locations for fast, secure access worldwide. AWS Storage Gateway acts as a hybrid bridge, seamlessly connecting on-premises storage to AWS cloud storage via file, volume, or tape gateways. AWS Snowball is a physical, secure appliance for efficiently transferring large data volumes (terabytes) when internet transfer is impractical. The host differentiates storage services (focused on data storage/retrieval) from database services (for structured data management with queries), introducing Amazon RDS as a managed relational database supporting various engines like MySQL and PostgreSQL, with features like Multi-AZ for high availability. Throughout, practical analogies are used to simplify concepts, stressing understanding over memorization for real-world application, alongside personal anecdotes about work and conference speaking.
Hello and welcome to another episode of the Identity Navigator. My name is Rohit. This is the third part of our AWS IM series in which we are discussing about basics of AWS AWS IM and comparison with on-premise IM concepts. Jumping to Cloud IM could be overwhelming. So that is why in this multi-part series, we are diving into Cloud Computing, AWS Cloud, AWS IM and comparing and contrasting the Cloud concepts with the legacy concepts, giving you a foundational knowledge of all things Cloud IM. In the last episode, we talked about some more storage services and storage services are like a virtual warehouse where you can store all kinds of digital stuff. It is like having a giant locker in the cloud where you can keep your files and photos, videos and more instead of storing them on your computer or phone. So last week, we deep-dived into S3, glacier, EC2 instance storage, elastic block storage, elastic file system. Today, let's look into a couple more storage services like Cloud Front, AWS Storage Gateway, Snowball. And then let's start looking into the third foundational AWS services, which is the database. Remember, we spoke about the four foundational database services and then we will be deep diving into AWS IM. Before we start looking into Cloud Front, on the personal level, it has been a super busy week. I am speaking at European identity and Cloud conference this week. So preparing for that is taking quite some time. However, preparation is not an issue. I would be remote as you know, this European conference is happening in Berlin. So an IM based out of East Coast in the US. So my session would be around five in the morning and that is what is stressing me out that, hey, I first need to wake up and then need to be ready with the presentation. So keeping the fingers crossed, let's see. Apart from that, what happened? I did have a call with the vendor this week at work and I was late to the call. I was running into some issues that needed my attention. These were production issues and then as I joined and I apologized that here I was running into some issues. I'm really sorry and what not. And the first thing that they said was hopefully access issues because this is what we saw. And I can tell you that that wasn't really maybe the best first impression. So, hey, if you are on the sales side, you all know that I have a ton of respect for you. All I did sales and consulting for more than 15 years, but don't wish ill on the prospects. That is all I could say. But I'm not an expert in sales. I will stick to the facts that I know about which is identity and access management. So let's look into CloudFront. So think of Amazon CloudFront like a super fast delivery service that ensures your online content like websites and videos and files reaches user quickly no matter where they are in the world. So how do we understand it using a real-world analogy? So imagine a global delivery network. Imagine you have a network of warehouses around the world. Each warehouse stores copies of your product. So when someone orders a product, it gets shipped from the warehouse closest to them. This way the delivery is much faster than shipping everything from a single distant location. So when you put your content like web page, videos or downloads on CloudFront, it works like those warehouses. CloudFront store copies of your content in many locations around the world called edge locations. When someone tries to access your content, CloudFront delivers it from the nearest edge location making it load much faster. So just like any other delivery service, also Amazon delivery service, your packages are secure and arrive on time. CloudFront ensures that your content is delivered quickly and securely. It protects your contents from attacks and ensures it is always available even during high traffic periods. So imagine during a big sale, many people order advance your global warehouses handles the rush without delay. Similarly, CloudFront can handle heavy traffic loads efficiently ensuring that many users can access your content simultaneously without slowdowns. So key features are the content is from the closest location to the user making it faster to access. It has low latency and high speed. Obviously, scalability would be another feature and then security is another thing that they speak about. So why use it? It's pretty obvious at this point for faster content delivery, for improved user experience, global reach and then reliability and security. Let's look into some of the other key facts. So this is a content delivery network service which means of which provides a means of distributing your source data of your web traffic closer to the end user requesting the content via AWS Edge Location as CacheData. As a result, it doesn't provide durability of your data. AWS Edge Locations are sites deployed in highly populated areas across the globe to cache data and reduce latency for and end user access. It uses distributions to control which source data it needs to distribute and to ware. And there are two delivery methods that exist to distribute this data. It could be a web distribution or it could be an RT MP distribution. So the distribution requires an origin containing your source data such as an S3 and data can be distributed using the. Edge location options like US Canada and Europe, US Canada Europe and Asia or any other Edge location. Additional encryption security can also be added by specifying an SSL certificate that must be used within the distribution. Pricing for a cloud front is based on data transfer costs and STTP request. So when you think about Amazon cloud front, think of a global delivery content network for your online content. Let's look into another storage service. AWS storage gateway. Now think of it like a bridge that connects your local storage to the cloud. Using you to store and retrieve data easily between your on-premise environments and AWS. Let's let's use a real word analogy. So imagine a virtual storage bridge. Imagine you have a storage room in your office, which is the local storage and you also rent a storage unit in a secure facility across town. This is cloud storage. So the AWS storage gateway acts like a delivery service that moves item between your office storage room and the remote storage unit seamlessly. So there are definitely type of this Amazon storage gateways, Amazon offers different type of delivery services and you can pick and choose depending on your needs like. Let's look into three of them file gateway. So this is like a file cabinet that stores file locally, but also send copies to the storage unit across town for backup. So more of a backup thing going on here and you can access this files locally and remotely. Volume gateway. Now think of it as having a virtual hard drive in your office that also keep copies of your data in the remote storage unit. So you can use it for data that you need to access frequently and then tape gateway. Imagine you have old tapes or archives in your storage room and you want to digitize them and keep backups in the remote storage unit. So this helps you move these tape backups to the cloud. Right.
Again, as we have spoken about, we don't need to memorize all of these things. When we actually need to work on it, we should be aware of, okay, there is a high level services like this that AWS provides. We should be able to identify what those services are from the service catalog and then use them as per our use cases. But memorizing them will serve no purposes. As all the other AWS service, the storage gateway is secure and reliable. This ensures that all your data transfers are secure and reliable. It also makes sure that your data is safely moved to and from the cloud without loss or corruption. And you don't need to change how you work in your office. The storage gateway integrates with your existing systems. So you can continue using your local storage while benefiting from cloud storage facilities like scalability and duration. Key features of AWS storage gateway. Obviously, this is a hybrid storage solution. It bridges your local storage and AWS cloud storage. It has got multiple gateway types. We discussed three of them, file, volume and tape. The data transfer is obviously secure and the integration is seamless. That means you could continue working the way that you are and you don't have to worry about the integrations or the data transfers happening behind the scene. But those behind the scene things would give you features like scalability and durability. So why use AWS storage gateway? Please extend local storage. That's number one. It is also used for backup and archive. Right. Could you think of a use case like ransomware? Could it be used in ransomware cases disaster recovery and then cost efficiency optimizing storage cost by using cloud storage for infrequently access data and backups? So other key facts are it allows you to provide a gateway between your own data center storage system such as your sand, NAS or DAS and Amazon S3 in Glacier on AWS. So sometimes I struggle with is this truly a storage. But then when we are thinking about storage services, I really wanted to bring up AWS storage gateway because it is not a storage. It's an storage service right? It's that bridge between your local storage and all the other types of AWS cloud storage that we have spoken about so far. So this is just a software application downloaded as a VM and installed within your data center. And then a local on premise cache is also provision for accessing your most recently access file. So AWS storage gateway acts like a bridge or delivery services that connects your local storage to AWS cloud storage. It helps you extend backup and archive your data seamlessly and securely providing a hybrid solution that combines the best of local and cloud storage. So for your on-premise applications that need to archive data for audit requirements or for any other requirements, which is very infrequently used if ever. Can you think of using something like an Amazon storage gateway, having a VM on premise, maybe adding it to AWS S3 or Glacier and then sending your data there for backup. So there could be multiple use cases. Obviously that you could think of in terms of what you are trying to solve for today. Let's look into one last storage services snowball. Think of AWS snowball like a secure heavy duty suitcase designed to move large amount of data between your office and the cloud. So you got the point as to why I broke down the last episodes and stop there and why we are talking about these now because these are not typically storages, but these are storage services that are absolutely necessary to know in order for us to be able to use in cases which are hybrid. And we are storing data both on local and on cloud systems, but also for high availability like in Netflix requiring a content delivery network or an Amazon shopping website or an e-commerce or shopping website requiring that. So it's just not about the storage, but also the services that are associated with storage. So how do we understand snowball using a real word analogy? So imagine a secure data suitcase. You think of snowball you think a suitcase. Imagine you have a lot of important documents, photos and videos that you need to send to a remote location. If you try to send all of this data over the internet, it would take forever and might not be secure. AWS snowball acts like a sturdy secure suitcase that you can fill with your data and ship to the cloud. This suitcase is designed to hold large amount of data. I say large is truly as large it could store terabytes of data. It is built to be tough and durable protecting your data during transport. It's like having a robust storage box that can withstand bumps and drops. The transfer is secure. This suitcase has a secure lock and encryption ensuring that your data is protected from unauthorized access and only you and the cloud storage service can unlock and access the data inside. It is super easy to use. You simply connect the suitcase to your local system, transfer your data into it and then ship it to the cloud provider. When the cloud provider receives it, they load your data into cloud storage and then erase this suitcase securely. And then it is very efficient for large data transfers. So using this suitcase is much faster and more efficient for transferring huge amount of data compared to sending it over the internet. It is perfect for situations where you have large data sets that need to be moved quickly and securely. It reminds me of early 2000s. When I used to work a lot with virtual machines and my friends and some colleagues of mine used to post a ton of hard disks or hard drives, external hard drives to each other just because it was easier for us to have those 200 or 300 gigs of hard drives because you know the download it sometimes stopped, it failed and it was just an easier way of doing it. Probably not the most efficient but it worked for us at that time. Now key features or AWS snowball is obviously it is high capacity storage. We spoke about terabytes. It is durable and secure. The data transfer is fast or faster than transferring large data volumes over the internet and it is easy to use. Just connect, transfer and ship. Now why use AWS snowball? Large data migration. So ideal for moving large data sets to the cloud quickly and securely. Data backup and archiving. So efficient for backing up or archiving significant amount of data to the cloud disaster recovery and then network limitations sometimes. Sometimes it, the DC could be in a location where there is slow or unreliable internet connection which cannot be used for large data transfers. Which was definitely true in my case when I was using usps to send out those VMs. One of the other key facts for AWS snowball is it could send the data from on-premise data center to Amazon S3 and also from Amazon S3 back to your data center using a physical appliance called or known as a snowball. So that suitcase is actually called a snowball. It is a two way traffic. You can send it from the cloud to your local system, from local system to your cloud. A snowball is basically an appliance. So it comes as either a 50 terabyte or an 80 terabyte storage device and is fully dust, water and temper resistant. It's been designed obviously for high speed data transfers. data to the world.
answer is automatically encrypted. This is also HIPAA compliant. So well done, Amazon. And as a general rule, if your data retrieval will take longer than a week using your existing connection method, then you should consider using AWS Snowball. So think of it like secure, high capacity, suitcase that helps you move large amount of data to the cloud quickly and safely. This is a physical appliance like AWS Storage Gateway was a software application. It is an hardware appliance that you would receive. So we spoke about multiple AWS services and also tried to create a word association. So we don't have to memorize all of this but get a sense of what are the different or what are some of the different types of storage services that are out there. So Amazon S3, we imagined a secure or storage locker. Amazon Glacier, imagine storage basement, EC2 instance storage, your computer hard drive, elastic block storage, an external hard drive, elastic file system, a shared office drive, cloud front, cloud delivery network, AWS Storage Gateway. We just discussed this, a virtual storage bridge and then Snowball. Can you think of what Snowball is? A secure data suitcase. Now we have been talking about differences or I have been asking you about differences between an AWS Storage and AWS Databases. Low and behold, we are at the point where we need to talk about this. I'm sure most of you have already figured it out but if not, no worries. Let's look into the purposes of these two distinct services. The storage services primarily focus on storing and retrieving data. Data could be your file, an object, a block storage with varying performance and cost characteristics. Think of this as a place to keep your stuff safe and accessible like closets, boxes or shared drives. Database services, on the other hand, focuses on managing data with specific data models. It could be relational, key value, document, graph and providing query capabilities, data integrity and transactional support. These are systems for organizing, managing and finding specific pieces of information like organized filing cabinets, index card systems or a ledger. In terms of data models, a storage service is a general purpose storage without intrinsic query capabilities. You can query but only up to a level. Whereas on the Database services, it is structured and unstructured data model but it will always have an advanced query and transactional capabilities. But these are all theoretical differences. In terms of actual use cases, think of storage services when you are thinking of backup, recovery, content storage, content delivery, files here and data leaks. Use these services when you need to store data, file backups or large amount of data. Whereas in Database applications, think of Database services when you are thinking about application data management, real-time analytics, complex queries and transactional systems. Use this when you need to manage and query information like keeping track of customers, sales or other structured data. It brings us to our third type of Amazon foundational service, which is Database. We will only be talking about a couple of them and there is a reason this would give you more than enough of an understanding to identify what exact AWS Database service would you need when you have a use case or problem to solve. Again, we are not looking to memorize any of this stuff. We are just trying to create a basic level of understanding so we can find our way as and when a work around these AWS services is given to us. So the first type of Database service that we will talk about is RDS or Amazon relational Database services or service. This is as the name specifies a relational Database service that provides a simple way to provision, create and scale a relational Database within AWS. relational Database comes from obviously set theory. You all know about it. It's definitely the RDS or the Amazon relational Database service. It is a managed service and you can select from a range of different Database engines. You can have it a MySQL, you can have a MariaDB, you can have PostgresSQL or an Amazon Aurora. Amazon Aurora is actually AWS own fork of MySQL or you can have an Oracle or a SQL server. So you can choose from multiple different flavors or different engines of relational Databases that you use in your everyday life in your on-premise environment today. Now when you create your RDS Database, you must select an instance to support your Database from a processing and memory perspective. However, if HA and resiliency is of importance when it comes to your Database, it could be a production system. Then you might want to consider a feature which is known as multi-az, which stands for multi-avability zones. When multi-az is configured, a secondary RDS instance is deployed within a different az when the same region as within the same region as the primary instance. So again, this is a mouthful. All you need to know is that if you need HA and resiliency, there is a feature. feature is called multi-az but you don't need to remember that. The primary purpose of the second instance, if you haven't already guessed, it's to provide a failover option and the replication of data happens synchronously. Now it is not possible to talk about Database and not talk about scaling. So when it comes to scaling your storage, there is, it's not a possibility that there is no feature. So there is definitely a feature which is called auto scaling. So my SQL post-griss SQL, MariaDB, Oracle, SQL server all use elastic block storage. Can you name the word that we used or associated with elastic block storage? It was an external hard drive. So for both data and log storage, however, Amazon Aurora uses a shared cluster storage architecture and does not use EBS or elastic block storage. At any point, you can scale your RDS database is vertically changing the size of your instance. From a backup perspective, by default, RDS provides an automatic feature. This is enabled on all new RDS databases, which backs up your RDS databases to Amazon S3. You are able to configure the level of retention in days from 0 to 35 and implement a level of encryption using the key management services, a service or KMS. So this is different from HA. This is a backup. You can also perform manual backups anytime you need to. This is thankfully called as snapshots. Overall, I believe Amazon has done a really good job at naming these services. Now this is a philosophical discussion. I don't want to get into it, but if you don't feel like it, please do let me know. I would be happy to change my mind. If you can justify it to me, but for the most part, as long as I as much as I have seen, I believe they have done a good job. Now these snapshots are the manual backups are not bound by retention periods, set an automatic backup configuration and you can delete it only through a manual process. So automatic
backup 35 0 to 35 manual backups or snapshots there is no retention period. Now when using a MySQL compatible or error database you can also use a feature called as Backtrack. I love this feature because this allows you to go back in time on the database to recover from an error or incident without having to perform a restore and create another database cluster. When it allows you to enter a number of hours of how far you would like to backtrack with a maximum of 72 hours. I personally feel like this should be a default feature for every relational database out there. It would have saved me so many rollbacks, procedures and backups and what not. Another database service that we will talk about is the Amazon DynamoDB. Now I specifically chose it because when we speak of database relational database is no longer the database of choice and Amazon DynamoDB is in no SQL database. That is why I specifically picked it up because we need to look into some of the relational databases but also in some of this category of database which is no SQL or the category should be called as a key value stores. This very much like RDS is a fully managed service. For disk space DynamoDB will automatically allocate space for your table as it grows. However, you do need to reserve capacity for input and output for reads and writes RCU and WCU which is read capacity unit and write capacity unit. But you can choose it from provisioned and an ad hoc or demand. The last main point of the configuration allows you to set encryption of your table which is enabled by default for data at rest. And I think I misspoke because the RDS is a managed service but the DynamoDB is a fully managed service and this is one of the biggest advantage of it. DynamoDB tables are schemaless so you don't have to define the exact data model in advance. The data model can change automatically to fit your application's needs. It is by design highly available and your data is automatically replicated across three different A's with an ad hoc or a reason. And it is designed to be fast, read and write just takes a few milliseconds to complete and DynamoDB will be fast no matter how large your table grows. And unlike relational databases which can slow down as the table gets large, DynamoDB performance is constant and stays consistent even with tables that are many terabytes in size. Now there are definitely some downsides otherwise why will not every application developer use this type of databases. So there are trade-offs obviously like everything else in software. Eventual consistency. Now since data is replicated over three A's in a region thus mismatch can occur. DynamoDB queries aren't as flexible as what you can do with SQL. It also has some strict limitations in the way that you are allowed to work with it. Two important limitations are the maximum record size of 400 kb and the limit of 20 global indexes and 5 secondary indexes per table. So don't get into that blank data approval of A we need to use no SQL because it just makes sense. So we need to use case and identify relational database could be making sense as well. And finally although DynamoDB performance can scale up as your needs grow your performance is limited to the amount of read and write throughput that you are provisioned for each table. So we are now or we have now so far discussed the three different AWS foundation services. We started out with compute thinking of it as CPU and RAM. We then spoke about storage. We have spoken about databases today. Let's pick up the network in the next episode and then we will look into AWS IAM. Also we are going to have a vendor spotlight very very soon. So if you are a vendor in the identity and access management or the security space and you would like to feature on this podcast please reach out to me. Thank you for listening. Please keep your review and your feedback coming. I read each one of them and they have been absolutely fabulous and very very kind. You can always reach out to me via LinkedIn or you can email me at the identity navigator at gmail.com. Until next time this is Rohit your identity navigator. (upbeat music)
Podcast Summary
Key Points:
The episode continues a series on AWS Identity Management, focusing on additional storage services: CloudFront (a global content delivery network), AWS Storage Gateway (a hybrid storage bridge), and AWS Snowball (a physical appliance for large data transfers).
It distinguishes between AWS storage services (for storing/retrieving data like files) and database services (for managing structured data with query capabilities), introducing Amazon RDS as a managed relational database service.
The host shares personal anecdotes about conference preparation and a vendor call, emphasizing practical insights over memorization and the importance of foundational understanding for real-world use cases.
Summary:
This episode, part of an AWS Identity Management series, explores additional AWS storage services. Amazon CloudFront is described as a global content delivery network that caches content at edge locations for fast, secure access worldwide. AWS Storage Gateway acts as a hybrid bridge, seamlessly connecting on-premises storage to AWS cloud storage via file, volume, or tape gateways.
AWS Snowball is a physical, secure appliance for efficiently transferring large data volumes (terabytes) when internet transfer is impractical. The host differentiates storage services (focused on data storage/retrieval) from database services (for structured data management with queries), introducing Amazon RDS as a managed relational database supporting various engines like MySQL and PostgreSQL, with features like Multi-AZ for high availability. Throughout, practical analogies are used to simplify concepts, stressing understanding over memorization for real-world application, alongside personal anecdotes about work and conference speaking.
FAQs
Amazon CloudFront is a content delivery network (CDN) service that distributes your content globally via edge locations. It stores copies of your content in multiple locations worldwide, delivering it from the nearest edge location to users for faster access and lower latency.
AWS Storage Gateway is a hybrid storage service that acts as a bridge between on-premises environments and AWS cloud storage. It offers three main types: File Gateway for file storage and backup, Volume Gateway for block storage, and Tape Gateway for archiving tape backups to the cloud.
AWS Snowball is a physical appliance used for securely transferring large amounts of data (terabytes) between on-premises systems and AWS cloud storage. It is ideal for data migration, backup, or archiving when internet transfer would be too slow or unreliable, typically if transfers would take over a week online.
Storage services focus on storing and retrieving data like files or objects, often for backups or content delivery. Database services manage structured or unstructured data with query capabilities, data integrity, and transactional support, used for application data management and real-time analytics.
Amazon RDS (Relational Database Service) is a managed service for provisioning and scaling relational databases in AWS. It supports multiple database engines like MySQL, PostgreSQL, and Amazon Aurora, and offers features like Multi-AZ deployment for high availability and auto-scaling for storage.
Use Amazon CloudFront for faster content delivery, improved user experience, global reach, reliability, and security. It efficiently handles high traffic loads and protects content from attacks, ensuring quick and secure access from anywhere in the world.
Chat with AI
Loading...
Pro features
Go deeper with this episode
Unlock creator-grade tools that turn any transcript into show notes and subtitle files.