Data Engineering on AWS

In this 3-day course we take a deep dive into designing, building, optimizing, and securing data engineering solutions on AWS.

The Data Engineering on AWS course is designed for professionals who want to architect and manage modern data solutions at scale. It starts with the role of the data engineer, data personas, data discovery, and the AWS tools for orchestration, security, monitoring, CI/CD, infrastructure as code, networking, and cost optimization.

From there, the course moves through the main building blocks. Participants design and implement a data lake, covering storage, ingestion, cataloging, transformation, and serving, and then optimize and secure it with open table formats and AWS Lake Formation. Day 2 focuses on data warehouses with Amazon Redshift, including Redshift Serverless, performance optimization, and access control, followed by batch data pipelines. On Day 3 participants optimize, orchestrate, and secure batch pipelines, and then turn to streaming data architectures, including their optimization, security, and compliance considerations.

The course includes presentations, demonstrations, group exercises, and many hands-on labs, such as setting up a data lake, setting up a data warehouse with Amazon Redshift Serverless, orchestrating Spark processing with AWS Step Functions, streaming analytics with Amazon Managed Service for Apache Flink, and access control with Amazon Managed Streaming for Apache Kafka.

The course is aimed at professionals interested in designing, building, optimizing, and securing data engineering solutions on AWS. It is recommended that attendees are familiar with basic machine learning concepts, have working knowledge of Python and libraries such as NumPy, Pandas, and Scikit-learn, and have a basic understanding of cloud computing and AWS. Familiarity with SQL and relational databases is helpful but not mandatory, and experience with Git is beneficial but not required.

← All AWS courses