Data Engineering on Google Cloud

In this 4-day course we’re taking a hands-on look at designing and building data processing systems on Google Cloud.

The course uses lectures, demos and labs to show how to design data processing systems, build end-to-end data pipelines, analyze data and implement machine learning, across structured, unstructured and streaming data. It brings together four Google Cloud courses: Introduction to Data Engineering, Build Data Lakes and Data Warehouses, Build Batch Data Pipelines, and Build Streaming Data Pipelines. Along the way you will work with the extract-and-load, ELT and ETL pipeline patterns, data lakehouse concepts with BigQuery and BigLake, batch pipelines and their orchestration, and streaming pipelines.

The course includes 18 labs, covering topics such as loading data into BigQuery, building SQL workflows in Dataform, batch pipelines in Cloud Data Fusion, and a streaming pipeline for a real-time dashboard with Dataflow.

The course is aimed at data engineers, data analysts and data architects. Attendees should understand data engineering principles such as ETL/ELT, data modeling and common formats (Avro, Parquet, JSON), be familiar with data warehouse and data lake concepts, be proficient in SQL and a programming language (Python recommended), and know the command line and core Google Cloud concepts.

← All Google Cloud courses