Serverless Data Processing with Dataflow

In this 3-day course we’re taking an advanced look at Dataflow for big data practitioners who want to take their data processing applications further.

The Serverless Data Processing with Dataflow course begins with the foundations, explaining how Apache Beam and Dataflow work together to meet your data processing needs without the risk of vendor lock-in. The section on developing pipelines covers how to convert business logic into data processing applications that run on Dataflow, and the course culminates in operations, with the key lessons for running a data application on Dataflow, including monitoring, troubleshooting, testing and reliability.

Across 21 modules and labs, participants select the right IAM permissions for a Dataflow job, implement best practices for a secure data processing environment, select and tune I/O for their pipelines and use schemas to simplify Beam code and improve performance. They also develop Beam pipelines using SQL and DataFrames and work on monitoring, troubleshooting, testing and CI/CD for Dataflow pipelines.

The course is aimed at data engineers, data analysts and data scientists aspiring to develop data engineering skills, and is available instructor-led and on demand. Participants should have completed Building Batch Data Pipelines and Building Resilient Streaming Analytics Systems.

← All Google Cloud courses