This whitepaper is for historical reference only. Some content might be outdated and some links might not be available.
Further reading
For additional information, refer to:
-
Storage Best Practices for Data and Analytics Applications (AWS whitepaper)
-
Deploy data lake ETL jobs using CDK Pipelines
(blog post) -
Orchestrate multiple ETL jobs using AWS Step Functions and AWS Lambda
(samples on GitHub) -
Orchestrate Apache Spark applications using AWS Step Functions and Apache Livy
(blog post) -
Building complex workflows with Amazon MWAA, AWS Step Functions, AWS Glue, and Amazon EMR
(blog post) -
Field Notes: How to Build an AWS Glue Workflow using the AWS Cloud Development Kit (AWS CDK)
(blog post) -
Setting up automated data quality workflows and alerts using AWS Glue DataBrew and AWS Lambda
(blog post) -
Discovering metadata with AWS Lake Formation: Part 1
(blog post) -
Discovering metadata with AWS Lake Formation: Part 2
(blog post)