View a markdown version of this page

Building a Unified Catalog with Amazon DataZone - Building a Unified Catalog with Amazon DataZone

Building a Unified Catalog with Amazon DataZone

Publication date: November 20, 2024 (Diagram history)

The unified data catalog acts as a central repository for all your organization's data assets. This architecture shows how Amazon DataZone breaks down data silos by bringing data from various sources together and fostering improved search functionality and trust in data.

Building a Unified Catalog with Amazon DataZone

Architecture diagram showing a unified catalog with Amazon DataZone, AWS Glue, Amazon S3, Amazon Redshift, and Amazon Aurora.

The following steps describe the architecture:

  1. Extract data using Amazon AppFlow, AWS Glue, Amazon Kinesis, Amazon Managed Streaming for Apache Kafka (Amazon MSK), or AWS Database Migration Service (AWS DMS).

  2. Based on end user requirements, store the data in Amazon S3, Amazon Redshift, or purpose-built databases like Amazon Aurora or DynamoDB. Design data lakes to store raw structured, semi-structured, and unstructured data at low cost.

  3. Crawl structured and semi-structured data from Amazon S3 using AWS Glue Crawler, which writes metadata to AWS Glue Data Catalog. Perform data quality checks on catalog tables at rest using AWS Glue Data Quality. Use Java Database Connectivity (JDBC) data sources to crawl and catalog the data.

  4. The Amazon DataZone domain acts as the central catalog repository hub. Use the AWS Glue Data Catalog data source to onboard existing data assets from data lakes and databases. With Amazon Redshift, automatically extract technical metadata of database tables and views to Amazon DataZone.

  5. Use the Amazon DataZone data portal to discover, catalog, share, and govern data in a self-serve fashion.

  6. Organize Amazon DataZone entities under different levels of hierarchy using domain units.

  7. Amazon DataZone data projects provide a collaborative space where members onboard data assets from their respective business units. Business glossaries define business terms associated with data assets. Metadata forms augment business context of asset metadata. Custom assets expand the catalog beyond predefined system assets.

  8. Catalog unstructured data assets stored in Amazon S3 and publish them to Amazon DataZone by tagging them with relevant custom asset types.

  9. Every data asset onboarded to Amazon DataZone is part of Inventory. Enrich the data asset with business catalogs, data lineage, and data quality to aid discoverability.

Further reading

For additional information, refer to the following resources:

Diagram history

To be notified about updates to this reference architecture diagram, subscribe to the RSS feed.

ChangeDescriptionDate

Initial publication

Reference architecture diagram first published.

November 20, 2024

Note

To subscribe to RSS updates, you must have an RSS plugin enabled for the browser you are using.