Building a Unified Catalog with Amazon DataZone
Publication date: November 20, 2024 (Diagram history)
The unified data catalog acts as a central repository for all your organization's data assets. This architecture shows how Amazon DataZone breaks down data silos by bringing data from various sources together and fostering improved search functionality and trust in data.
Building a Unified Catalog with Amazon DataZone
The following steps describe the architecture:
-
Extract data using Amazon AppFlow, AWS Glue, Amazon Kinesis, Amazon Managed Streaming for Apache Kafka (Amazon MSK), or AWS Database Migration Service (AWS DMS).
-
Based on end user requirements, store the data in Amazon S3, Amazon Redshift, or purpose-built databases like Amazon Aurora or DynamoDB. Design data lakes to store raw structured, semi-structured, and unstructured data at low cost.
-
Crawl structured and semi-structured data from Amazon S3 using AWS Glue Crawler, which writes metadata to AWS Glue Data Catalog. Perform data quality
checks on catalog tables at rest using AWS Glue Data Quality. Use Java Database Connectivity (JDBC) data sources to crawl and catalog the data. -
The Amazon DataZone domain acts as the central catalog repository hub. Use the AWS Glue Data Catalog data source to onboard existing data assets from data lakes and databases. With Amazon Redshift, automatically extract technical metadata of database tables and views to Amazon DataZone.
-
Use the Amazon DataZone data portal
to discover, catalog, share, and govern data in a self-serve fashion. -
Organize Amazon DataZone entities under different levels of hierarchy using domain units.
-
Amazon DataZone data projects provide a collaborative space where members onboard data assets from their respective business units. Business glossaries define business terms associated with data assets. Metadata forms augment business context of asset metadata. Custom assets expand the catalog beyond predefined system assets.
-
Catalog unstructured data assets stored in Amazon S3 and publish them to Amazon DataZone by tagging them with relevant custom asset types.
-
Every data asset onboarded to Amazon DataZone is part of Inventory. Enrich the data asset with business catalogs, data lineage, and data quality to aid discoverability.
Further reading
For additional information, refer to the following resources:
Diagram history
To be notified about updates to this reference architecture diagram, subscribe to the RSS feed.
| Change | Description | Date |
|---|---|---|
Initial publication | Reference architecture diagram first published. | November 20, 2024 |
Note
To subscribe to RSS updates, you must have an RSS plugin enabled for the browser you are using.