Transportation and Logistics Data Lake
Publication date: March 15, 2020 (Diagram history)
This architecture shows how to break data silos and create a single repository for transportation and logistics data from multiple sources. You can use a fully managed, pay-per-use, and scalable architecture to remove the complexity and costs of managing on-premises databases.
Transportation and Logistics Data Lake
The following steps describe the architecture:
-
Ingest data from multiple sources: historical and large databases with AWS Snowball, real-time data with Amazon Kinesis, ERP data synchronization with AWS DataSync, and mobile app data with Amazon AppSync.
-
Automatically extract, transform, and load raw data with Lake Formation. Discover schemas and create a data catalog for further data exploration.
-
Create an Amazon Simple Storage Service data lake where you can store, organize, and access business-critical data securely. Use Amazon EMR for fine-grained control over extraction, transformation, and loading (ETL) jobs. Use AWS Glue to manage your data catalog.
-
Turn raw data into actionable knowledge with Athena and Amazon Quick Sight. Disseminate key business intelligence insights across your organization. Use SageMaker AI to deliver ML-based predictive and prescriptive analytics.
-
Use AWS fully managed databases to deliver specific capabilities: Amazon Neptune graph database for network optimization, Amazon Aurora for fast relational queries, and Amazon DynamoDB for interactive mobile and web apps with real-time updates. Avoid data latency with live data stream ingestion from Amazon Kinesis.
Further reading
For additional information, refer to the following resources:
Diagram history
To be notified about updates to this reference architecture diagram, subscribe to the RSS feed.
| Change | Description | Date |
|---|---|---|
Initial publication | Reference architecture diagram first published. | March 15, 2020 |
Note
To subscribe to RSS updates, you must have an RSS plugin enabled for the browser you are using.