View a markdown version of this page

Optimizing data layout in Amazon S3 - Amazon Aurora

Optimizing data layout in Amazon S3

Aurora PostgreSQL scans columnar data in Amazon S3, so your data layout determines how much data each query reads. Follow these practices.

  • Avoid many small files: Consolidate data into fewer, larger files. Reading many small files adds per-file metadata and Amazon S3 request overhead.

  • Use a reasonable row group size: Row groups let the engine read and skip data in chunks using min/max statistics. Very small row groups add per-group overhead, while very large ones reduce how precisely the engine can skip data that a query does not need.

  • Partition on your filter columns: Partitioning splits data into separate paths by column value, so a query that filters on a partition column skips entire partitions without reading them.

For detailed guidance, see Optimize data in the Amazon Athena User Guide.