
Building Modern Data Lakes with Amazon S3 Tables and Apache Iceberg
Learn how Amazon S3 Tables and Apache Iceberg are changing modern data engineering on AWS, from scalable table storage and automated maintenance to real-time streaming and analytics
Building Modern Data Lakes with Amazon S3 Tables and Apache Iceberg
Modern data engineering is moving beyond traditional file-based data lakes. Teams increasingly need data platforms that can handle large datasets, support analytics and AI workloads, and reduce the operational effort required to maintain pipelines.
One of the interesting developments in the AWS data ecosystem is Amazon S3 Tables, which provides managed storage for Apache Iceberg tables on Amazon S3.
Why S3 Tables + Iceberg matters
Apache Iceberg is an open table format designed for large analytic datasets. It provides table-level capabilities such as schema evolution, snapshots, and time-travel queries while working with engines such as Spark.
Amazon S3 Tables builds on this approach by providing purpose-built Iceberg table storage with automated maintenance capabilities such as file compaction and snapshot management.
This gives data engineers a simpler architecture:
Data Sources → Ingestion → S3 Tables / Iceberg → AWS Glue Catalog → Athena / EMR / Redshift → Analytics & AI
What changes for a Data Engineer?
In a traditional data lake, engineers may need to manage many separate concerns:
- File organization
- Table metadata
- Small-file problems
- Compaction jobs
- Catalog integration
- Data processing
- Governance
- Table metadata
- Small-file problems
- Compaction jobs
- Catalog integration
- Data processing
- Governance
With S3 Tables, AWS manages important table-maintenance operations, allowing engineers to focus more on data pipelines, quality, governance, and business requirements.
AWS Glue integration
AWS Glue can run ETL jobs against S3 Tables using Apache Spark. AWS currently recommends integrating S3 Tables with the AWS Glue Data Catalog when building production pipelines that use services such as Athena, EMR, or other AWS analytics services. S3 Tables support in AWS Glue requires Glue 5.0 or later.
A simplified pipeline could look like this:
Source Systems
↓
Amazon Kinesis / Applications / Databases
↓
Ingestion
↓
Amazon S3 Tables
(Apache Iceberg)
↓
AWS Glue
↓
Athena / EMR / Redshift
↓
Dashboards / Analytics / AI
↓
Amazon Kinesis / Applications / Databases
↓
Ingestion
↓
Amazon S3 Tables
(Apache Iceberg)
↓
AWS Glue
↓
Athena / EMR / Redshift
↓
Dashboards / Analytics / AI
Real-time data is becoming more interesting
A particularly interesting recent development is Kinesis Data Streams streaming tables.
AWS announced on August 28, 2026 that Kinesis Data Streams can continuously deliver streaming data into Apache Iceberg tables in Amazon S3 Tables through a fully serverless capability. AWS says this can remove the need for self-managed Iceberg delivery pipelines and includes intelligent inline compaction.
This creates a powerful pattern for near-real-time analytics:
Applications / Events
↓
Amazon Kinesis Data Streams
↓
S3 Tables
Apache Iceberg
↓
Athena / EMR / Redshift
↓
Real-Time Analytics
↓
Amazon Kinesis Data Streams
↓
S3 Tables
Apache Iceberg
↓
Athena / EMR / Redshift
↓
Real-Time Analytics
For Data Engineers, this means streaming data can move closer to the lakehouse without requiring a large collection of custom delivery and maintenance components.
S3 Tables and AI-ready analytics
Modern data platforms are also being designed for AI workloads.
AWS has enabled Amazon Quick to use S3 table buckets directly as a data source, allowing users to work with Iceberg data for dashboards, conversational analytics, and agentic AI workloads without requiring an intermediate warehouse or OLAP layer in that architecture.
This is important because AI applications need reliable and governed access to current data.
A modern architecture can therefore look like:
Operational Data
↓
Streaming / Zero-ETL
↓
Amazon S3 Tables
↓
Apache Iceberg
↓
Governed Data Access
↓
Analytics + AI Applications
↓
Streaming / Zero-ETL
↓
Amazon S3 Tables
↓
Apache Iceberg
↓
Governed Data Access
↓
Analytics + AI Applications
Where AWS Glue fits
AWS Glue remains important because it can provide the ETL and catalog capabilities around the lakehouse.
For example:
# Conceptual Spark example
df = spark.read \
.format("iceberg") \
.load("catalog.database.customer_data")
.format("iceberg") \
.load("catalog.database.customer_data")
df.createOrReplaceTempView("customers")
result = spark.sql("""
SELECT customer_id, COUNT(*) AS orders
FROM customers
GROUP BY customer_id
""")
SELECT customer_id, COUNT(*) AS orders
FROM customers
GROUP BY customer_id
""")
The exact configuration depends on the AWS Glue, catalog, permissions, and table setup being used. AWS provides multiple supported ways to connect Glue and Iceberg clients to S3 Tables.
What should beginners learn?
For someone starting Data Engineering on AWS, I would recommend learning these concepts together:
1. Amazon S3
Understand object storage, data lakes, prefixes, lifecycle policies, and partitioning.
Understand object storage, data lakes, prefixes, lifecycle policies, and partitioning.
2. Apache Iceberg
Learn tables, snapshots, schema evolution, partitioning, and time travel.
Learn tables, snapshots, schema evolution, partitioning, and time travel.
3. AWS Glue
Learn ETL jobs, Spark, Data Catalog, crawlers, and production pipeline design.
Learn ETL jobs, Spark, Data Catalog, crawlers, and production pipeline design.
4. Amazon Athena
Learn SQL analytics directly on lake data.
Learn SQL analytics directly on lake data.
5. Amazon Kinesis
Learn streaming ingestion and near-real-time pipelines.
Learn streaming ingestion and near-real-time pipelines.
6. Governance
Understand IAM, AWS Lake Formation, permissions, and data discovery.
Understand IAM, AWS Lake Formation, permissions, and data discovery.
A portfolio project idea
A great beginner-to-intermediate project would be:
Real-Time E-Commerce Lakehouse on AWS
E-Commerce Events
↓
Kinesis Data Streams
↓
S3 Tables / Apache Iceberg
↓
AWS Glue
↓
Athena
↓
QuickSight / Analytics
↓
Kinesis Data Streams
↓
S3 Tables / Apache Iceberg
↓
AWS Glue
↓
Athena
↓
QuickSight / Analytics
You could generate events such as:
- Customer orders
- Product views
- Payments
- Cart activity
- Product inventory
- Product views
- Payments
- Cart activity
- Product inventory
Then build analytics such as:
- Daily revenue
- Top products
- Customer activity
- Conversion trends
- Real-time order monitoring
- Top products
- Customer activity
- Conversion trends
- Real-time order monitoring
This project demonstrates multiple real-world Data Engineering skills instead of showing only isolated AWS services.
Final thoughts
The important lesson is that modern AWS Data Engineering is moving toward managed, open-format lakehouse architectures.
Amazon S3 Tables brings managed Apache Iceberg table storage to S3, AWS Glue provides ETL and catalog integration, Kinesis can deliver streaming data into S3 Tables, and analytics services can work directly with the resulting data.
For students and aspiring Data Engineers, this is a useful architecture to understand because it combines storage, ETL, streaming, open table formats, governance, analytics, and AI into one modern data platform.
The future Data Engineer is not only moving data from one system to another—they are building reliable platforms that make data continuously available for analytics and intelligent applications.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article