Agents Meet the Lakehouse: 5 AWS Trends Data Engineers Should Watch in Late 2026
AWS is pushing agents into production while blurring the line between databases, object storage and the data lake. Here is what the latest Summit and September launches mean for anyone building data pipelines.
I'm finishing an MSc in Data Science and AI and working towards a data engineering role, so I've been reading AWS announcements with one question in mind: what changes for the people who build and run data pipelines? Looking across the AWS Summit New York 2026 launches and the steady stream of September updates, five trends stand out.
1. Agents are moving from demos to production
The biggest theme of 2026 is that AWS wants agents running in production, not just in notebooks. At the New York Summit, Amazon Bedrock AgentCore picked up generally available capabilities for connecting agents to organisational and web knowledge, monitoring them in production, and applying governance at scale. AgentCore Harness lets you define a production agent through configuration rather than writing orchestration code, and a managed Web Search tool gives agents current web knowledge.
The open-source Strands Agents SDK also gained better context management, an isolated execution environment (Strands Shell) and Strands Evals for chaos testing and red teaming.
Why it matters for data engineers: an agent is only as good as the data it can reach. Somebody has to own the connectors, freshness, access controls and lineage behind those agents, and that work looks a lot like data engineering.
2. Managed RAG is becoming a platform feature
Amazon Bedrock Managed Knowledge Base (GA) bundles native data connectors, smart parsing and an agentic retriever. AWS has also previewed AWS Context, which is described as mapping data relationships into knowledge graphs so agents can work with governed information.
On the storage side, Amazon S3 Vectors added metadata pre-filtering, which returns more relevant matches for RAG and semantic search. The pattern is clear: vector search, chunking and retrieval are moving out of custom code and into managed services. The skill that stays valuable is designing good metadata, document structure and evaluation, because the managed pieces still depend on clean inputs.
3. Object storage is getting smarter
S3 keeps doing more than storing bytes:
- S3 Annotations (GA) lets you attach up to 1 GB of queryable context to an object, which is useful for tagging documents with business context for search and agents.
- S3 Tables now supports more of the Apache Iceberg V3 specification, including geometry and geography types, nanosecond timestamps and column defaults.
Together, these make S3 feel less like a dumb landing zone and more like the centre of the analytics platform.
4. The database and the lakehouse are merging
One of the more interesting September updates lets Aurora PostgreSQL query Apache Iceberg and Parquet data directly, without building an ETL pipeline first. Aurora Serverless also improved its scaling, so it can add capacity much faster to handle bursty AI workloads.
For years the standard answer was "copy operational data into the warehouse, then join it." Direct lake queries from the operational database won't replace every pipeline, but they change the decision: sometimes the right move is no new pipeline at all. Knowing when to federate and when to materialise is becoming a core data engineering judgement call.
5. AI is coming for the developer workflow too
Beyond data, AWS is applying agents to software delivery. The AWS DevOps Agent (preview) reviews release readiness and tests changes; AWS Transform (preview) continuously checks code against baselines and opens remediation pull requests; and the AWS Security Agent in AWS Continuum does threat modelling and PR scanning. Kiro, AWS's spec-driven AI IDE, even has an iOS app in gated preview for managing sessions remotely.
For data teams, this means pipeline code, infrastructure-as-code and data quality checks will increasingly be reviewed and patched with AI assistance. Clear specs and good tests become more important, not less.
What I'm doing about it
- Learning Iceberg properly because it now sits underneath S3 Tables, Athena, Glue and even Aurora queries.
- Building one small RAG pipeline end to end with S3, a vector index and Bedrock, paying attention to metadata and evaluation rather than just getting an answer out.
- Treating governance as a feature, with access control, lineage and data contracts, because agents will only be trusted with data that is well governed.
Final thought
The common thread is that AI on AWS is becoming a data problem. Agents, RAG and AI coding tools all depend on well-modelled, well-governed and fresh data. That's good news for data engineers: the job isn't disappearing, it's becoming the foundation everything else stands on.
Which of these trends are you already seeing in your work? I'd love to hear in the comments.
Sources: AWS News Blog, "Top announcements of the AWS Summit in New York 2026", and AWS What's New updates from September 2026. Check each service's page for current availability in your Region.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article