Data engineering with AI agents: a playbook
How I'd let an agent build pipelines without trusting it too much.
There's a lot of noise right now about AI agents writing data pipelines. Most of what I see is a demo: someone hands a language model database access, it generates some Spark, and it works on the happy path. It's a nice demo, but I wouldn't run it in production.
This post is about the opposite approach: don't try to make the agent smarter, make its job smaller. A fixed architecture, a set of tested building blocks, a contract that says exactly what to build, and people checking the steps that need judgment. The agent only works inside those rules.
I'll go through the pieces one by one, then follow a single table, orders, from the first request to production, so you can see how they fit together.
What gets decided up front
Before any agent runs, people decide the parts that must not vary:
- The global architecture. A layered lakehouse: raw / bronze / silver / gold, or landing / bronze / curated. The names don't matter much, the discipline does. Each layer has a fixed contract about what it holds and how much you can trust it.
- The abstraction layer. A library of tested classes the agent composes, instead of raw Spark it makes up.
- The deployment path. CI and infrastructure code. People write it once and it works the same way for every pipeline, so I won't spend much time on it.
The agent is what makes this worth doing. It can go from a contract to a working pipeline in minutes, and it's just as fast on the fiftieth table as on the first. But it's only safe to use because of the structure around it: the tested building blocks and the rules above. That structure is what most of this post is about.
The global architecture: raw data as a safety net
The raw zone is designed once by a person, and agents never change it. It follows two rules:
- Append-only. Nothing in raw is ever updated or deleted, so it's always an exact copy of what the source sent.
- Partitioned by arrival date, not by when the event happened:
raw/source=salesforce/dataset=orders/ingest_date=2026-01-15/
For this post, the reason these matter is simple: they make mistakes cheap. If a pipeline the agent built turns out to be wrong, nothing is lost. You fix it and replay the affected days from raw, and since each day is its own partition, a re-run replaces that partition instead of rewriting a whole table. That's what makes it reasonable to let an agent build on top of it. The worst case is a re-run, not lost data.
The abstraction layer
This is the part that actually makes agents safe. We don't let agents write Spark, with each one reinventing what a bronze-to-silver step looks like. Those steps are nearly identical across pipelines anyway: read a source, standardize it, apply changes in order, enforce quality, write a table. So we write that logic once, as tested classes, and register them in a catalog the agent draws from.
RawToBronze # land + parse + capture rescued columns
Standardize # cast types, rename, normalize formats
MaskColumns # hash or redact columns marked as personal data
ApplyCDC # apply an ordered change feed, keep latest (SCD type 1)
ApplySCD2 # apply an ordered change feed, keep history
QualityGate # nulls, ranges, uniqueness, referential
Think about the level an agent could generate at:
- Raw Spark. The agent can produce code that's subtly wrong, hard to maintain, and untested. Correctness lives in the agent's output, which isn't deterministic. Unsafe.
- Calls to tested abstractions. Correctness lives in your tested class. The agent picks the class and fills in the parameters. The abstraction is the test boundary.
So the agent's job shrinks from "write a correct pipeline" to "wire known-good blocks in the right order with the right parameters". It ends up producing config, not logic. That's really the main point of this post.
The silver cleansing you always end up needing (fixing nulls, bad formats, slowly changing dimension errors) is just another abstraction here, QualityGate or Standardize, configured by the contract and never improvised by the agent.
The contract
The agent needs a bounded spec of what to build. That's the interface contract: a document a person writes, or at least reviews, describing the source, the target, and the rules. Everything else is generated from it, so if the contract is right, most of what comes after is right too.
source:
location: s3://raw/salesforce/orders/
format: json
target:
table: catalog.silver.orders
primary_keys: [order_id]
sequence_by: updated_at # orders change events
scd: type_2 # keep history
cdc: true # source is a change feed
schedule: "@hourly" # also used to generate the DAG
privacy:
gdpr: personal_data # this table holds personal data
schema:
- {name: order_id, type: string, nullable: false}
- {name: customer_email, type: string, pii: true, mask: keyed_hash}
- {name: amount, type: decimal(18,2)}
- {name: status, type: string, allowed: [new, paid, cancelled]}
- {name: updated_at, type: timestamp}
quality:
- not_null: [order_id, amount]
- unique_current: [order_id] # one current row per order
Two terms in there drive the choice of abstractions:
- CDC (change data capture). The source emits row-level inserts, updates, and deletes instead of full snapshots. The
sequence_bycolumn orders those changes, so an older change that arrives late never overwrites a newer one. - SCD (slowly changing dimension). How you handle an attribute that changes. Type 1 overwrites and keeps only the current value. Type 2 keeps history: a new row per change, with validity dates and a current flag. That's why the quality rule checks uniqueness on current rows only. The contract says which type; the abstraction implements it.
Notice the contract uses decimal for money and is explicit about keys and the sequence column. If you leave those out, the agent will guess, and it'll often guess wrong.
The privacy block works the same way. The contract says the table holds personal data under GDPR and which columns get masked. The agent doesn't decide what's sensitive, and it can't forget a column either: validation rejects any plan where a pii column reaches silver without a MaskColumns step. Raw keeps the original value, locked down, and silver only sees a keyed hash. Hashing is pseudonymization, not anonymization, so the silver table still counts as personal data and keeps its classification.
How the agent finds the abstractions
The piece people tend to skip is how the agent knows which abstractions exist and when to use each one. My answer: expose every abstraction as a described, machine-readable capability, basically a typed tool definition, and require the agent to choose only from that catalog.
Each entry carries what the agent needs to choose well: a name, a one-line "use this when", typed parameters, and preconditions (ApplySCD2 requires a sequence_by column, for example). The agent gets the contract plus the catalog and is asked for a plan: which abstractions, in what order, with which parameters. Not code.
For the orders contract, the plan looks something like this:
pipeline: silver_orders
steps:
- use: RawToBronze
with: {source: s3://raw/salesforce/orders/, format: json}
- use: Standardize
with: {schema: contract.schema}
- use: MaskColumns
with: {columns: [customer_email], method: keyed_hash}
- use: QualityGate
with: {not_null: [order_id, amount]}
- use: ApplySCD2
with: {keys: [order_id], sequence_by: updated_at,
partition_by: date(updated_at),
target: catalog.silver.orders}
- use: QualityGate
with: {unique_current: [order_id]}
That's all the agent produces at this stage. No Spark, no SQL. Each step names a block from the catalog and fills in parameters taken from the contract. Reviewing it means reading a dozen lines against the contract, not auditing a pipeline.
Two properties make this safe:
- Closed vocabulary. The agent composes only from the catalog and can't invent a transform. If a contract needs something the catalog doesn't have, the agent escalates to a person instead of improvising. That's on purpose.
- Validation before generation. The plan is checked against each abstraction's preconditions before anything is generated, so a bad wiring is caught structurally rather than at runtime.
If you've built agentic systems, this is the familiar tool-use pattern: you aren't asking the model to be creative, you're asking it to pick correctly from a vetted menu. One thing that's easy to forget: the descriptions in the catalog need to be maintained like code. If a description is vague, the agent's choices will be vague too.
Anthropic's post on building effective agents has a useful distinction here. It separates workflows, where the model and its tools run along predefined code paths, from agents that decide their own process. What I'm describing is closer to a workflow, on purpose: the path is fixed and the model fills in the steps. The same post suggests putting as much effort into the agent-computer interface, meaning the tools and their docs, as you would into a user interface. For this setup, that interface is the catalog.
Where the humans stand
The agent suggests things; people and CI decide. In practice that means an agent (say a Claude agent with the abstraction catalog as its tools) that opens pull requests instead of committing to main, and a person checking the steps that need judgment. The walkthrough below shows exactly where those checks sit. Here's the split of work:
| Layer / artifact | Who does what |
|---|---|
| Raw zone design | Written by people, deterministic, fixed |
| Abstraction library | Written and tested by people |
| Contract | Written by a person, or drafted by the agent and approved by a person |
| Bronze to silver pipeline | Generated by the agent (config over tested blocks), reviewed by a person |
| Quality, SCD, CDC, privacy rules | Declared in the contract, implemented by abstractions |
| CI and deployment | Written by people once, same for every pipeline |
| DAGs / orchestration | Generated by the agent from the contract, unit-tested by a person |
| Gold layer | Written by people, with an assistant at most |
One table, start to finish
Here's the whole thing in order, using the orders table from above.
- The request. Someone on the finance side needs order history in silver, including how an order's status changed over time. At this point it's a business need, not a spec.
- The contract. The agent drafts the contract from that request:
order_idas the key,updated_atto order changes, SCD type 2 because they want history, decimal for amounts,customer_emailmarked as personal data, an hourly schedule. A person reads it and signs off, because business intent is the one thing the agent must never assume. This step is never skipped. - The plan. The agent reads the contract and the catalog and outputs the plan shown earlier. Still no code.
- Validation. The plan is checked automatically: every step is in the catalog,
ApplySCD2has thesequence_byit needs, every column it mentions exists in the contract, everypiicolumn is masked. A plan that fails here never reaches a person. - Plan review. A person checks the choices: SCD2 keyed on
order_id, ordered byupdated_at, partitioned by event date. That takes a few minutes. - Generation. The agent turns the plan into pipeline config plus an Airflow DAG and opens a PR. The config only points at tested classes; nothing is hand-written Spark.
- CI and deploy. CI runs the same way for every pipeline: tests against sample data, DAG unit tests (the kind Astronomer documents for Airflow), and a preview of the infrastructure change with a cost estimate. This is where you'd catch an agent that dropped a table or skipped a rule. Cluster size is capped by policy, so a generated pipeline can't quietly ask for a huge cluster. A person approves the PR and CI promotes it from dev to prod.
- Running. Every hour the pipeline picks up what landed in raw and updates the silver table. If a bug shows up in a block later, it gets fixed once and the affected days are replayed from raw.
The agent does three things here: drafts the contract (2), writes the plan (3), and generates the config (6). Those three are where most of the hands-on time used to go. Everything else is a person or code people wrote.
A failure that gets through
None of this is bulletproof. Here's an example of something that could still get through.
Picture a contract for a customers source that lists twelve columns. The source recently gained a thirteenth, loyalty_tier, and nobody updated the contract. The agent generates a perfectly valid pipeline: right abstractions, schema validation passes (the contract is internally consistent), tests pass (they only check the columns the contract declares). It runs green for three weeks. The problem surfaces when an analyst asks why loyalty segmentation is empty.
The agent didn't do anything wrong. It built exactly what the contract said. The gap is that nothing checked the contract against the real source schema.
The good news is that nothing was lost. RawToBronze captures columns the contract doesn't know about, so loyalty_tier sat in bronze the whole time. Nobody was looking at it. Once the contract is updated, the three weeks of silver can be backfilled from bronze.
What's missing is a guardrail: a schema drift check that compares the live source schema to the contract at ingestion and flags any difference for review, instead of waiting for an analyst to notice. The agent can only be as good as the contract, so you need something that checks the contract against the real data. A better model wouldn't fix this. A simple schema check would.
When not to use this
This isn't where you start. It pays off only where a pattern repeats and a tested abstraction for it already exists. So:
- Young projects. Early on you don't know your patterns yet, you're still finding them. Building a catalog before you understand the shape of your pipelines is premature. Write pipelines by hand until the patterns settle, then extract abstractions, then let agents compose them.
- One-off pipelines. If it won't be repeated, there's no abstraction, and building one for a single use isn't worth it.
- New logic with no safe abstraction. That's why the gold layer stays human-written, with a coding assistant at most. Gold is where business logic lives. Every mart is different, so there's no repeated pattern to constrain against, the cost of a subtle mistake is highest, and there's no tested abstraction to hold the agent to.
The general rule: agents pay off where the pattern repeats and a tested abstraction exists. Anything new, one-off, or not well understood yet, a person should write.
The takeaway
To sum up: this doesn't depend on a good prompt or a smarter model. It works because most of the hard decisions are made before the agent runs. The raw zone is set up so you can replay safely, the logic lives in tested building blocks, the contract says exactly what to build, the agent can only pick from a fixed list, and people review the steps that need judgment. So when the agent gets something wrong, it gets caught.
The tools will probably change a lot. I think the idea still holds: people set the rules and check the important steps, and the agent does the repetitive wiring.
Sources and further reading
Agentic design
- Anthropic, Building effective agents. Workflows vs agents, and why tool design matters as much as the prompt.
- Anthropic, Writing effective tools for agents. How to name, scope, and describe tools so a model picks the right one.
- Model Context Protocol. An open standard for exposing tools to models, one way to publish a catalog like the one above.
Change data capture
- Debezium documentation. Log-based CDC for Postgres, MySQL, and others.
- Delta Lake change data feed. CDC on the lakehouse side, between your own tables.
- Martin Kleppmann, Designing Data-Intensive Applications, chapter 11. The section on change data capture explains why the log is the source of truth.
Slowly changing dimensions and modeling
- Ralph Kimball and Margy Ross, The Data Warehouse Toolkit. Where the SCD types come from.
- Kimball Group, Dimensional modeling techniques. Short reference pages for SCD types 0 to 7.
- Databricks, Medallion architecture. The bronze / silver / gold layering.
Orchestration
- Apache Airflow, Best practices, including how to test DAGs.
- Astronomer, Testing Airflow DAGs.