ETL & Data Pipelines
Pipelines are not judged by how cleverly they move data, but by whether anyone can trust the numbers that come out the other end.

TL;DR
Data pipelines are judged not by clever transformations but by whether people can trust the numbers they produce. Modern pipelines load raw data first and transform it in the warehouse (ELT), rely on idempotent, re-runnable steps, fail loudly with tests and alerts, and favor boring, proven tooling.
On this page
The plumbing nobody sees
Every dashboard the executive team stares at, every churn model the data scientists tune, every revenue figure the finance team closes the quarter on - all of it sits on top of a data pipeline. The pipeline is plumbing. And like plumbing, you only think about it when it fails: when the dashboard is empty on Monday morning, when last quarter’s number quietly changes, when two reports that should agree do not.
This is the central tension of data engineering. The work that gets praised is the clever transformation, the slick model, the beautiful chart. But the work that actually matters is the unglamorous discipline of making the plumbing reliable. A pipeline that loads fast but loads wrong is worse than no pipeline at all, because people make decisions on it without knowing it lies.
How we stopped transforming first
For most of the field’s history, the pattern was ETL: extract data from sources, transform it on a dedicated server, then load only the clean result into the warehouse. This made sense when warehouses were expensive and slow. You rationed their compute and you did not waste storage on raw junk.
Then the ground shifted. Cloud warehouses - Snowflake, BigQuery, Redshift - separated storage from compute and made both cheap and elastic. Suddenly there was no reason to throw away raw data, and every reason to keep it. The pattern flipped to ELT: extract, load the raw data immediately, and transform it inside the warehouse with SQL.
The reordering of three letters sounds trivial. It was not. Keeping raw data means that when you discover a bug in a transformation - and you always do - you fix the SQL and re-run, rather than scrambling to re-extract data that may no longer exist at the source. It means a new question from the business can be answered by writing new SQL against data you already have, not by building a new extraction. ELT traded a little storage cost for an enormous amount of flexibility, and the trade was overwhelmingly worth it. Tools like dbt arrived to make those in-warehouse transformations into version-controlled, tested, documented code rather than clicks in a GUI.
The property that holds everything up
If I had to name the one idea that separates pipelines that work from pipelines that hurt, it would not be ETL or ELT. It would be idempotency: the property that running a step twice produces the same result as running it once.
This sounds like a pedantic detail until you realize that pipelines re-run constantly. A task crashes and Airflow retries it. A bug is found and you backfill three months of history. A source sends yesterday’s data twice. In every one of these cases, a non-idempotent pipeline turns a survivable event into corrupted data. The order count doubles. The revenue figure is wrong. Trust evaporates.
Idempotency is not hard to achieve - you upsert on a key instead of blindly inserting, or you overwrite whole partitions so re-running a day is safe. But it has to be designed in from the start, because retries and backfills are useless, even dangerous, without it. The reason to obsess over it is simple: it is what makes a pipeline safe to fix. And you will always need to fix it.
Fail loudly
The most dangerous failure in a data pipeline is not the one that throws an error. It is the one that succeeds while loading zero rows. The job goes green, the dashboard goes blank, and nobody notices for a week.
This is why mature teams build pipelines that fail loudly. They add data-quality tests that assert their assumptions: this key is unique, this column is never null, this table has roughly the number of rows we expect, every foreign key points at a real parent. They run those tests as part of the pipeline, and they let a failed test fail the run, so bad data never reaches the people downstream. They wire up alerting that fires on outcomes, not just crashes. A pipeline that screams when something is wrong is infinitely more valuable than one that fails in polite silence.
The corollary is that you must decide, deliberately, what a failure should do. Should a broken quality check halt the pipeline and serve yesterday’s known-good data, or should it load anyway and raise a flag? There is no universal answer. A financial table blocks. An experimental dashboard might warn. The discipline is in choosing on purpose rather than by accident.
The boring path is the right one
Put all of this together and a pattern emerges. The teams whose data everyone trusts are rarely the ones with the most sophisticated streaming architecture or the cleverest transformations. They are the ones who chose batch over streaming because batch was enough, who made every task idempotent, who tested their data like they test their code, who built observability and lineage from day one, and who kept an eye on cost so the pipeline did not quietly bankrupt them.
It is a boring path. It is also the right one. Data engineering rewards the unglamorous virtues - reliability, repeatability, honesty about failure - far more than it rewards cleverness. Build the plumbing so well that nobody ever has to think about it. That is the whole job.
Key takeaways 5
- Pipelines are plumbing: noticed only when they break.
- ELT loads raw data first and transforms it in the warehouse.
- Idempotency means a step can re-run safely without duplicating data.
- Fail loudly: test data and alert on problems.
- Boring, proven tools beat clever custom code.
Watch & learn
Frequently asked questions
What is the difference between ETL and ELT?
ETL transforms data before loading it into the warehouse. ELT loads raw data first and transforms it inside the warehouse, which keeps raw history and uses the warehouse's computing power.
What does idempotent mean in data pipelines?
An idempotent step produces the same result no matter how many times it runs, so failed jobs can be retried safely without creating duplicates.
How do I make data pipelines reliable?
Use idempotent loads, data quality tests, monitoring and alerting, clear ownership, version-controlled code and orchestration that handles retries and dependencies.
Go deeper with the free masterclass
Workshop, PDF handbook and curated resources for “ETL & Data Pipelines”.
Related articles

IIoT: Sensor-to-Cloud Data Pipelines
Everyone wants the dashboard and the digital twin. Almost nobody wants to talk about the pipes that feed them - which is exactly why so many Industry 4.0 projects stall.

Data Warehousing & Analytics Engineering
The warehouse stopped being a place to store data and became a place to write code - and that single shift created a whole new craft.

Practical Big Data Analytics
Everyone wants to say they work with big data. The freeing truth is that almost nobody does - and that means your laptop is more powerful than you think.

Comments
No comments yet. Start the conversation.