Practical Big Data Analytics
Everyone wants to say they work with big data. The freeing truth is that almost nobody does - and that means your laptop is more powerful than you think.

TL;DR
Most teams don't have big data, and that's good news. Before reaching for clusters, apply an honest test: does the data fit on one machine? Columnar formats like Parquet and engines like DuckDB handle surprisingly large datasets on a laptop. Climb the tooling ladder step by step instead of leaping to distributed systems.
On this page
There is a quiet status game in analytics. Saying you work with “big data” sounds serious. It implies clusters, terabytes, and engineering muscle. So teams reach for the vocabulary, and then for the tools, long before the data demands either. I want to make an unfashionable argument: you probably do not have big data. And that is genuinely good news, because it means most of your problems have simpler, cheaper solutions than you have been told.
What “big” actually means
The classic way to describe big data is through a set of Vs. It started with three - volume, velocity, and variety - and later picked up veracity and value. Volume is sheer size. Velocity is how fast new data arrives. Variety is the mix of tables, logs, text, and images you have to juggle. Veracity is how much you can trust it, and value is whether any of it helps you decide anything.
Notice that only one of those, volume, is about being large. And even volume has a specific meaning here. Data earns the “big” label when it no longer fits or processes comfortably on a single machine. Not when it is annoying in a spreadsheet. Not when it has a few million rows. When it overflows one computer.
By that definition, the overwhelming majority of datasets are not big at all. They are medium at most.
The honest test
Ask yourself three questions about your dataset.
Does it fit in memory, or at least on the disk of one modern machine, today? Will it still fit in roughly a year? And is a typical query taking minutes rather than hours?
If the answers are yes, yes, and yes, you do not need a distributed system. You need better tools on the single machine you already own. A laptop with sixteen gigabytes of memory can chew through tens of millions of rows without complaint. That covers a startling share of real analytical work - sales records, event logs, survey results, product usage. The data feels intimidating because the CSV is slow to open, not because it is actually too large to handle.
Why CSV is fooling you
Here is the trap many people fall into. They open a multi-gigabyte CSV in pandas, watch it crawl and devour memory, and conclude that the data has outgrown their tools. They start shopping for a cluster.
But the slowness is usually the file format, not the size. A CSV is a row-oriented text file. Every value is stored as text, with no type information, and to read a single column you must scan the entire file. It compresses poorly and parses slowly.
Switch to a columnar format like Parquet and the picture changes completely. Parquet stores each column together, with proper types and built-in compression. A multi-gigabyte CSV often shrinks to a fraction of its size. Better still, a query can read only the columns it needs and skip everything else, and it can skip whole blocks of data that cannot possibly match your filter. That last trick is called predicate pushdown, and it means the engine avoids reading data it does not need at all.
Pair Parquet with an in-process query engine such as DuckDB, which runs SQL directly over Parquet files with no server to manage, and a great deal of so-called big data work collapses into something that finishes in seconds on a laptop. No cluster. No new infrastructure. Just a smarter file and a smarter reader.
The cost of pretending
Reaching for distributed compute when you do not need it is not free. A Spark cluster has to be provisioned, configured, secured, and paid for. Jobs that would be a one-line groupby on your machine become exercises in tuning partitions and chasing network shuffles. Debugging gets harder because the work is spread across machines you cannot easily inspect. The learning curve is real, and so is the operational drag.
All of that effort can be worth it when the data is truly enormous. When you are processing many terabytes spread across many machines, distributed frameworks are the right answer, and they are genuinely impressive. But spending that cost to process forty million rows that would fit on a thumb drive is the analytics equivalent of renting a freight train to carry your groceries.
The ladder, not the leap
The healthier mental model is a ladder rather than a leap. The bottom rungs are a spreadsheet for quick sharing and pandas for flexible work in Python. Above them sits columnar storage with Parquet and a tool like DuckDB, which still runs happily on one machine. Only above that do you find distributed engines like Spark, and finally a managed data warehouse for many analysts and governed, historical tables.
The skill is not knowing how to use the top of the ladder. The skill is knowing which rung your problem actually requires, and refusing to climb higher until the current rung genuinely hurts. Most of the time, you will be surprised how far the lower rungs carry you.
So the next time someone asks whether you work with big data, you can answer honestly. Probably not - and because of that, your work is faster, cheaper, and simpler than the hype would have you believe. That is not a limitation to apologize for. It is an advantage to enjoy.
Key takeaways 5
- Most organizations' data isn't truly "big".
- If it fits on one machine, use single-machine tools.
- CSV wastes space and time; columnar formats like Parquet are far better.
- Premature distributed systems add cost and complexity.
- Scale tools up a step at a time as data actually grows.
Watch & learn
Frequently asked questions
What counts as big data?
Data is "big" when its volume, velocity or variety exceeds what a single machine and conventional tools can handle. Many datasets called big data fit comfortably on a modern laptop or server.
Why use Parquet instead of CSV?
Parquet is a compressed, columnar format: files are much smaller, queries read only the columns they need and data types are preserved, making analysis faster and cheaper.
What is DuckDB?
DuckDB is an in-process analytical database that runs SQL queries very fast on local files such as CSV and Parquet, without needing a server or cluster.
Go deeper with the free masterclass
Workshop, PDF handbook and curated resources for “Practical Big Data Analytics”.
Related articles

Data Cleaning & Wrangling with pandas
Analysts love to talk about models and dashboards, but the real job is mostly cleaning. Here is why that unglamorous work deserves more respect - and a sharper method.

Data Analysis: From Spreadsheets to Python
You already think in rows, columns, and pivot tables. Here is why those same instincts make pandas the natural next step, and when it is worth the switch.

Data Warehousing & Analytics Engineering
The warehouse stopped being a place to store data and became a place to write code - and that single shift created a whole new craft.

Comments
No comments yet. Start the conversation.