Tech Insights

Intro to Statistics for Analysts

You already work with data every day. Here are the five statistical ideas that will stop the numbers from quietly misleading you.

Statistical charts used by a data analyst

TL;DR

Analysts don't need a semester of theory to avoid being misled; they need five ideas used consistently. Describe data honestly (medians for skewed data), distrust the sample before the result, treat correlation as a clue rather than proof, read p-values narrowly and lead with confidence intervals and practical importance.

On this page

Plenty of capable analysts have never sat through a statistics course. They write SQL fluently, build dashboards, and ship reports — and then, at the moment of interpretation, fall back on instinct. Instinct is exactly where statistics is most needed, because data has a talent for telling flattering stories that are not true. You do not need a semester of theory to defend yourself. You need five ideas, applied consistently.

Start by describing, honestly

The first temptation is to average everything. But the average is the most over-trusted number in analytics. The moment your data has a long tail — revenue per customer, time on page, deal size, response latency — the mean drifts toward the extremes and ends up describing nobody.

A team reports that the average account is worth 159,000 dollars. Impressive, until you notice one whale and six modest accounts. The median tells the honest story: a typical account is worth 38,000. The fast check costs nothing: compare the mean and the median. If they disagree, the data is skewed and the median is your friend. And pair every measure of center with a measure of spread — a standard deviation or an interquartile range — because “the average is 50” hides whether everything clusters at 50 or scatters from 0 to 100. Above all, plot the shape. A histogram takes seconds and reveals the outliers and hidden subgroups that summary numbers bury.

Distrust the sample before you trust the result

Almost every dataset you touch is a sample standing in for a larger population you actually care about. The arithmetic only means something if the sample represents that population. It usually does not, and the failure is rarely random.

The classic trap is the satisfaction survey: eight percent respond, they average 4.6 out of 5, and someone writes “users rate us 4.6.” But the ninety-two percent who ignored the survey are not a random ninety-two percent — they are the less engaged, the indifferent, the quietly unhappy. Your 4.6 describes respondents, not users. The same blind spot hides churned accounts when you study only active ones, and abandoned carts when you study only completed checkouts.

Here is the part that surprises people: more data does not fix this. A bigger biased sample is just a more confident wrong answer. Bias is systematic, and volume does not wash it out. The only defense is to ask, every single time, who or what is missing from this data and why.

Correlation is a clue, not a verdict

When two things move together, the mind leaps to “one causes the other.” Resist it. A correlation has at least four explanations: X causes Y, Y causes X, some third factor drives both, or it is pure coincidence. The third one — the confounder — quietly ruins most business analyses.

Suppose customers who adopt a new feature retain far longer. Ship the feature to everyone? Not so fast. Your most engaged users were always going to adopt features and always going to stick around. Engagement causes both behaviors; the feature may contribute nothing. The correlation is real, the causal story is invented. The only clean way to know is a randomized experiment: give the feature to a random half and compare. Randomization is what breaks the link to confounders, which is why the A/B test, not the correlation, earns the right to say “caused.”

A p-value is narrower than you think

Hypothesis testing answers one question: could this pattern be random noise? You assume there is no effect (the null hypothesis), then ask how surprising your data would be if that were true. The p-value measures that surprise — it is the probability of seeing data at least this extreme assuming nothing is really going on.

Notice what it is not. A p-value of 0.03 does not mean a 97 percent chance the effect is real. It does not mean the result is important. And the famous 0.05 cutoff is a convention, not a law — 0.049 and 0.051 are essentially the same evidence. The real dangers are behavioral: testing twenty metrics and reporting the one that crossed the line, or refreshing an A/B test until it briefly looks significant and stopping there. Both manufacture false positives. Decide your metric and your sample size before you look.

Lead with the interval, then ask if it matters

The fix for the p-value’s narrowness is the confidence interval. Instead of “the lift was significant,” report “the lift was 2.1 percent, with a 95 percent interval from 0.3 to 3.9 percent.” Now you can see the size of the effect and the size of your uncertainty in one line. If that interval excludes zero, the result is significant; if it includes zero, it is not — and you have learned far more than a bare verdict could tell you.

Then ask the question that statistics alone cannot answer: is the effect big enough to act on? A 0.05 percent lift can be statistically significant with enough traffic and still be too small to bother shipping. A 4 percent lift with an interval spanning minus 1 to plus 9 percent is not significant yet, but the upside is worth gathering more data. Decide in advance how large an effect would change your decision, and judge results against that threshold, not against zero.

None of this is advanced. Describe honestly, distrust your sample, treat correlation as a clue, read the p-value for exactly what it says, and lead with the interval and the practical threshold. Do these five things every time and you will avoid the mistakes that sink most analyses — and your numbers will finally start telling the truth.

Key takeaways 5

  1. Averages mislead on skewed data; look at medians and distributions.
  2. A biased sample can't be fixed by more analysis.
  3. Correlation is a clue, not proof of causation.
  4. A p-value says less than most people think.
  5. Report confidence intervals and ask whether the effect matters in practice.

Watch & learn

Introduction to Statistics and Data AnalysisSteve Brunton · YouTube

Frequently asked questions

When should I use the median instead of the mean?

Use the median when data is skewed or has outliers, such as revenue per customer, response times or deal sizes, because a few extreme values can pull the mean far from a typical case.

What does a p-value actually mean?

A p-value is the probability of seeing results at least as extreme as yours if there were really no effect. It does not tell you the probability that your hypothesis is true or how large the effect is.

Why doesn't correlation imply causation?

Two variables can move together because of a third factor, reverse causation or chance. Experiments or careful causal methods are needed to show one causes the other.

Tech InsightsScience VaultProjects & Practice#statistics#data-analysis#hypothesis-testing#p-values#confidence-intervals

Comments

No comments yet. Start the conversation.

Comments are reviewed before they appear. Be kind; one link max.

Go deeper with the free masterclass

Workshop, PDF handbook and curated resources for “Intro to Statistics for Analysts”.

Open AL Academy ↗
Keep reading

Related articles