Observability: Logs, Metrics & Traces
Monitoring tells you the building is on fire; observability tells you which wire shorted, in which room, for which tenant - without you having ever predicted that wire.
The email always sounds the same. "Checkout is slow, but only for customers in Germany, paying with one particular card type, since roughly this morning." You read it twice, open your dashboards, and find nothing. CPU is calm. Memory is calm. The uptime check is green. Every graph you built is reporting health while a real customer is staring at a spinner.
This is the moment that separates monitoring from observability, and once you have lived it, the distinction stops being academic.
The question you did not think to ask
Monitoring is the discipline of known-unknowns. You sit down ahead of time, enumerate the ways your system can fail - disk fills, queue backs up, latency creeps - and you build a graph and an alert for each. It is a genuinely good practice, and it works beautifully right up until your system stops being one application on one box.
The German-card-type bug is a different animal. Nobody built a dashboard for it because nobody imagined it. Observability is the property that lets you ask that new question of your running system without shipping new code to answer it. You slice your telemetry along a dimension you never anticipated - country, then card type, then deploy version - and the answer is already in there, waiting.
The reason this got hard is structural. A monolith has a small, countable set of failure modes. A modern request touches a dozen services, two managed databases, a queue, and a third-party API you do not control. The request is not slow "in the app." It is slow somewhere in a chain no single machine can see, and the container you would have SSHed into has already been rescheduled.
Three signals, one identifier
The working answer is the three pillars, and the trick is that they are not three competing tools. They are three resolutions of the same picture.
Metrics are your cheap, always-on trend line. A counter costs roughly the same whether one request or a million hit it, which is why you build dashboards and alerts on metrics and nothing else. They tell you that the error rate jumped at 14:05. They will never tell you why.
Traces are the map. A trace is one request's journey rendered as a waterfall of nested spans, and in a single glance it tells you that of an 820 ms checkout, 540 ms was the payment charge and almost all of that was the external provider. No amount of staring at per-service dashboards reconstructs that decomposition. The trace localises the pain to one hop.
Logs are the close-up. Once the trace has pointed you at the failing span, the log line for that exact request carries the precise exception, the order ID, the amount, the stack. Logs are the most familiar pillar and the most abused, because the instinct is to write them for human eyes - prose you can only grep with fragile regex - when you should be emitting structured events where every field is a queryable dimension.
What stitches the three together is one identifier. The trace ID is generated at the edge, propagated to every downstream call via the traceparent header, and stamped onto every log line. That single seam is what lets you alert on a metric, jump to the trace, and read the logs for that request in three clicks instead of three hours.
Cardinality is the whole game
Here is the uncomfortable truth underneath all of it. The questions that actually matter in production are high-cardinality: which user, which request, which deploy, which card type. And high cardinality is exactly what traditional metrics punish, because every unique combination of labels is a separate time series to store. Put user_id in a metric label and you can multiply your series count by millions overnight.
So you carry the high-cardinality detail in logs and traces, you keep your metric labels disciplined and low-cardinality, and you accept that this tension - you need cardinality to debug but it is expensive to store as metrics - drives most of the architecture decisions you will make. The frontier idea, wide events, is just this tension pushed to its conclusion: one very wide, high-cardinality record per unit of work, from which you derive metrics by aggregating and traces by linking.
Alert on pain, not on twitches
None of this matters if your pager cries wolf. The most common self-inflicted wound in this whole field is alerting on causes - CPU over 80 percent, a pod restarted, disk at 90 - most of which require no human action. Engineers learn to ignore the pager, and then they miss the one that mattered.
The discipline is to alert on symptoms. Define an SLO - 99.9 percent of requests under 300 ms over 28 days - and page on the rate at which you are burning the error budget that target implies. Fast burn, where a month of budget evaporates in hours, pages a human now. Slow burn files a ticket. Everything else - the hot CPU, the restart - lives on the dashboard you consult after you are paged, to find the cause.
That is the quiet promise of observability done well. Not more graphs. Fewer pages, faster diagnosis, and the ability to answer the question in the customer email you never thought to ask.
No comments:
Post a Comment