senior loop
Deep Dives/Observability Basics

Observability Basics

A system you cannot see inside is a system you cannot operate. Metrics, traces, and logs answer three different questions, and averages will lie to you about all of them.

FundamentalsOperations~10 min · 5 sections

Prerequisites: A service in production, or the intention of having one.

After this: Say how you would know the system is broken, in percentiles, before the interviewer asks.

Suggested first pass: Read sections 1–5, answer each section in your own words, then use the remaining failure modes and exercises as the advanced pass.

AnswersCostFails at
MetricsIs something wrong?cheapWhy. Detail is aggregated away
TracesWhere did the time go?moderate, sampledPer-request detail
LogsWhat happened to this request?expensive at volumeFinding the right lines

Metrics are numbers over time: request rate, error rate, latency, queue depth. Easy to graph and alert on. Logs are individual events with detail. Traces follow one request across every service it touches, with timing per hop.

The workflow runs in that order. Metrics alert you, traces localise the problem, logs explain it. A design that mentions all three in that order sounds like someone who has been on call.

Next deep dive
Idempotency & Exactly-Once Effects in Payments
~35 min