About this article
As the ninth installment of the “DevOps Architecture” category in the series “Architecture Crash Course for the Generative-AI Era,” this article explains monitoring and observability.
What answers “is it running?” is monitoring, what answers “why did it break?” is observability. This article covers the 3 pillars of monitoring (Metrics/Logs/Traces), the 4 golden signals (Latency/Traffic/Errors/Saturation), OpenTelemetry, AIOps, and operational implementation (system-architecture-stage monitoring requirements live in the separate “System Architecture” article).
Before you read this
This article is mostly about the flow of building, testing, releasing and monitoring a service. If IT vocabulary is unfamiliar, reading the primer "From Development to Operations" first makes it far easier to follow. You can also look anything up in the glossary as you read.
What are monitoring and observability, anyway?
Imagine driving a car. Speedometer, fuel gauge, engine warning light — without dashboard instruments, you can’t tell your speed, remaining fuel, or engine status. “Noticing only after the engine seizes” is too late.
Monitoring is your system’s dashboard. It’s the mechanism of continuously measuring values like CPU usage, error rate, and response time, and firing alerts when something is off. Observability goes one step further, referring to the ability to investigate “why it broke” after the fact.
Without monitoring, you first learn about incidents from user complaints. Without observability, even when you notice an incident, you can’t identify the cause, and recovery takes hours.
Why it’s needed
First, because damage expands while nobody notices. The gap between finding out from your monitoring and finding out from a customer is the gap between minutes and hours. Second, because tracking a cause is hard in microservices. With dozens of services on a request path, no single log tells you where the time went. Third, because an SLO or SLA has to be expressed in numbers, and numbers require measurement.
The three pillars — metrics, logs and traces
The 3 fundamental data types of observability. Any one alone is insufficient, and the whole picture appears only when combining all 3. “3 pillars” is now classic, with the modern mainstream being multi-faceted approaches adding Events and Profiles.
| Type | Content | Representatives |
|---|---|---|
| Metrics | Numeric time-series data | Prometheus, Datadog |
| Logs | String event records | Loki, Elasticsearch |
| Traces | Request paths | Jaeger, Tempo |
Metrics show “what’s happening,” logs show “what was recorded,” traces show “how it moved.”
Metrics are the numeric view.
Numerically-quantified time-series data, recording CPU usage, request count, error rate, latency, etc. Storage-efficient and good for aggregation/alerting - the basis of monitoring.
| Typical metrics | Content |
|---|---|
| System | CPU, memory, disk I/O, network |
| App | Request count, error rate, latency |
| Business | Order count, signup count, revenue |
| USE | Utilization, Saturation, Errors |
| RED | Rate, Errors, Duration |
USE/RED are famous metric-design frameworks used as guides for “what to measure.”
Logs are the textual record of events.
Time-stamped text events recording detailed info output by apps. Structured logs (JSON format) are the modern standard, linkable with metrics and traces. Logs are covered in detail in the next article.
| Log type | Use case |
|---|---|
| App logs | Business-processing records |
| Access logs | HTTP requests |
| Audit logs | Permission ops, important changes |
| System logs | OS / middleware events |
Logs occur in massive amounts, so storage cost easily becomes a problem - retention period, compression, and sampling strategies are operational topics.
Traces follow a request across services.
Data tracking the process of one request transiting multiple services. In microservices, with chains like “the order API calls the inventory API and the payment API,” you can visualize where it’s slow and where it failed.
Traces have Spans (individual processing) connected by Trace ID, expressing the whole flow as a DAG. Representative tech is OpenTelemetry (industry standard), sent to Jaeger, Tempo, Datadog APM, etc.
Trace: order-abc123
|- Span: POST /order (100ms)
| |- Span: check stock (30ms)
| |- Span: charge card (50ms)
| |- Span: Stripe API (45ms)
The main tools — open source and SaaS
OSS-based observability has Prometheus + Grafana + Loki + Tempo + OpenTelemetry as the de facto standard combination. Developing under CNCF, with high future prospects.
| Tool | Role |
|---|---|
| Prometheus | Metric collection, storage, alerting |
| Grafana | Visualization dashboards |
| Loki | Log aggregation (Grafana Labs) |
| Tempo | Distributed tracing (Grafana Labs) |
| OpenTelemetry | Standard for measurement-data collection |
| Alertmanager | Alert notifications |
Integrated as the LGTM stack (Loki, Grafana, Tempo, Mimir), rapidly spreading recently.
On the SaaS side the options are as follows.
If you don’t want to build the foundation in-house, using observability SaaS is a quick choice. Pricing is high, but you get high-feature monitoring foundation with near-zero operational burden.
| Service | Characteristics |
|---|---|
| Datadog | Strongest features, expensive |
| New Relic | Veteran, all-feature integrated |
| Splunk | King of log analytics |
| Dynatrace | AI-driven auto-analysis |
| Honeycomb | Strong on high cardinality |
| Grafana Cloud | Managed LGTM stack |
Datadog has the strongest features, but often hits hundreds of thousands to millions of yen monthly - small scales realistically use Grafana Cloud or New Relic free tier.
OpenTelemetry (measurement standard)
The industry standard unifying collection of metrics, logs, and traces. Born in 2019 from the merger of OpenTracing and OpenCensus, advancing from CNCF Incubating to Graduated - the modern common standard.
Using OpenTelemetry, instrumentation code can be written vendor-neutrally. Even when switching from Datadog to Grafana, no need to rewrite app-side code. Just swap send destination on the Collector side, so vendor-lock-in avoidance kicks in not as desk argument but at implementation level.
Furthermore, with unified instrumentation API, using multiple tools together becomes easy. For example: production metrics to Datadog, long-term log archive to S3 + Loki, dev-env traces to Jaeger - distribute by use case from the same instrumentation code. Custom SDKs require redoing instrumentation per tool, so this difference becomes non-negligible at scale.
| Feature | Content |
|---|---|
| Vendor-neutral | Send to any backend |
| Language support | Almost all major languages |
| Auto-instrumentation | Major-framework-supporting |
| Unified measurement | Metrics + logs + traces integrated |
For new builds, OpenTelemetry is the top candidate. Avoids vendor lock-in.
What to alert on is the decision that matters. — fire on user impact
Just collecting monitoring data is meaningless - it becomes operationally useful only with visualization and alerting. The standard operation is overviewing state with Grafana or each SaaS dashboard, and notifying via Slack or PagerDuty when thresholds are exceeded.
| Dashboard type | Content |
|---|---|
| Service health | Per-API state, error rate |
| Infrastructure | CPU, memory, network |
| Business | Revenue, user count, conversion |
| SLO | SLI actual vs target |
Alerts have the dichotomy of too few = miss it, too many = paralyzed, so per-importance routing (Slack for warnings, PagerDuty for critical) matters.
Alert design — fire on user impact
Conditions for good alerts: (1) response definitely needed, (2) response possible, (3) should move now - 3 satisfied. Alerts off these produce alert fatigue and cause missing real critical events.
| Good alerts | Bad alerts |
|---|---|
| SLO violation | CPU over 80% (auto-recovers) |
| Sudden error rate spike | One-shot error |
| Clear user impact | ”Somehow slow” |
| Has response procedures | Unclear who does what |
Alerts should fire on user impact, not “symptoms.” High CPU itself doesn’t impact users.
“Whether the machine is suffering” and “whether humans are suffering” are different things - the standard lesson of monitoring design. Cases where teams set CPU usage 80% alerts to feel safe, with no one’s dashboard red but Twitter flowing with “can’t log in” reports - are commonly heard. The cause is DB connection pool exhaustion, with CPU actually idle and only the app having everyone wait - the typical pattern.
An implementation example of an SLO burn-rate alert
The modern way for “SLO-violation-based” alerts is firing by error-budget consumption speed (burn rate). Not mere threshold exceeding, but a mechanism detecting “at this pace, the budget will be exhausted” - recommended as standard by the Google SRE Workbook.
| Severity | Condition | Destination |
|---|---|---|
| Critical (immediate response) | 2% budget consumed in 1 hour (burn rate > 14.4x) | PagerDuty |
| High (within hours) | 5% budget consumed in 6 hours (burn rate > 6x) | PagerDuty |
| Warning (within business hours) | 10% budget consumed in 3 days (burn rate > 1x) | Slack |
# Prometheus example (availability SLO 99.9%, monthly budget 43.2 min)
alert: ErrorBudgetBurnRateCritical
expr: (1 - availability_slo:ratio_rate5m) > (14.4 * 0.001)
and (1 - availability_slo:ratio_rate1h) > (14.4 * 0.001)
for: 2m
“Fire when CPU exceeds 80%” is outdated. Fire by burn rate - the modern standard.
Three scenarios
If you are building solo or on a small SaaS
Cloud-native CloudWatch or Cloud Monitoring plus UptimeRobot runs from a few dollars a month. Add the Grafana Cloud free tier when you want your own dashboards. With no dedicated SRE, standardising only the instrumentation through the OpenTelemetry SDK and leaving the backend to a SaaS is the configuration with the least operational load.
If you are a small or mid-size SaaS
Grafana Cloud (the LGTM stack) with OpenTelemetry is the main candidate. The LGTM stack carries a real learning cost, and it comes in far cheaper than Datadog and can be run by two or three SREs. If you use Datadog or New Relic, billing climbs sharply somewhere past a hundred hosts, which is the point to review it.
If you are a large enterprise
With a team that can run it, self-hosting Prometheus, Grafana, Loki, Tempo and OpenTelemetry avoids vendor lock-in and keeps confidential data inside — at the price of five or more dedicated SREs. For more than a hundred microservices on a SaaS, the choice becomes an enterprise plan from Datadog or Dynatrace.
Monitoring-cost / alert-operation numerical gates
Note: Industry baseline values as of April 2026. Will become outdated as technology and the talent market shift, so requires periodic updates.
For monitoring, not “install and feel safe” - the key of operations is tracking cost and signal quality numerically.
| Metric | Recommended | What to do if exceeded |
|---|---|---|
| Monthly monitoring foundation cost | 5-10% of infrastructure cost | Sampling, retention shortening |
| Production DEBUG-log output | Forbidden | Narrow to INFO+ |
| Alerts fired / week | 10 or fewer | Noise reduction, move to SLO-based |
| Alert response rate | 90%+ | Delete if firing but no one looks |
| MTTA | Within 5 min | Review on-call regime |
| MTTR | Within 30 min | Maintain Runbooks |
| Lighthouse Observability score | 90 or more | Review metric design |
| Structured-log rate | 100% | Don’t accept non-JSON |
| OpenTelemetry adoption rate | 100% for new | Avoid vendor-specific SDKs |
Alerts firing 10+ /week is a “noise-ization sign.” Datadog over $3k/month is a guideline for over-investment at startup scale. Continuing to output DEBUG logs in production easily hits CloudWatch Logs $10k+/month - the typical accident.
Monitoring cost has 10% of infrastructure cost as upper bound. Cut via sampling and retention if exceeded.
AI decision axes — Prepare for root-cause analysis by AI
AI-driven root-cause analysis (RCA) has become practical
Datadog’s “Watchdog RCA” and New Relic’s “AI Insights” provide features that cross-analyze multiple signals (logs, metrics, traces) to estimate incident root causes. The prerequisite for these to work is that the 3 pillars (logs, metrics, traces) are correlated via trace_id.
With unified instrumentation via OpenTelemetry, the log → trace → metrics correlation is automatically built on the tool side, maximizing AI RCA accuracy.
Observability backend selection and AI feature gaps
As of 2026, AI feature maturity varies by observability tool. Datadog covers anomaly detection, RCA, and Runbook recommendation with AI; Honeycomb broadens AI intervention space with BubbleUp (high-cardinality analysis); Grafana Cloud has implemented LLM-based log-query generation. “How far AI-powered diagnostic support goes” is now an evaluation axis in tool selection.
Pitfalls and forbidden moves
Here are the six most dangerous ways monitoring goes wrong. Every one has the structure of having been configured but not actually operated.
| Forbidden move | Why it is bad → what to do instead |
|---|---|
| Alerting only on static CPU and memory thresholds | false positives accumulate until nobody looks when it fires → switch to SLO burn rate |
| Leaving DEBUG-level logs on in production | the logging bill passes five figures a month → restrict to INFO and above and monitor the cost |
| Sending every alert to one Slack channel | a real incident is buried in everyday noise → route by severity, with the serious ones to PagerDuty |
| Tracking performance on averages alone | the slowest one percent of users is invisible → track P95 and P99 |
| Running microservices with no tracing | you cannot tell which service is slow and MTTR runs into hours → instrument with OpenTelemetry |
| Running monitoring on the same network as production | the same structure as the six-hour Facebook outage → keep monitoring independent of what it watches |
Instrumenting with a proprietary SDK is a breeding ground for lock-in as well. A future move means rewriting everything, so for anything new OpenTelemetry is the only choice.
Author’s note - cases of “all-green dashboards” with fires beneath
Cases of “monitored but couldn’t notice” have become standard talking points in operational fields.
At a certain SaaS, dashboards of CPU, memory, and network all stayed green while Twitter had hundreds of “can’t log in” reports flowing - a near-miss. The cause was DB connection pool exhaustion, with CPU actually idle and only the app having everyone wait. The trap was “feeling safe looking at infrastructure metrics,” with not measuring SLO (user impact) the root cause.
Another, the October 2021 Facebook/Instagram 6-hour outage, where BGP-config error made servers invisible from outside, but also internal monitoring tools and entry/exit systems all depended on the same network, so engineers couldn’t enter the data center and recovery was delayed - told as a case of “no monitoring of monitoring.” Estimated $60M+ in ad-revenue loss alone.
Both have design gaps in “what to measure” and “monitoring system independence” as lethal blows, slapping home that firing on user impact and separating monitoring from target are both required.
What to decide - what is your project’s answer?
For each of the following, try to articulate your project’s answer in 1-2 sentences. Starting work with these vague always invites later questions like “why did we decide this again?”
- Measurement SDK (OpenTelemetry recommended / vendor-specific)
- Backend (OSS LGTM / Datadog / New Relic)
- Metric design (USE, RED, SLI)
- Log strategy (collection scope, retention)
- Distributed tracing (all requests / sampling)
- Alert design (SLO-based, channel separation)
- Dashboards (service health, business)
Related Articles
Summary
This article covered monitoring and observability, including the 3 pillars, OpenTelemetry, main tools, SLO burn-rate alerts, and structured data for AI diagnosis.
Unify measurement with OpenTelemetry, decide SaaS vs OSS by ops regime, alerts on user-impact basis, organize structured data AI can read. That is the practical answer for monitoring and observability in 2026.
Next time we’ll cover log design (structured logs, PII protection, retention).
Back to series TOC -> ‘Architecture Crash Course for the Generative-AI Era’: How to Read This Book
I hope you’ll read the next article as well.
Also popular with readers
📚 Series: Architecture Crash Course for the Generative-AI Era (68/95)