What's in this module

  1. Observability definition & the 3 pillars
  2. Why containers make observability hard
  3. Metrics: Prometheus + Grafana
  4. CloudWatch Container Insights
  5. Managing logs: control plane & container logs
  6. Log routing: Fluent Bit, OpenSearch, Fargate
  7. Application tracing: ADOT + X-Ray
  8. Knowledge checks

TOPIC 1 Observability definition & the 3 pillars

Observability: "A measure of how well we can understand a system from the work it does, and how to make it better." In Kubernetes, observability rests on three pillars — metrics, logs, and traces — each answering a different question about your system's health and behavior.

Animation · The 3 observability pillars
📊

Metrics

Numerical measurements sampled over time. CPU=45%, memory=2.1 GB, req/s=200. Used for dashboards, alarms, and autoscaling.

📋

Logs

Discrete timestamped events. "ERROR: connection refused to db:5432". Used for debugging specific failures after they occur.

🔗

Traces

End-to-end request journeys across multiple services. Shows latency per hop: frontend → auth → DB. Used for root-cause analysis.

What is the course definition of observability?
"A measure of how well we can understand a system from the work it does, and how to make it better." Observability is not just monitoring — it is the ability to ask arbitrary questions about system state and get answers.
Name the 3 pillars of observability.
1. Metrics — numerical measurements over time (CPU, memory, request rates)
2. Logs — discrete events (text output from containers and services)
3. Traces — end-to-end request flow across multiple services
What are the five values of observability insight?
1. Customer experience — know when users are impacted
2. Performance and cost — optimise resource use
3. Trends — detect slow degradation over time
4. Troubleshooting — diagnose incidents faster
5. Learning and improvement — understand system behaviour
Which pillar answers "what happened at 14:32:07 on pod web-abc"?
Logs — logs are discrete timestamped events. You search logs to find the exact error message, stack trace, or state change at a specific moment in a specific container.
Which pillar answers "which downstream service is causing slow checkout?"
Traces — distributed tracing follows a single request across all services and records latency per segment. You can see that checkout calls payment-service which takes 800 ms, vs. 5 ms for everything else.

TOPIC 2 Why containers make observability hard

Traditional monitoring was designed for long-lived servers. Containers break those assumptions in four important ways, each requiring a different observability strategy.

Name four reasons containers make observability harder than VMs.
1. Microservice complexity — many services interacting unpredictably
2. Metric volume — hundreds of pods each emit many metrics
3. Transient containers — container stops → evidence disappears
4. OS is just one factor — container behaviour also shaped by orchestration, networking, sidecars
Why are transient containers a problem for log analysis?
When a container crashes and restarts, any logs stored inside the container are lost. You must use an external log forwarder (like Fluent Bit) to stream logs off-node to durable storage before the container terminates — otherwise the evidence disappears with it.
Why doesn't traditional OS-level monitoring scale for microservices?
Traditional monitoring watches a handful of servers. A microservice cluster may have hundreds of pods, each emitting CPU, memory, network, and custom application metrics. Polling each pod individually is impractical — you need a system (like Prometheus) designed to scrape thousands of endpoints at scale.

TOPIC 3 Metrics: Prometheus + Grafana

Prometheus is the de-facto metrics system for Kubernetes. It scrapes /metrics endpoints from pods and the control plane, stores time-series data, and answers queries via PromQL. Grafana visualises Prometheus data in dashboards.

Animation · Prometheus + Grafana metrics pipeline

Prometheus metric format

  • metric_name{"tag"="value"} value
  • e.g. http_requests_total{method="GET"} 1234
  • Plain text, exposed at /metrics endpoint
  • Prometheus scrapes on a configurable interval

Control plane metrics

  • Access via: kubectl get --raw /metrics
  • Includes API server latency, etcd operations
  • Scheduler + controller-manager metrics available
  • Prometheus can scrape these same endpoints
What is the Prometheus metrics text format?
metric_name{"tag"="value"} value

Example: cpu_usage{pod="web-abc",ns="prod"} 0.45

It is plain text served at an HTTP /metrics endpoint. Prometheus polls these endpoints on a schedule and stores the time-series data.
What is PromQL?
PromQL (Prometheus Query Language) is the query language used to select, filter, and aggregate Prometheus time-series data. Examples: rate(http_requests_total[5m]) for request rate, or sum by (pod) (container_memory_usage_bytes) for per-pod memory.
How does Prometheus collect metrics from the Kubernetes control plane?
The same way it scrapes pods: via an HTTP /metrics endpoint. You can access control plane metrics directly with kubectl get --raw /metrics. Prometheus is configured with a scrape job targeting the API server, scheduler, and controller-manager endpoints.
What role does Grafana play alongside Prometheus?
Grafana is a dashboarding tool that queries Prometheus (via PromQL) and renders time-series data as graphs, gauges, and heatmaps. Prometheus stores and queries data; Grafana visualises it. They are often deployed together in a cluster using the kube-prometheus-stack Helm chart.
How does Prometheus discover which pods to scrape?
Prometheus uses Kubernetes service discovery — it queries the K8s API to find pods, services, and endpoints. Pods annotated with prometheus.io/scrape: "true" are automatically included. ServiceMonitor CRDs (from the Prometheus Operator) provide a declarative way to configure scrape targets.

TOPIC 4 CloudWatch Container Insights

CloudWatch Container Insights is the AWS-native observability solution for EKS. It collects metrics and logs from your cluster without requiring you to manage Prometheus. An agent installs on each node and ships data to CloudWatch where you can build dashboards and alarms.

COLLECTS
CPU & Memory
Per cluster, node, pod, container
COLLECTS
Disk & Network
Disk I/O, bytes sent/received
COLLECTS
Container Logs
stdout/stderr shipped to CloudWatch Logs
INTEGRATES
CloudWatch Alarms
Set alarms on any collected metric
What does CloudWatch Container Insights collect?
Metrics: CPU, memory, disk, network, container data
Logs: container stdout/stderr

Aggregated at: cluster / node / pod / task / service level
How is CloudWatch Container Insights installed on EKS nodes?
As a DaemonSet agent installed on each node. The CloudWatch Agent (or the newer Container Insights for EKS add-on) runs one pod per node, collects local metrics and logs, and ships them to CloudWatch. It requires appropriate IAM permissions on the node role.
What are the aggregation levels supported by Container Insights?
Container Insights aggregates metrics at five levels:
Cluster → Node → Pod → Task → Service

This lets you drill from "why is the cluster CPU high?" all the way down to a specific container in a specific pod.
Container Insights vs Prometheus — when would you choose each?
Container Insights: simpler setup, AWS-native, integrates with CloudWatch alarms and Logs Insights. Good for teams already in the CloudWatch ecosystem.

Prometheus: richer query language (PromQL), community ecosystem, custom app metrics, works across multiple clouds. Requires more operational overhead.

TOPIC 5 Managing logs: control plane & container logs

EKS produces two categories of logs with different collection mechanisms. Control plane logs are produced by Kubernetes system components and must be explicitly enabled in the EKS console. Application logs are container stdout/stderr, collected by a DaemonSet agent on each node.

Control plane logs (5 types)

  • API server — all API requests
  • Audit — who did what, when
  • Authenticator — IAM auth events
  • Controller manager — reconciliation loops
  • Scheduler — pod placement decisions

Enabled per-type in EKS Console → Cluster → Logging tab

Application / container logs

  • Container stdout and stderr
  • Written to node filesystem by container runtime
  • Collected by DaemonSet agent (Fluent Bit)
  • Shipped to: CloudWatch, OpenSearch, S3, Kinesis
  • Container crash = log survives in agent buffer
Name the 5 EKS control plane log types.
1. API server
2. Audit
3. Authenticator
4. Controller manager
5. Scheduler

Each must be individually enabled in the EKS Console (Logging tab) or via the API. They are sent to CloudWatch Logs.
How are application container logs collected in EKS?
Containers write logs to stdout/stderr. The container runtime writes these to the node's filesystem. A DaemonSet agent (typically Fluent Bit) runs on each node, tails those log files, and ships them to a configured destination. This ensures logs survive container restarts.
Where do EKS control plane logs go once enabled?
Amazon CloudWatch Logs. Each control plane log type gets its own log group: /aws/eks/<cluster-name>/cluster. From CloudWatch, you can use Logs Insights for querying, set metric filters, or export to S3 for long-term retention.
What are the three stages of the log routing workflow?
1. Log collection and forwarding — agent on node reads and forwards logs
2. Log aggregation — logs from all nodes combined in one store
3. Log analysis — search, filter, visualise (e.g. OpenSearch Dashboards)

TOPIC 6 Log routing: Fluent Bit, OpenSearch & Fargate

Fluent Bit is the lightweight log forwarder deployed as a DaemonSet on EC2 nodes. It reads container logs from the node and routes them to configurable destinations. For Fargate pods, log routing is configured via a dedicated aws-observability namespace ConfigMap.

Animation · Fluent Bit log routing pipeline
Destination Use case Notes
Amazon OpenSearch Full-text search and Dashboards UI Best for interactive log exploration
Amazon S3 Long-term archival, Athena queries Low cost for large volumes
CloudWatch Logs AWS-native alerting, Logs Insights Integrates with alarms and dashboards
Kinesis Data Firehose High-throughput streaming to any destination Used in Lab 4; bridges to S3, OpenSearch, Redshift
What is Fluent Bit and how is it deployed in EKS?
Fluent Bit is a lightweight, high-performance log forwarder. In EKS it is deployed as a DaemonSet — one pod per node. It tails container log files from the node filesystem and routes them to configured output plugins (OpenSearch, S3, CloudWatch, Kinesis Firehose).
How does Fargate log routing differ from EC2 node log routing?
On EC2, Fluent Bit runs as a DaemonSet you deploy. On Fargate, there are no nodes to deploy DaemonSets on. Instead, Fargate reads log routing config from a ConfigMap named aws-logging in a namespace called aws-observability (which must be labelled aws-observability: enabled).
What namespace and ConfigMap name are required for Fargate log routing?
Namespace: aws-observability
(must have label aws-observability: enabled)

ConfigMap name: aws-logging
Contains an output.conf key with Fluent Bit output plugin config (e.g. data_firehose).
In the Fargate log routing ConfigMap, what OUTPUT plugin sends logs to Kinesis Data Firehose?
[OUTPUT]
Name data_firehose
Match *
region us-west-2
delivery_stream my-stream-firehose


The plugin name is data_firehose. You specify the target delivery stream and region.
What are the four common AWS destinations Fluent Bit can route container logs to?
1. Amazon OpenSearch — full-text search and dashboards
2. Amazon S3 — archival storage
3. CloudWatch Logs — AWS-native monitoring
4. Kinesis Data Firehose — streaming to multiple downstream targets

TOPIC 7 Application tracing: ADOT + AWS X-Ray

Distributed tracing solves a problem that logs and metrics cannot: understanding the causal chain of a single user request as it flows through many microservices. AWS Distro for OpenTelemetry (ADOT) collects trace spans from your applications and forwards them to AWS X-Ray, which visualises the end-to-end flow with latency per segment.

Animation · ADOT + X-Ray distributed trace flow
Why is traditional debugging insufficient for microservices?
Traditional debugging (inspecting a single service) doesn't scale because a request may touch dozens of services. You cannot tell which service is slow, which failed, or how errors propagate, by looking at one service's logs in isolation. Distributed tracing provides the full picture across all services.
What is AWS Distro for OpenTelemetry (ADOT)?
ADOT is AWS's distribution of the OpenTelemetry Collector. It runs as a sidecar or DaemonSet in your cluster, collects trace spans (and metrics) from instrumented applications, and forwards them to AWS X-Ray (and other backends). It is vendor-neutral: OpenTelemetry SDKs work with many backends.
What four insights does distributed tracing provide that logs cannot?
1. Service discovery — automatically maps which services exist and talk to each other
2. Individual operation insights — latency per operation within a service
3. Service-isolated issues — pinpoint exactly which service is the bottleneck
4. Root-cause analysis — trace the error back to its origin across the call chain
What does AWS X-Ray show you?
X-Ray provides:
• A service map — visual graph of all services and their connections
• Trace timeline — Gantt-style view of latency per segment in a single request
• Bottleneck identification — which segment contributes most latency
• Error rates per service segment
What is the full observability tooling stack summarised at module end?
Metrics: Prometheus + Grafana or CloudWatch Container Insights
Logs: Fluent Bit → OpenSearch / S3 / CloudWatch / Kinesis
Traces: ADOT → AWS X-Ray

For Fargate logs: aws-observability namespace ConfigMap

TOPIC 8 Knowledge checks

Quick-fire recall. Try to answer from memory, then flip to verify.

KC 1 — Which tool scrapes /metrics endpoints and stores time-series data?
A. Grafana   B. Prometheus   C. X-Ray   D. Fluent Bit
✓ B — Prometheus

Prometheus scrapes /metrics endpoints. Grafana (A) visualises Prometheus data. X-Ray (C) handles traces. Fluent Bit (D) forwards logs.
KC 2 — What command exposes control plane metrics directly?
A. kubectl describe metrics   B. kubectl get --raw /metrics   C. aws eks get-metrics
✓ B — kubectl get --raw /metrics

The --raw flag bypasses the kubectl object model and returns the raw HTTP response from the API server — in this case, the Prometheus-format metrics endpoint.
KC 3 — Fargate log routing is configured in which namespace?
aws-observability — the namespace must exist and be labelled aws-observability: enabled. Inside it, a ConfigMap named aws-logging contains the Fluent Bit output configuration.
KC 4 — Which AWS service visualises distributed traces sent by ADOT?
AWS X-Ray. ADOT (AWS Distro for OpenTelemetry) is the collector that gathers spans from your instrumented services. X-Ray is the storage and visualisation backend that shows you the service map and per-segment latency.
KC 5 — EKS control plane logs are enabled per-type in which console location?
EKS Console → your cluster → Logging tab. You toggle each of the 5 log types (API server, Audit, Authenticator, Controller manager, Scheduler) individually. Once enabled, they stream to CloudWatch Logs under /aws/eks/<cluster>/cluster.
KC 6 — Lab 4 routes container logs through which services in order?
Fluent Bit DaemonSet → Kinesis Data Firehose → Amazon S3

Lab 4 also configures OpenSearch for log search, and uses Prometheus + Grafana for metrics on the same cluster.
KC 7 — Which log type records "who did what, when" in the Kubernetes API?
Audit logs — they record every request to the Kubernetes API server including the caller's identity, the resource accessed, and the action performed. Essential for security forensics and compliance.
KC 8 — CloudWatch Container Insights aggregates at which 5 levels?
Cluster → Node → Pod → Task → Service

This hierarchy lets you start at the cluster level and drill down to a single container to find the source of a performance issue.