What's in this module
TOPIC 1 Observability definition & the 3 pillars
Observability: "A measure of how well we can understand a system from the work it does, and how to make it better." In Kubernetes, observability rests on three pillars — metrics, logs, and traces — each answering a different question about your system's health and behavior.
Metrics
Numerical measurements sampled over time. CPU=45%, memory=2.1 GB, req/s=200. Used for dashboards, alarms, and autoscaling.
Logs
Discrete timestamped events. "ERROR: connection refused to db:5432". Used for debugging specific failures after they occur.
Traces
End-to-end request journeys across multiple services. Shows latency per hop: frontend → auth → DB. Used for root-cause analysis.
2. Logs — discrete events (text output from containers and services)
3. Traces — end-to-end request flow across multiple services
2. Performance and cost — optimise resource use
3. Trends — detect slow degradation over time
4. Troubleshooting — diagnose incidents faster
5. Learning and improvement — understand system behaviour
TOPIC 2 Why containers make observability hard
Traditional monitoring was designed for long-lived servers. Containers break those assumptions in four important ways, each requiring a different observability strategy.
2. Metric volume — hundreds of pods each emit many metrics
3. Transient containers — container stops → evidence disappears
4. OS is just one factor — container behaviour also shaped by orchestration, networking, sidecars
TOPIC 3 Metrics: Prometheus + Grafana
Prometheus is the de-facto metrics system for Kubernetes. It scrapes /metrics endpoints from pods and the control plane, stores time-series data, and answers queries via PromQL. Grafana visualises Prometheus data in dashboards.
Prometheus metric format
metric_name{"tag"="value"} value- e.g.
http_requests_total{method="GET"} 1234 - Plain text, exposed at
/metricsendpoint - Prometheus scrapes on a configurable interval
Control plane metrics
- Access via:
kubectl get --raw /metrics - Includes API server latency, etcd operations
- Scheduler + controller-manager metrics available
- Prometheus can scrape these same endpoints
metric_name{"tag"="value"} valueExample:
cpu_usage{pod="web-abc",ns="prod"} 0.45It is plain text served at an HTTP
/metrics endpoint. Prometheus polls these endpoints on a schedule and stores the time-series data.rate(http_requests_total[5m]) for request rate, or sum by (pod) (container_memory_usage_bytes) for per-pod memory./metrics endpoint. You can access control plane metrics directly with kubectl get --raw /metrics. Prometheus is configured with a scrape job targeting the API server, scheduler, and controller-manager endpoints.prometheus.io/scrape: "true" are automatically included. ServiceMonitor CRDs (from the Prometheus Operator) provide a declarative way to configure scrape targets.TOPIC 4 CloudWatch Container Insights
CloudWatch Container Insights is the AWS-native observability solution for EKS. It collects metrics and logs from your cluster without requiring you to manage Prometheus. An agent installs on each node and ships data to CloudWatch where you can build dashboards and alarms.
Logs: container stdout/stderr
Aggregated at: cluster / node / pod / task / service level
Cluster → Node → Pod → Task → Service
This lets you drill from "why is the cluster CPU high?" all the way down to a specific container in a specific pod.
Prometheus: richer query language (PromQL), community ecosystem, custom app metrics, works across multiple clouds. Requires more operational overhead.
TOPIC 5 Managing logs: control plane & container logs
EKS produces two categories of logs with different collection mechanisms. Control plane logs are produced by Kubernetes system components and must be explicitly enabled in the EKS console. Application logs are container stdout/stderr, collected by a DaemonSet agent on each node.
Control plane logs (5 types)
- API server — all API requests
- Audit — who did what, when
- Authenticator — IAM auth events
- Controller manager — reconciliation loops
- Scheduler — pod placement decisions
Enabled per-type in EKS Console → Cluster → Logging tab
Application / container logs
- Container stdout and stderr
- Written to node filesystem by container runtime
- Collected by DaemonSet agent (Fluent Bit)
- Shipped to: CloudWatch, OpenSearch, S3, Kinesis
- Container crash = log survives in agent buffer
2. Audit
3. Authenticator
4. Controller manager
5. Scheduler
Each must be individually enabled in the EKS Console (Logging tab) or via the API. They are sent to CloudWatch Logs.
/aws/eks/<cluster-name>/cluster. From CloudWatch, you can use Logs Insights for querying, set metric filters, or export to S3 for long-term retention.2. Log aggregation — logs from all nodes combined in one store
3. Log analysis — search, filter, visualise (e.g. OpenSearch Dashboards)
TOPIC 6 Log routing: Fluent Bit, OpenSearch & Fargate
Fluent Bit is the lightweight log forwarder deployed as a DaemonSet on EC2 nodes. It reads container logs from the node and routes them to configurable destinations. For Fargate pods, log routing is configured via a dedicated aws-observability namespace ConfigMap.
| Destination | Use case | Notes |
|---|---|---|
| Amazon OpenSearch | Full-text search and Dashboards UI | Best for interactive log exploration |
| Amazon S3 | Long-term archival, Athena queries | Low cost for large volumes |
| CloudWatch Logs | AWS-native alerting, Logs Insights | Integrates with alarms and dashboards |
| Kinesis Data Firehose | High-throughput streaming to any destination | Used in Lab 4; bridges to S3, OpenSearch, Redshift |
aws-logging in a namespace called aws-observability (which must be labelled aws-observability: enabled).aws-observability(must have label
aws-observability: enabled)ConfigMap name:
aws-loggingContains an
output.conf key with Fluent Bit output plugin config (e.g. data_firehose).
[OUTPUT]
Name data_firehose
Match *
region us-west-2
delivery_stream my-stream-firehoseThe plugin name is
data_firehose. You specify the target delivery stream and region.
2. Amazon S3 — archival storage
3. CloudWatch Logs — AWS-native monitoring
4. Kinesis Data Firehose — streaming to multiple downstream targets
TOPIC 7 Application tracing: ADOT + AWS X-Ray
Distributed tracing solves a problem that logs and metrics cannot: understanding the causal chain of a single user request as it flows through many microservices. AWS Distro for OpenTelemetry (ADOT) collects trace spans from your applications and forwards them to AWS X-Ray, which visualises the end-to-end flow with latency per segment.
2. Individual operation insights — latency per operation within a service
3. Service-isolated issues — pinpoint exactly which service is the bottleneck
4. Root-cause analysis — trace the error back to its origin across the call chain
• A service map — visual graph of all services and their connections
• Trace timeline — Gantt-style view of latency per segment in a single request
• Bottleneck identification — which segment contributes most latency
• Error rates per service segment
Logs: Fluent Bit → OpenSearch / S3 / CloudWatch / Kinesis
Traces: ADOT → AWS X-Ray
For Fargate logs:
aws-observability namespace ConfigMap
TOPIC 8 Knowledge checks
Quick-fire recall. Try to answer from memory, then flip to verify.
A. Grafana B. Prometheus C. X-Ray D. Fluent Bit
Prometheus scrapes /metrics endpoints. Grafana (A) visualises Prometheus data. X-Ray (C) handles traces. Fluent Bit (D) forwards logs.
A. kubectl describe metrics B. kubectl get --raw /metrics C. aws eks get-metrics
The
--raw flag bypasses the kubectl object model and returns the raw HTTP response from the API server — in this case, the Prometheus-format metrics endpoint.aws-observability — the namespace must exist and be labelled aws-observability: enabled. Inside it, a ConfigMap named aws-logging contains the Fluent Bit output configuration./aws/eks/<cluster>/cluster.Lab 4 also configures OpenSearch for log search, and uses Prometheus + Grafana for metrics on the same cluster.
This hierarchy lets you start at the cluster level and drill down to a single container to find the source of a performance issue.