What's in this module

  1. Horizontal Pod Autoscaler (HPA)
  2. Vertical Pod Autoscaler (VPA)
  3. Cluster Autoscaler (CA)
  4. Karpenter
  5. HPA vs VPA vs CA vs Karpenter
  6. CI/CD pipelines
  7. GitOps with Argo CD
  8. Knowledge checks

TOPIC 1 Horizontal Pod Autoscaler (HPA)

HPA scales the number of Pod replicas in or out based on observed metrics (CPU utilisation, memory, or custom metrics). It does not change the resource requests of individual pods — it just adds or removes them.

Animation · HPA (horizontal) vs VPA (vertical) — side-by-side
What does HPA scale — pods or resources?
Pods (replicas). HPA scales the number of Pod instances horizontally — adding more identical pods under load, removing them when load drops. It does not resize CPU/memory on existing pods.
Write the HPA kubectl command for a deployment "myapp" targeting 50% CPU, min 3, max 10.
kubectl autoscale deployment myapp --cpu-percent=50 --min=3 --max=10

This creates an HPA object that watches average CPU utilisation across all pods and scales between 3 and 10 replicas to keep it near 50%.
What metrics can HPA use to make scaling decisions?
Built-in: CPU utilisation, memory utilisation
Custom metrics: any metric from Kubernetes Metrics API (e.g. requests-per-second, queue depth)
External metrics: metrics from outside the cluster (e.g. SQS queue length via KEDA)
What must pods have set for HPA CPU scaling to work?
Pods must have CPU resource requests set in their spec. HPA measures utilisation as a percentage of the requested CPU. Without requests, the Metrics Server cannot compute utilisation and HPA will not scale.

TOPIC 2 Vertical Pod Autoscaler (VPA)

VPA scales the CPU and memory resources available to individual pods — making each pod bigger or smaller rather than adding more of them. It is made of three deployments and has four operating modes.

VPA — 3 deployments

  • vpa-recommender — watches pod usage, produces resource recommendations
  • vpa-updater — evicts pods that need updated resources (in Recreate/InPlaceOrRecreate mode)
  • vpa-admission-controller — mutates new pods at creation time to apply recommended resources

VPA vs HPA — key distinction

  • HPA: changes replica count → more pods
  • VPA: changes resource requests → bigger pods
  • Do not combine HPA (CPU) + VPA (CPU) on same deployment — they conflict
  • Safe combo: HPA on custom metrics + VPA for right-sizing
Off
Recommendations generated but never applied. Read-only advisory mode.
Initial
Resources set at pod creation only. No evictions after pod is running.
Recreate
Evicts pods and recreates them with updated resource requests. Causes restarts.
InPlaceOrRecreate
Updates resources in-place if possible; falls back to evict+recreate. Preferred for minimal disruption.
Which VPA component generates resource recommendations?
vpa-recommender. It continuously watches historical and current pod CPU/memory usage and produces a recommendation. The recommender writes its output to the VPA object's status.recommendation field.
Which VPA mode causes the least disruption to running workloads?
Off (zero disruption — read-only) or InPlaceOrRecreate (tries in-place resize first, only evicts if in-place is impossible). For production workloads, start with Off to observe recommendations before enabling any mutation.
What does vpa-admission-controller do?
It is a MutatingAdmissionWebhook that intercepts pod creation requests and rewrites the resource requests/limits to match VPA's current recommendation — before the pod is scheduled. This means new pods start right-sized even before the updater acts.
Why can't you combine VPA (CPU) and HPA (CPU) on the same Deployment?
They fight each other: HPA adds replicas when CPU is high; VPA simultaneously increases each pod's CPU request. The Metrics Server sees changing denominators, causing oscillation. Safe combo: HPA on custom/external metrics + VPA for CPU/memory right-sizing.

TOPIC 3 Cluster Autoscaler (CA)

Cluster Autoscaler scales EC2 nodes in or out by adjusting existing Auto Scaling Groups. When pods are Pending due to insufficient node capacity, CA asks the ASG to add nodes. When nodes are underutilised, CA removes them.

Animation · Cluster Autoscaler scale-out flow

CA best practices

  • Use auto-discovery setup (preferred over static config)
  • Adjust min/max directly on the ASG — not via CA config
  • Use MixedInstancesPolicy for On-Demand + Spot mix
  • Use Expanders for multi-node-group clusters
  • Set --balance-similar-node-groups=true for balanced AZ spread

CA Expanders

  • random — default; pick randomly from eligible groups
  • most-pods — pick the group that can schedule the most pending pods
  • least-waste — pick the group with least CPU/memory waste
  • price — pick the cheapest option (cloud-provider support needed)
  • priority — user-defined priority labels on node groups
What triggers Cluster Autoscaler to add nodes?
Pods stuck in Pending state because no existing node has sufficient CPU/memory or matches the pod's node selector/affinity constraints. CA simulates scheduling and determines which node group to expand.
In the course's CA scale-out scenario (3 nodes, ASG max=6, then 14 replicas): what happens to the 2 pods after CA hits the ASG max?
They remain Pending indefinitely. CA cannot exceed the ASG's maximum. Once 6 nodes are running, CA has hit the hard limit — the 2 extra pods cannot be scheduled until you raise the ASG max or reduce the replica count.
What is MixedInstancesPolicy in the context of Cluster Autoscaler?
An ASG configuration that lets a single node group use a mix of On-Demand and Spot instances. CA can scale this group normally — it doesn't need to know which individual instance type is launched. Reduces cost while maintaining capacity.
What does the CA "auto-discovery" setup do?
CA tags ASGs with k8s.io/cluster-autoscaler/<cluster-name>=owned and k8s.io/cluster-autoscaler/enabled=true. The CA controller discovers all matching ASGs automatically — no need to list ASG names statically in the CA deployment config.

TOPIC 4 Karpenter

Karpenter is an open-source node provisioner maintained by AWS that scales nodes without Auto Scaling Groups. It provisions exactly the right node for each pending pod — faster and more flexible than Cluster Autoscaler.

Animation · Karpenter vs Cluster Autoscaler — key differences
What is the most fundamental difference between Karpenter and Cluster Autoscaler?
ASG dependency. CA requires pre-configured Auto Scaling Groups — it can only scale groups that exist. Karpenter provisions EC2 instances directly via the EC2 API, choosing instance type and size based on the pending pod's exact requirements — no ASG needed.
When should you choose Karpenter over Cluster Autoscaler?
Choose Karpenter when:
• Spiky demand — fast node provisioning needed
• Diverse compute — need many instance families/sizes
• Cost optimisation — want Karpenter to auto-select cheapest instance that fits
• Spot workloads — Karpenter handles spot interruption and rebalancing natively
What annotation prevents Karpenter from evicting a critical pod?
karpenter.sh/do-not-evict: "true" placed on the pod annotation. This tells Karpenter's consolidation logic to skip evicting that pod when it tries to remove underutilised nodes. Use for stateful workloads, long-running jobs, or anything that cannot tolerate rescheduling.
What are "layered constraints" in Karpenter?
Karpenter applies constraints from multiple layers: cluster-level NodePool constraints → workload pod spec (node selector, affinity, tolerations). Each layer narrows the provisioning decision. For example: the NodePool allows only us-east-1a + us-east-1b, and the pod's affinity further restricts to GPU nodes.
Who maintains Karpenter and what is its open-source status?
Karpenter is an open-source project originally created and currently maintained by AWS. It was donated to the CNCF ecosystem. It works with any Kubernetes cluster on any cloud, though it has native AWS EC2 integration via the karpenter-provider-aws.

TOPIC 5 Four autoscalers — side-by-side comparison

EKS has four complementary autoscalers that operate at different layers of the stack. They are not mutually exclusive — production clusters typically combine HPA + Karpenter (or HPA + CA).

Autoscaler What scales Mechanism K8s object Best for
HPA Pod replicas (out/in) Metrics Server → CPU/memory/custom metric → adjust replica count HorizontalPodAutoscaler Stateless services with variable traffic
VPA Pod resources (up/down) Recommender → Updater evicts → Admission Controller injects new requests VerticalPodAutoscaler Right-sizing; stateful pods; initial resource tuning
Cluster Autoscaler EC2 nodes (out/in) Watches Pending pods → calls ASG to add/remove nodes None (controller deployment) Standard EKS clusters with managed node groups + ASGs
Karpenter EC2 nodes (out/in) Watches Pending pods → directly provisions EC2 via NodePool + NodeClaim NodePool, NodeClaim (CRDs) Spiky demand; diverse instance types; cost-optimised Spot
Which two autoscalers operate at the node level (not pod level)?
Cluster Autoscaler and Karpenter both scale EC2 nodes. The difference: CA works through existing ASGs, Karpenter provisions nodes directly. HPA and VPA operate at the pod level.
Can HPA and Karpenter be used together? If so, how do they cooperate?
Yes — this is the recommended production pattern. HPA scales pods in response to traffic. When HPA adds pods that cannot fit on existing nodes, those pods go Pending. Karpenter detects Pending pods and provisions the right-sized node. They work at different layers and complement each other.
Which autoscaler has the most complex architecture with 3 separate deployments?
VPA — it needs vpa-recommender, vpa-updater, and vpa-admission-controller all running. HPA is built into the K8s controller manager. CA and Karpenter each deploy as a single controller.

TOPIC 6 CI/CD pipelines

CI/CD (Continuous Integration / Continuous Delivery) automates the path from code commit to production. The course defines eight stages that every robust pipeline covers.

1
📋
Plan
2
💻
Code
3
🔨
Build
4
🧪
Test
5
📦
Release
6
🚀
Deploy
7
⚙️
Operate
8
📊
Monitor
Name the 8 stages of a CI/CD pipeline in order.
Plan → Code → Build → Test → Release → Deploy → Operate → Monitor

Memory trick: "Please Can Bob Take Really Detailed Operations Metrics"
What is the difference between CI (Continuous Integration) and CD (Continuous Delivery)?
CI — automatically builds and tests code on every commit (Plan → Test stages). Ensures the codebase is always in a releasable state.

CD — automatically delivers validated builds to a staging or production environment (Release → Deploy stages). May include a manual approval gate before production.
In an EKS CI/CD pipeline, where does the container image go after the Build stage?
The built and tested image is pushed to Amazon ECR (Elastic Container Registry). The Deploy stage then references the new ECR image tag in a Kubernetes Deployment manifest — either directly (imperative) or via a GitOps manifest update (declarative).

TOPIC 7 GitOps with Argo CD

GitOps treats Git as the single source of truth for both application and infrastructure configuration. Argo CD is the GitOps controller that continuously syncs your EKS cluster to match the desired state declared in Git — and automatically reverts unauthorized changes.

Animation · GitOps with Argo CD — app change + drift detection

GitOps — 4 principles

  • Git as single source of truth — all desired state lives in Git
  • Declarative — describe what you want, not how to get there
  • Approved changes applied automatically — merged = deployed
  • Software agents ensure state — controllers reconcile cluster to Git continuously

Argo CD — how it works

  • Watches a Git repo for manifest changes
  • Compares live cluster state to desired Git state
  • Syncs the cluster when drift is detected
  • Reverts unauthorized kubectl edits automatically
  • Provides a visual dashboard of sync status
Walk through the GitOps app code change scenario step by step.
1. Developer pushes code to app Git repo
2. CI pipeline builds image, runs tests
3. CI pushes new image to ECR
4. CI updates image tag in manifest Git repo
5. Argo CD detects manifest change
6. Argo CD syncs → rolls out new pods in EKS cluster
What happens if someone runs "kubectl edit deployment myapp" directly on a GitOps-managed cluster?
Argo CD detects drift — it continuously compares live cluster state to Git. The direct kubectl change creates a mismatch. On the next sync cycle (or immediately if auto-sync is on), Argo CD reverts the cluster to the state declared in Git, undoing the manual change.
Why is GitOps described as "declarative" rather than "imperative"?
Declarative = you describe the desired end state (e.g. "3 replicas of v2.1.0"). The GitOps controller figures out how to get there.
Imperative = you issue commands (e.g. "kubectl set image", "kubectl scale"). GitOps eliminates one-off imperative commands because they bypass Git and get reverted.
What are the two Git repositories in a GitOps setup and what does each contain?
App code repo — application source code. Changes here trigger the CI pipeline.
Manifest repo (infra/config repo) — Kubernetes YAML manifests (Deployments, Services, etc.) + infrastructure-as-code. Changes here are watched by Argo CD and applied to the cluster.
How does Argo CD enforce the GitOps "approved changes applied automatically" principle?
When auto-sync is enabled in Argo CD, any merge to the watched Git branch automatically triggers a sync to the cluster — no human operator needs to run kubectl or approve a deploy. The Git merge/pull request process IS the approval gate. Argo CD then ensures the cluster matches.

TOPIC 8 Knowledge checks & quick-fire recall

Official course knowledge-check questions plus quick-fire recall cards. Try to answer before flipping.

KC — You have a stateless web app with unpredictable traffic spikes. Which autoscaler should you use first?
A. VPA   B. HPA   C. Karpenter only   D. CA only
✓ B — HPA

HPA scales pod replicas in response to CPU/traffic. For spiky traffic on stateless apps, HPA is the primary tool. Karpenter/CA add nodes when HPA can't fit the extra pods.
KC — Which VPA component actually evicts pods that need updated resources?
A. vpa-recommender   B. vpa-admission-controller   C. vpa-updater
✓ C — vpa-updater

vpa-recommender generates recommendations. vpa-admission-controller injects resources at pod creation. vpa-updater is the only one that evicts running pods to apply new resource settings.
KC — Cluster Autoscaler cannot add more nodes. What is the most likely cause?
The Auto Scaling Group's maximum size has been reached. CA cannot scale beyond the ASG max. Fix: increase the max capacity on the ASG directly. Also check: IAM permissions for CA to modify the ASG, and that the CA pod is running and healthy.
KC — In GitOps, a developer directly edits a running pod using kubectl. What happens next?
Argo CD detects the drift between live cluster state and the Git manifest. On the next sync cycle (or immediately with auto-sync), it reverts the cluster to the Git-declared state. The manual change is overwritten.
Quick-fire: What does the VPA "Off" mode do?
Generates recommendations only — no changes are applied to pods. Use it to observe what VPA would recommend before enabling mutations. Zero risk to running workloads.
Quick-fire: What annotation prevents Karpenter from removing a node with a critical pod?
karpenter.sh/do-not-evict: "true" on the pod. This signals Karpenter's consolidation loop to leave the node running even if it's underutilised.
Quick-fire: Name the 4 GitOps principles.
1. Git as single source of truth
2. Declarative (not imperative)
3. Approved changes applied automatically
4. Software agents ensure consistent state (drift detection + auto-reconcile)
Quick-fire: Which autoscaler does NOT use Auto Scaling Groups?
Karpenter. It provisions EC2 instances directly via the EC2 Fleet API, choosing the right instance type based on pending pod requirements. Cluster Autoscaler requires pre-configured ASGs.
Quick-fire: What are the 3 deployments that make up VPA?
1. vpa-recommender — generates resource recommendations
2. vpa-updater — evicts pods needing resource changes
3. vpa-admission-controller — mutates new pods at creation
Quick-fire: In the GitOps Argo CD flow, what triggers Argo CD to sync the cluster?
A change in the manifest Git repo. Argo CD polls the repo (or uses webhooks) and detects that the desired state (Git) differs from the current cluster state. It then applies the diff to bring the cluster in line with Git.