How does the Horizontal Pod Autoscaler work?
Interviewer Expectations for this Question
01
Core Competency
Assesses fundamental understanding of Kubernetes conventions, runtime behavior, and memory/performance considerations.
02
Evaluation Criteria
Hiring managers look for precision, avoidance of ambiguous jargon, and ability to explain trade-offs under real production conditions.
Comprehensive Model Answer
Verified Solution
The Horizontal Pod Autoscaler, or HPA, scales the number of Pod replicas based on observed metrics.
- The metrics-server, or a custom or external adapter, supplies metrics. The HPA controller polls every 15 seconds by default.
- The desired replica count is currentReplicas multiplied by the ratio of the current metric to the target metric, clamped by minReplicas and maxReplicas.
- Metrics can be CPU, memory, or custom metrics such as requests per second or queue depth.
kubectl autoscale deployment web --cpu-percent=70 --min=2 --max=10
kubectl get hpa
Important points:
- Pods must have resource requests set or CPU utilisation is undefined.
- Add stabilisation windows and scale-down delay to avoid flapping.
- HPA scales Pods, not nodes, so combine it with the cluster autoscaler so new Pods can schedule.
- For event-driven workloads, KEDA scales on external metrics such as queue length.
Candidate Response Strategy & Interview Tips
- Start with a concise one-sentence summary: Deliver a direct, confident answer first before expanding into nuances.
- Demonstrate real-world trade-offs: Discuss where this approach excels and when you would avoid it in production systems.
- Discuss complexity & edge cases: Proactively explain time/space complexity or boundary conditions (null values, scale limits).
- Prepare for interviewer follow-ups: Technical hiring panels frequently probe deeper into concurrency, backward compatibility, or alternative libraries.