VPA Recommendation Mode: A Practical Rightsizing Workflow
Abdul Ali
August 28, 2025 · Updated September 27, 2026 · 8 min read

In short
VPA's update mode decides whether a recommendation restarts running pods, and jumping straight to an automated mode is how rightsizing becomes an incident. Off only records recommendations, Initial applies them to new pods, Recreate evicts running pods, and InPlaceOrRecreate resizes them in place on Kubernetes 1.33 and later but still evicts when it can't, while Auto is deprecated. Roll out per workload: start in Off, review the recommendations, move stateless workloads to Initial, and automate only with Pod Disruption Budgets in place.
The Vertical Pod Autoscaler (VPA) is a Kubernetes component that automatically adjusts the CPU and memory requests of containers to match actual usage, scaling limits along with them. Its Recommender does the analysis, but the setting that actually determines what happens to your running pods is the recommendation mode (the updateMode field in the VPA object), and picking the wrong one is how a rightsizing project turns into an unplanned incident.
For the full breakdown of VPA’s three components (Recommender, Updater, Admission Controller) and how the Recommender’s histogram-based algorithm works, see A Guide to Kubernetes VPA. This article assumes that background and focuses specifically on choosing and rolling out a recommendation mode safely.
The Recommendation Modes
VPA’s updatePolicy.updateMode field controls what the Updater is allowed to do with a recommendation once the Recommender has generated it:
Off: Recommendations are calculated and can be inspected viakubectl describe vpa, but nothing is applied automatically. Zero risk of disruption, but zero automation either.Initial: Recommendations are applied only at pod creation. Running pods are left alone; only new pods (from a deploy, scale-up, or restart) get the updated resources.Recreate: Recommendations are applied to running pods by evicting them through the Eviction API, so Pod Disruption Budgets are respected, and letting the workload controller recreate them with the new values. This is how you get rightsizing applied to pods that are already running, without waiting for their next natural restart.InPlaceOrRecreate: The Updater first tries to resize the running pod in place, using Kubernetes’ in-place pod resize, and only evicts it when an in-place resize isn’t possible. It requires Kubernetes 1.33 or later (in-place pod resize went beta in 1.33 and GA in 1.35) and has been generally available and enabled by default since VPA 1.6.Auto: Deprecated since VPA 1.5 and currently treated as an alias forRecreate. It still works, but new VPA objects should nameRecreateorInPlaceOrRecreateexplicitly, and existing ones should be migrated.
The mode is set in the VPA object itself:
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: example-vpa
namespace: default
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: example-deployment
updatePolicy:
updateMode: "Off" # Off | Initial | Recreate | InPlaceOrRecreate
A Practical Rollout Workflow
The mistake teams make with VPA isn’t choosing the wrong mode, it’s skipping straight to an automated mode on day one. Recreate evicts pods, and even InPlaceOrRecreate falls back to eviction when a resize can’t be done in place. A database or stateful service restarting on a recommendation nobody’s reviewed is exactly the kind of self-inflicted incident VPA is supposed to prevent, not cause. A safer rollout looks like this:
-
Start in
Off. Deploy the VPA object against your workload withupdateMode: "Off". A first recommendation appears within minutes, but it only reflects what the Recommender has seen so far. Let the workload run through at least a full weekly cycle, including its busiest periods, before you rely on it. (By default the Recommender builds its history from live metrics and saves it in checkpoint objects; if you configure Prometheus as its history source, it can backfill up to 8 days of past usage at startup.) -
Read the recommendations before trusting them.
kubectl describe vpa example-vpaThe output includes
target,lowerBound, andupperBoundfor CPU and memory. Compare these against what the workload is actually configured with today, a first recommendation that looks wildly different from current usage is worth investigating before you apply anything, not after. -
Move stateless, replicated workloads to
Initial. For a workload with multiple replicas and no meaningful state,Initialis a low-risk way to start getting rightsized pods without forcing a restart of anything currently running, new pods just come up correctly sized as normal deploys and scale events happen. -
Use
RecreateorInPlaceOrRecreatedeliberately, with Pod Disruption Budgets in place. For workloads that need the recommendation applied to already-running pods, preferInPlaceOrRecreateon clusters that support it, since most changes then happen without a restart. Either way, set a PDB first, because any change that can’t be made in place still goes through eviction. -
Graduate a workload to an automated mode only after you’ve validated it under
OfforInitial. By the time you switch, the recommendations shouldn’t be a surprise, you’ve already seen them.
What Actually Happens When an Automated Mode Applies a Change
The two automated modes take different paths, and knowing both is why the cautious rollout above matters.
Recreate (and the deprecated Auto): eviction
- The VPA Recommender reads usage metrics and generates a recommendation.
- The VPA Updater sees that a running pod’s resources are out of date and evicts it through the Eviction API, which honors any Pod Disruption Budget.
- The workload controller (for example, the Deployment’s ReplicaSet) notices the pod is gone and creates a replacement to restore the desired replica count.
- During creation, the VPA Admission Controller intercepts the pod creation request and injects the updated CPU and memory values.
- The new pod comes up with the recommended resources already applied.
InPlaceOrRecreate: resize first, evict as a fallback
- The Updater asks the kubelet to resize the running pod’s CPU and memory in place. For most changes the containers keep running; whether a container restarts depends on its
resizePolicy. - If the resize can’t be done in place, the Updater falls back to the eviction path above. That happens when the node doesn’t have room for the new size, when a resize stays deferred for more than a few minutes or runs for over an hour, when the change would move the pod to a different QoS class, and in some memory-limit decreases.
For a stateless pod behind a Service, an eviction is invisible to users. For a pod holding open connections, in-flight work, or local state, it’s a restart like any other, which is exactly why Off and Initial exist as lower-risk stepping stones. Also keep in mind that an in-place resize changes what the container is allowed to use, not how the application inside it behaves: a JVM with a fixed heap size, for example, won’t use extra memory until it restarts.
Common Pitfalls
- Switching every workload to an automated mode from the start. Treat mode as a per-workload decision, not a cluster default. A stateless API service and a database have very different restart tolerances.
- Leaving
Autoin existing manifests. It still behaves likeRecreate, but it’s deprecated. Replace it withRecreate, or withInPlaceOrRecreateif your clusters run Kubernetes 1.33 or later. - Skipping Pod Disruption Budgets. Without a PDB, the only brakes on evictions are the Updater’s own defaults: it won’t touch workloads with fewer than 2 replicas, and it evicts at most half of a workload’s replicas at a time. That’s rarely the right limit for a specific service.
- Assuming in-place means zero restarts.
InPlaceOrRecreatestill evicts when an in-place resize isn’t possible, so it needs the same PDBs and the same review asRecreate. - Never revisiting
Offmode recommendations.Offis only useful if someone actually looks at the output. Left unattended, it collects data nobody acts on. - Assuming
Initialis enough for cost savings. If your workload rarely restarts or redeploys,Initialalone won’t rightsize the pods that are already running. That requiresRecreateorInPlaceOrRecreate.
Rightsizing with Randoli
Randoli’s Cost Management for Kubernetes automates the parts of this workflow that are otherwise manual: it installs and configures VPA for you (unless your cluster already has it) and uses it as the source of its rightsizing data, applies the same cautious-rollout logic (starting conservative, validating before automating), and surfaces rightsizing recommendations alongside cost and the rest of your observability data instead of a separate kubectl describe call.
If you’re already running VPA yourself in Recreate, InPlaceOrRecreate, or Auto mode, talk to our team before installing Randoli to avoid the two systems fighting over the same pods.
Appendix: The Histogram Calculation, Bucket by Bucket
The Recommender’s histogram-based algorithm (introduced at a high level in A Guide to Kubernetes VPA) works like this under the hood:
- Histogram creation. For each container, VPA keeps a histogram of CPU usage samples and a histogram of memory peaks (the highest memory usage in each 24-hour interval). The buckets grow exponentially, so small values get fine-grained buckets and large values get wide ones.
- Decaying weights. Every sample adds weight to the bucket it falls into, and that weight halves every 24 hours by default. Recent behavior counts most, which is how VPA stays responsive to changing usage without whipsawing on every spike.
- Percentiles. The recommendation’s
targetcomes from the 90th percentile of each histogram,lowerBoundfrom the 50th, andupperBoundfrom the 95th. - Safety margin and minimums. VPA adds a 15% safety margin on top of those percentiles, and never recommends less than 25 millicores of CPU or 250 MB of memory per pod. After an OOM kill, it bumps the memory recommendation up (by 20%, or at least 100 MB).
- Limits. Limits aren’t calculated from a percentile. When VPA applies a new request, it scales the limit proportionally so the container keeps its original limit-to-request ratio. If you only want VPA to manage requests, set
controlledValues: RequestsOnlyin the resource policy.
Bucket math. Bucket indices are represented by N, with bucket size increasing exponentially: bucketSize = 0.01 * (1.05^N). For CPU, measured in cores, bucket 0 covers [0, 0.01), bucket 1 covers [0.01, 0.0205), and so on. A usage sample is placed into the bucket matching its value and adds its (decayed) weight to that bucket.
To find a percentile, say the 90th, VPA accumulates bucket weights from smallest to largest until the running total reaches 90% of the total weight across all buckets; the bucket where that threshold is crossed sets the percentile value.
Worked example. Say a container’s CPU usage samples (all with equal weight, for simplicity) fall into these ranges over a period:
| CPU usage | Samples | Running total |
|---|---|---|
| 0–100m | 10 | 10% |
| 100–200m | 20 | 30% |
| 200–300m | 50 | 80% |
| 300–400m | 15 | 95% |
| 400–500m | 5 | 100% |
- The 50th percentile (
lowerBound) is crossed in the 200–300m range, since the running total goes from 30% to 80% there. - The 90th percentile (
target) is crossed in the 300–400m range, from 80% to 95%. That’s where the request is based, plus VPA’s 15% safety margin. - The 95th percentile (
upperBound) is reached at the top of the 300–400m range.
Note that the busiest range, 200–300m with half of all samples, is not where the target lands. The 90th percentile deliberately sizes for the busier 10% of the time, not the most common load.
Related reading
- A Guide to Kubernetes VPA — VPA’s components and how the Recommender works, the background this article builds on.
- Rightsizing Kubernetes Workloads: A Practical Guide — rightsizing beyond VPA, across container, pod, namespace, and cluster levels.
- Guide to Kubernetes Cost Management — how rightsizing fits into the broader Kubernetes cost picture.
- Cost & FinOps — rightsizing recommendations, idle-workload detection, and chargeback without assembling VPA and OpenCost yourself.
