controllers

package
v0.0.0-...-773ecd9 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 14, 2026 License: MIT Imports: 38 Imported by: 0

Documentation

Overview

PDB-floor policy helpers.

These are the policy-side counterparts to the actuation helpers in pdbmutation.go: pure functions over the EvictionAutoScaler that decide *whether* a PDB floor should be pinned and publish that intent on Status.PDBFloorPinned. They never touch the PDB — the PDBToEvictionAutoScaler reconciler actuates the marker.

PDB-floor mutation.

During a PDB-blocked drain the controller pins the target's PDB to an absolute minAvailable floor (the partner's required-healthy count at the pre-surge baseline, Status.MinReplicas) so a replica surge converts into DisruptionsAllowed instead of being tracked away by a relative floor, then restores the partner's spec when the drain finishes. The floor is captured once and persisted (on the PDB annotation and Status.PDBFloorPinned); a partner overwriting the PDB mid-drain is honored, not clobbered. Helpers here are pure functions over the PDB; the reconcile loop performs the client Update.

Index

Constants

View Source
const (
	ResourceTypeDeployment = "Deployment"
	ResourceTypePDB        = "PodDisruptionBudget"
	APIVersionAppsV1       = "apps/v1"
	APIVersionPolicyV1     = "policy/v1"
)

Owner-reference GVK parts used when wiring PDB <-> EvictionAutoScaler ownership.

View Source
const (
	// EASSurgeFinalizer is added to the EvictionAutoScaler when it first surges a target and is
	// retained thereafter — it is NOT dropped when a surge is reverted during normal operation.
	// It is removed only at deletion time, once reconcileSurgeTeardown has reverted any surge the
	// CR still owns, so a mid-drain CR delete is held until that teardown runs. Removing it
	// releases the CR for garbage collection.
	EASSurgeFinalizer = "eviction-autoscaler.azure.com/surge-revert"

	// PDBFloorFinalizer is placed on the EvictionAutoScaler while its partner PDB carries a
	// floor mutation, so a mid-drain CR delete is held until the PDB actuator restores the
	// partner PDB (or determines restore is moot). Removing it releases the CR for garbage
	// collection.
	PDBFloorFinalizer = "eviction-autoscaler.azure.com/pdb-floor"
)

Finalizers the controller places on an EvictionAutoScaler. Both live here so the full set of teardown guarantees is discoverable in one place, even though each is actuated by a different reconciler: the EvictionAutoScaler reconciler writes the Deployment/HPA/KEDA surge and owns EASSurgeFinalizer; the PDB actuator writes the partner PDB and owns PDBFloorFinalizer. Each controller owns the finalizer for the object it writes.

View Source
const (
	// AnnotationOriginalPDBSpec holds the JSON snapshot of the partner's original
	// disruption fields; its presence marks a PDB as mutated by us.
	AnnotationOriginalPDBSpec = "eviction-autoscaler.azure.com/original-pdb-spec"

	// AnnotationPinnedFloor records the pinned floor on the PDB so it survives a lost
	// CR status write (Status.PDBFloorPinned).
	AnnotationPinnedFloor = "eviction-autoscaler.azure.com/pinned-floor"
)
View Source
const ControllerName = "EvictionAutoScaler"
View Source
const EvictionSurgeReplicasAnnotationKey = "evictionSurgeReplicas"
View Source
const NodeNameIndex = "spec.nodeName"
View Source
const OriginalMinReplicasAnnotationKey = "eviction-autoscaler.azure.com/original-min-replicas"
View Source
const PDBCreateAnnotationKey = "eviction-autoscaler.azure.com/pdb-create"
View Source
const PDBOwnedByAnnotationKey = "ownedBy"

Variables

This section is empty.

Functions

func CreatePDBForDeployment

func CreatePDBForDeployment(ctx context.Context, c client.Client, deployment *v1.Deployment) error

CreatePDBForDeployment creates a PDB for the given deployment with standard configuration

func HasAutoscaler

func HasAutoscaler(ctx context.Context, c client.Client, namespace, targetName, targetKind string) (bool, error)

HasAutoscaler returns true if an HPA or KEDA ScaledObject targets this workload. Returns an error on real API failures (not errNotFound) so the caller can retry.

func ParseZeroSurgeOverride

func ParseZeroSurgeOverride(raw string) (*intstr.IntOrString, error)

ParseZeroSurgeOverride parses the fleet-wide zero-maxSurge override value (the ZERO_SURGE_OVERRIDE controller env var). The value is an int-or-percentage resolved against minReplicas at drain time, mirroring Kubernetes maxSurge — e.g. "25%" or an absolute "10". An empty string, or a value that resolves to zero ("0"/"0%"), returns (nil, nil) so the feature stays off; a negative or malformed value returns an error so startup fails fast rather than misbehaving mid-drain.

func ResolveMinReplicas

func ResolveMinReplicas(ctx context.Context, c client.Client, namespace, targetName, targetKind string, deployReplicas int32) (int32, bool, error)

ResolveMinReplicas returns the effective minimum replica count for a workload. Priority: KEDA ScaledObject minReplicaCount > standalone HPA minReplicas > deployment.spec.replicas.

The strategies are mutually exclusive in detectSurgeApplier: when a KEDA ScaledObject is present, only the KEDA strategy is used. If a standalone HPA also targets the same deployment, detectSurgeApplier rejects the configuration with an error (unsupported). This function mirrors that precedence for baseline calculation but does not enforce the rejection — that is done by detectSurgeApplier.

KEDA-managed HPAs are filtered out by isKEDAManagedHPA, so only user-created standalone HPAs are considered at tier 2.

The returned bool indicates whether an autoscaler (KEDA or HPA) was found. When true, the int32 is the autoscaler's floor (which may be 0 for KEDA scale-to-zero). When false, the int32 is the deployReplicas fallback. Returns an error on real API failures so the caller can retry rather than using a wrong value.

Types

type AutoscalerToPDBReconciler

type AutoscalerToPDBReconciler struct {
	client.Client
	Scheme *runtime.Scheme
	Filter filter
}

AutoscalerToPDBReconciler watches HPA and KEDA ScaledObject changes and updates the PDB minAvailable to match their min replicas floor. This ensures the PDB stays correct when an autoscaler's minReplicas/minReplicaCount changes without a corresponding deployment spec change.

A single controller handles both HPA and ScaledObject because the reconcile logic is identical (resolve the target deployment from scaleTargetRef → resolve min replicas floor → update PDB). Two separate controllers would duplicate this logic without any benefit.

func (*AutoscalerToPDBReconciler) Reconcile

Reconcile is triggered when an HPA or ScaledObject changes. The request key is the autoscaler's namespace/name. We resolve the target deployment from its scaleTargetRef.

func (*AutoscalerToPDBReconciler) SetupWithManager

func (r *AutoscalerToPDBReconciler) SetupWithManager(mgr ctrl.Manager) error

SetupWithManager registers watches on HPA and KEDA ScaledObject resources. Events are enqueued with the autoscaler's own key; Reconcile resolves the target deployment from the autoscaler's scaleTargetRef.

type DeploymentSurgeApplier

type DeploymentSurgeApplier struct {
	// contains filtered or unexported fields
}

func (*DeploymentSurgeApplier) ApplySurge

func (d *DeploymentSurgeApplier) ApplySurge(ctx context.Context, surgeReplicas int32) error

func (*DeploymentSurgeApplier) IsSurgeActive

func (d *DeploymentSurgeApplier) IsSurgeActive() bool

func (*DeploymentSurgeApplier) Name

func (d *DeploymentSurgeApplier) Name() string

func (*DeploymentSurgeApplier) RecordedBaseline

func (d *DeploymentSurgeApplier) RecordedBaseline() (int32, bool)

func (*DeploymentSurgeApplier) RecordedSurge

func (d *DeploymentSurgeApplier) RecordedSurge() (int32, bool)

func (*DeploymentSurgeApplier) RevertSurge

func (d *DeploymentSurgeApplier) RevertSurge(ctx context.Context, originalMinReplicas int32) error

type DeploymentToPDBReconciler

type DeploymentToPDBReconciler struct {
	client.Client
	Scheme   *runtime.Scheme
	Recorder record.EventRecorder
	Filter   filter
}

DeploymentToPDBReconciler reconciles a Deployment object and ensures an associated PDB is created and deleted

func (*DeploymentToPDBReconciler) Reconcile

Reconcile watches for Deployment changes (created, updated, deleted) and creates or deletes the associated PDB. creates pdb with minAvailable to be same as replicas for any deployment

func (*DeploymentToPDBReconciler) SetupWithManager

func (r *DeploymentToPDBReconciler) SetupWithManager(mgr ctrl.Manager) error

SetupWithManager sets up the controller with the Manager.

type DeploymentWrapper

type DeploymentWrapper struct {
	// contains filtered or unexported fields
}

func (*DeploymentWrapper) AddAnnotation

func (d *DeploymentWrapper) AddAnnotation(status, newReplicas string)

AddAnnotation add new status annotation

func (*DeploymentWrapper) GetMaxSurge

func (d *DeploymentWrapper) GetMaxSurge() intstr.IntOrString

func (*DeploymentWrapper) GetReplicas

func (d *DeploymentWrapper) GetReplicas() int32

func (*DeploymentWrapper) Obj

func (d *DeploymentWrapper) Obj() client.Object

func (*DeploymentWrapper) ReadyReplicas

func (d *DeploymentWrapper) ReadyReplicas() int32

func (*DeploymentWrapper) RemoveAnnotation

func (d *DeploymentWrapper) RemoveAnnotation(status string)

RemoveAnnotation will delete specific status annotation

func (*DeploymentWrapper) SetReplicas

func (d *DeploymentWrapper) SetReplicas(replicas int32)

type EvictionAutoScalerReconciler

type EvictionAutoScalerReconciler struct {
	client.Client
	Scheme   *runtime.Scheme
	Recorder record.EventRecorder
	Filter   filter
	// ZeroSurgeOverride lets the controller surge a workload whose maxSurge
	// resolves to 0 — an explicit maxSurge: 0 (common under safe-deployment
	// guidance) or a Recreate strategy — which otherwise cannot surge and
	// would degrade. Note: an unset RollingUpdate strategy is NOT treated as
	// zero — Kubernetes defaults it to 25% at admission time, and GetMaxSurge
	// returns that default. When non-nil, its value is applied as the drain surge for such
	// workloads: an int-or-percentage resolved against minReplicas, mirroring
	// Kubernetes' own maxSurge semantics — e.g. "25%" (rounded up) or an absolute
	// "10". The actual surge stays demand-driven (minReplicas + displaced) and is
	// capped at this amount, so larger drains proceed in waves. It is a fleet-wide,
	// install-time knob (the ZERO_SURGE_OVERRIDE controller env var); nil (the
	// default) preserves today's degrade-on-zero behavior, so Cosmic — not
	// individual workload owners — decides whether it applies.
	ZeroSurgeOverride *intstr.IntOrString
	// PDBFloorMutationEnabled is the master switch for the PDB-floor pinning feature.
	// It ships OFF (dormant) and is wired from the ENABLE_PDB_FLOOR_MUTATION controller
	// env var at startup (main.go). Held as a per-reconciler field (not a package var)
	// so config is declared at construction and tests set it per instance.
	PDBFloorMutationEnabled bool
}

EvictionAutoScalerReconciler reconciles a EvictionAutoScaler object

func (*EvictionAutoScalerReconciler) Reconcile

func (*EvictionAutoScalerReconciler) SetupWithManager

func (r *EvictionAutoScalerReconciler) SetupWithManager(mgr ctrl.Manager) error

type HPASurgeApplier

type HPASurgeApplier struct {
	// contains filtered or unexported fields
}

func (*HPASurgeApplier) ApplySurge

func (h *HPASurgeApplier) ApplySurge(ctx context.Context, surgeReplicas int32) error

func (*HPASurgeApplier) IsSurgeActive

func (h *HPASurgeApplier) IsSurgeActive() bool

func (*HPASurgeApplier) Name

func (h *HPASurgeApplier) Name() string

func (*HPASurgeApplier) RecordedBaseline

func (h *HPASurgeApplier) RecordedBaseline() (int32, bool)

func (*HPASurgeApplier) RecordedSurge

func (h *HPASurgeApplier) RecordedSurge() (int32, bool)

func (*HPASurgeApplier) RevertSurge

func (h *HPASurgeApplier) RevertSurge(ctx context.Context, originalMinReplicas int32) error

type KEDASurgeApplier

type KEDASurgeApplier struct {
	// contains filtered or unexported fields
}

func (*KEDASurgeApplier) ApplySurge

func (k *KEDASurgeApplier) ApplySurge(ctx context.Context, surgeReplicas int32) error

func (*KEDASurgeApplier) IsSurgeActive

func (k *KEDASurgeApplier) IsSurgeActive() bool

func (*KEDASurgeApplier) Name

func (k *KEDASurgeApplier) Name() string

func (*KEDASurgeApplier) RecordedBaseline

func (k *KEDASurgeApplier) RecordedBaseline() (int32, bool)

func (*KEDASurgeApplier) RecordedSurge

func (k *KEDASurgeApplier) RecordedSurge() (int32, bool)

func (*KEDASurgeApplier) RevertSurge

func (k *KEDASurgeApplier) RevertSurge(ctx context.Context, originalMinReplicas int32) error

type NodeReconciler

type NodeReconciler struct {
	client.Client
	Scheme   *runtime.Scheme
	Recorder record.EventRecorder
}

EvictionAutoScalerReconciler reconciles a EvictionAutoScaler object

func (*NodeReconciler) Reconcile

func (r *NodeReconciler) Reconcile(ctx context.Context, req ctrl.Request) (ctrl.Result, error)

Reconcile is the main loop of the controller. It will look for unschedulded nodes and for every pod on the node

func (*NodeReconciler) SetupWithManager

func (r *NodeReconciler) SetupWithManager(mgr ctrl.Manager) error

type PDBToEvictionAutoScalerReconciler

type PDBToEvictionAutoScalerReconciler struct {
	client.Client
	Scheme   *runtime.Scheme
	Recorder record.EventRecorder
	Filter   filter
}

PDBToEvictionAutoScalerReconciler reconciles a PodDisruptionBudget object.

func (*PDBToEvictionAutoScalerReconciler) Reconcile

Reconcile reads the state of the cluster for a PDB and creates/deletes EvictionAutoScalers accordingly.

func (*PDBToEvictionAutoScalerReconciler) SetupWithManager

func (r *PDBToEvictionAutoScalerReconciler) SetupWithManager(mgr ctrl.Manager) error

type StatefulSetWrapper

type StatefulSetWrapper struct {
	// contains filtered or unexported fields
}

func (*StatefulSetWrapper) AddAnnotation

func (s *StatefulSetWrapper) AddAnnotation(status, newReplicas string)

AddAnnotation will reset and add new annotation map every time this func is called

func (*StatefulSetWrapper) GetMaxSurge

func (s *StatefulSetWrapper) GetMaxSurge() intstr.IntOrString

func (*StatefulSetWrapper) GetReplicas

func (s *StatefulSetWrapper) GetReplicas() int32

func (*StatefulSetWrapper) Obj

func (s *StatefulSetWrapper) Obj() client.Object

func (*StatefulSetWrapper) ReadyReplicas

func (s *StatefulSetWrapper) ReadyReplicas() int32

func (*StatefulSetWrapper) RemoveAnnotation

func (s *StatefulSetWrapper) RemoveAnnotation(status string)

RemoveAnnotation will delete specific status annotation

func (*StatefulSetWrapper) SetReplicas

func (s *StatefulSetWrapper) SetReplicas(replicas int32)

type SurgeApplier

type SurgeApplier interface {
	// ApplySurge sets the minimum replica count to surgeReplicas.
	// Callers may invoke this multiple times; implementations must be idempotent.
	ApplySurge(ctx context.Context, surgeReplicas int32) error
	// RevertSurge restores the original minimum replica count.
	RevertSurge(ctx context.Context, originalMinReplicas int32) error
	// IsSurgeActive returns true if a surge is currently in progress on the target.
	// Used during generation tracking to distinguish our own scaling from external changes.
	IsSurgeActive() bool
	// RecordedSurge returns the replica count recorded by the last ApplySurge (from the
	// evictionSurgeReplicas annotation) and whether it is present. Used by the
	// bail-on-replica-change guard to detect an external replica edit mid-surge.
	RecordedSurge() (int32, bool)
	// RecordedBaseline returns the pre-surge baseline recorded by ApplySurge (from the
	// original-min-replicas annotation) and whether it is present. Used to recover the
	// true baseline for an EvictionAutoScaler that lost its Status.MinReplicas (e.g. a
	// freshly recreated CR that started at 0 while a surge was already active).
	RecordedBaseline() (int32, bool)
	// Name returns a human-readable name for logging
	Name() string
}

SurgeApplier abstracts the mechanism for temporarily increasing minimum replicas. Exactly one implementation is used per deployment, determined by detectSurgeApplier:

  • KEDASurgeApplier: when a KEDA ScaledObject targets the deployment
  • HPASurgeApplier: when a standalone HPA targets the deployment (no KEDA)
  • DeploymentSurgeApplier: when neither KEDA nor HPA is present

KEDA + standalone HPA on the same target is unsupported and rejected by detectSurgeApplier.

For autoscaler strategies (HPA, KEDA): the autoscaler floor is raised first, then deployment replicas are set directly for immediate effect. On failure, the reconcile loop retries ApplySurge idempotently until the deployment write succeeds.

type Surger

type Surger interface {
	//GetGeneration() int64
	GetReplicas() int32
	// ReadyReplicas returns the number of currently Ready pods (status.readyReplicas), used to
	// measure how much of a requested surge has actually materialized (pods scheduled once the
	// Cluster Autoscaler brings up nodes), as opposed to the desired spec.replicas.
	ReadyReplicas() int32
	SetReplicas(int32)
	GetMaxSurge() intstr.IntOrString
	Obj() client.Object
	//Update(ctx context.Context, obj Object, opts ...UpdateOption) error
	AddAnnotation(string, string)
	RemoveAnnotation(string)
}

func GetSurger

func GetSurger(kind string) (Surger, error)

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL