cloudprovider

package
v0.2.5 Latest Latest
Warning

This package is not in the latest version of its module.

Go to latest
Published: Sep 15, 2026 License: Apache-2.0 Imports: 36 Imported by: 0

Documentation

Index

Constants

View Source
const (
	ImageDrift         cloudprovider.DriftReason = "ImageDrift"
	VSwitchDrift       cloudprovider.DriftReason = "VSwitchDrift"
	SecurityGroupDrift cloudprovider.DriftReason = "SecurityGroupDrift"
	NodeClassDrift     cloudprovider.DriftReason = "NodeClassDrift"
)
View Source
const (
	ResourcePodsForExclusiveEniEnabled corev1.ResourceName = "alibabacloud.com/exclusive-eni"
	DefaultPodsLimit                                       = 110
)

Variables

This section is empty.

Functions

func BuildInstanceTags added in v0.2.5

func BuildInstanceTags(nodeClaim *coreapis.NodeClaim, nodeClass *v1alpha1.ECSNodeClass) map[string]string

Types

type CloudProvider

type CloudProvider struct {
	// contains filtered or unexported fields
}

CloudProvider implements the Karpenter CloudProvider interface for Alibaba Cloud This is the main integration point between Karpenter Core and Alibaba Cloud

func New

func New(
	instanceTypeProvider *instancetype.Provider,
	instanceProvider *instance.Provider,
	recorder events.Recorder,
	kubeClient client.Client,
	imageFamilyProvider *imagefamily.Provider,
	securityGroupProvider *securitygroup.Provider,
	vswitchProvider *vswitch.Provider,
	instanceProfileProvider *instanceprofile.Provider,
	pricingProvider *pricing.Provider,
	launchTemplateProvider *launchtemplate.Provider,
	bootstrapProvider *bootstrap.Provider,
	networkConfig *cluster.NetworkConfig,
) *CloudProvider

New creates a new CloudProvider with injected dependencies

func (*CloudProvider) Create

func (c *CloudProvider) Create(ctx context.Context, nodeClaim *coreapis.NodeClaim) (*coreapis.NodeClaim, error)

Create creates a new instance for the given NodeClaim.

Karpenter contract (upstream: cloudprovider/types.go)

  • Return cloudprovider.NewNodeClassNotReadyError() when the NodeClass is not ready. Karpenter treats this as a non-fatal signal and retries; any other error is treated as a hard failure.
  • Return cloudprovider.NewCreateError() with a reason/message for structured failures (e.g. quota exhaustion, unsupported config) so the event is surfaced on the NodeClaim.
  • The returned NodeClaim must have Labels populated (at minimum topology.kubernetes.io/zone, node.kubernetes.io/instance-type, karpenter.sh/capacity-type, kubernetes.io/arch, kubernetes.io/os) because Karpenter core reads those labels to build NodePool scheduling constraints.
  • Set the NodeClass hash annotation (v1alpha1.AnnotationECSNodeClassHash) on the returned NodeClaim so that IsDrifted can detect config changes.

func (*CloudProvider) Delete

func (c *CloudProvider) Delete(ctx context.Context, nodeClaim *coreapis.NodeClaim) error

Delete terminates the cloud instance backing the given NodeClaim.

Karpenter contract (upstream: cloudprovider/types.go)

  • Return cloudprovider.NewNodeClaimNotFoundError() when the instance is already gone. Karpenter keeps retrying Delete until it receives this error, so returning nil when the instance does not exist will cause an infinite retry loop.
  • Return nil only after deletion has been successfully *triggered* (not necessarily completed); Karpenter will confirm via Get/List once the node disappears.

func (*CloudProvider) Get

func (c *CloudProvider) Get(ctx context.Context, providerID string) (*coreapis.NodeClaim, error)

Get retrieves the current state of the instance

func (*CloudProvider) GetInstanceTypes

func (c *CloudProvider) GetInstanceTypes(ctx context.Context, nodePool *coreapis.NodePool) ([]*cloudprovider.InstanceType, error)

GetInstanceTypes returns instance types supported by this provider for the given NodePool.

Karpenter contract (upstream: cloudprovider/types.go, controllers/nodeclaim/disruption/drift.go)

## Offerings must include temporarily-unavailable entries (Available: false)

The drift controller calls GetInstanceTypes() every 30 min (starting 1 h after node creation) and runs instanceTypeNotFound(), which checks:

  1. Is the node's instance type present in the returned list? (lo.Find by name)
  2. Does any Offering match the node's (zone, capacityType) labels? (Offerings.HasCompatible)

HasCompatible does NOT filter by Offering.Available — it only checks requirements. Therefore: if zone inventory drops to zero and we omit the offering entirely, the drift controller sees "no compatible offering" and marks the node as InstanceTypeNotFound, triggering an unnecessary replacement.

Correct behaviour: always emit an Offering for every known (zone, capacityType) combination, even when inventory is zero. Set Offering.Available = false for out-of-stock combinations.

Callers that create NEW nodes use Offerings.Available() to filter — they correctly skip Available=false offerings. Only drift detection uses HasCompatible, which must find them.

## Offering.Available controls scheduling, not drift

  • Scheduling path: Offerings.Available().HasCompatible(req) — needs Available=true.
  • Drift path: Offerings.HasCompatible(req) — ignores Available.
  • Price sort: also filters by Available=true (cloudprovider/types.go:228).

## InstanceType.Requirements must cover all known zones and capacity types

Requirements on the InstanceType struct are used as a pre-filter by the scheduler. They should be built from ALL offerings (including Available=false ones) so that the scheduler does not silently exclude an instance type from the result set during drift checks.

func (*CloudProvider) GetSupportedNodeClasses

func (c *CloudProvider) GetSupportedNodeClasses() []status.Object

GetSupportedNodeClasses returns the CloudProvider NodeClass that implements status.Object

func (*CloudProvider) IsDrifted

func (c *CloudProvider) IsDrifted(ctx context.Context, nodeClaim *coreapis.NodeClaim) (cloudprovider.DriftReason, error)

IsDrifted reports whether the running node has drifted from its desired configuration.

Karpenter contract (upstream: controllers/nodeclaim/disruption/drift.go)

## Call order

The drift controller calls static checks first (NodePool hash, Requirements), then calls GetInstanceTypes() + instanceTypeNotFound(), and finally calls IsDrifted() here. IsDrifted is NOT called when the instance type can no longer be found in GetInstanceTypes(), so provider-level drift checks run only for nodes whose instance type is still known.

## Return value semantics

  • Return ("", nil): node is not drifted — drift condition is cleared.
  • Return (reason, nil): node is drifted for reason — drift condition is set and the node becomes a disruption candidate.
  • Return ("", err): transient error — drift condition unchanged, reconciler retries.

## What to check here vs. in GetInstanceTypes

IsDrifted should cover cloud-provider-specific config drift (image, vSwitch, security group, NodeClass hash). Instance-type availability drift is handled upstream via instanceTypeNotFound; do not duplicate it here. Duplicating it risks false positives when inventory is temporarily zero (the upstream check is guarded by a 1-hour delay + 30-min cache; this one is not).

## Reconcile frequency

The drift reconciler re-queues every 5 minutes, so IsDrifted is called frequently. Avoid expensive or rate-limited API calls on the hot path; cache where possible.

func (*CloudProvider) List

func (c *CloudProvider) List(ctx context.Context) ([]*coreapis.NodeClaim, error)

List retrieves all instances managed by Karpenter

func (*CloudProvider) Name

func (c *CloudProvider) Name() string

Name returns the CloudProvider implementation name.

func (*CloudProvider) RepairPolicies

func (c *CloudProvider) RepairPolicies() []cloudprovider.RepairPolicy

RepairPolicies returns the repair policies for Alibaba Cloud nodes

Jump to

Keyboard shortcuts

? : This menu
/ : Search site
f or F : Jump to
y or Y : Canonical URL