Documentation
¶
Index ¶
- Constants
- func BuildInstanceTags(nodeClaim *coreapis.NodeClaim, nodeClass *v1alpha1.ECSNodeClass) map[string]string
- type CloudProvider
- func (c *CloudProvider) Create(ctx context.Context, nodeClaim *coreapis.NodeClaim) (*coreapis.NodeClaim, error)
- func (c *CloudProvider) Delete(ctx context.Context, nodeClaim *coreapis.NodeClaim) error
- func (c *CloudProvider) Get(ctx context.Context, providerID string) (*coreapis.NodeClaim, error)
- func (c *CloudProvider) GetInstanceTypes(ctx context.Context, nodePool *coreapis.NodePool) ([]*cloudprovider.InstanceType, error)
- func (c *CloudProvider) GetSupportedNodeClasses() []status.Object
- func (c *CloudProvider) IsDrifted(ctx context.Context, nodeClaim *coreapis.NodeClaim) (cloudprovider.DriftReason, error)
- func (c *CloudProvider) List(ctx context.Context) ([]*coreapis.NodeClaim, error)
- func (c *CloudProvider) Name() string
- func (c *CloudProvider) RepairPolicies() []cloudprovider.RepairPolicy
Constants ¶
const ( ImageDrift cloudprovider.DriftReason = "ImageDrift" VSwitchDrift cloudprovider.DriftReason = "VSwitchDrift" SecurityGroupDrift cloudprovider.DriftReason = "SecurityGroupDrift" NodeClassDrift cloudprovider.DriftReason = "NodeClassDrift" )
const ( ResourcePodsForExclusiveEniEnabled corev1.ResourceName = "alibabacloud.com/exclusive-eni" DefaultPodsLimit = 110 )
Variables ¶
This section is empty.
Functions ¶
func BuildInstanceTags ¶ added in v0.2.5
Types ¶
type CloudProvider ¶
type CloudProvider struct {
// contains filtered or unexported fields
}
CloudProvider implements the Karpenter CloudProvider interface for Alibaba Cloud This is the main integration point between Karpenter Core and Alibaba Cloud
func New ¶
func New( instanceTypeProvider *instancetype.Provider, instanceProvider *instance.Provider, recorder events.Recorder, kubeClient client.Client, imageFamilyProvider *imagefamily.Provider, securityGroupProvider *securitygroup.Provider, vswitchProvider *vswitch.Provider, instanceProfileProvider *instanceprofile.Provider, pricingProvider *pricing.Provider, launchTemplateProvider *launchtemplate.Provider, bootstrapProvider *bootstrap.Provider, networkConfig *cluster.NetworkConfig, ) *CloudProvider
New creates a new CloudProvider with injected dependencies
func (*CloudProvider) Create ¶
func (c *CloudProvider) Create(ctx context.Context, nodeClaim *coreapis.NodeClaim) (*coreapis.NodeClaim, error)
Create creates a new instance for the given NodeClaim.
Karpenter contract (upstream: cloudprovider/types.go) ¶
- Return cloudprovider.NewNodeClassNotReadyError() when the NodeClass is not ready. Karpenter treats this as a non-fatal signal and retries; any other error is treated as a hard failure.
- Return cloudprovider.NewCreateError() with a reason/message for structured failures (e.g. quota exhaustion, unsupported config) so the event is surfaced on the NodeClaim.
- The returned NodeClaim must have Labels populated (at minimum topology.kubernetes.io/zone, node.kubernetes.io/instance-type, karpenter.sh/capacity-type, kubernetes.io/arch, kubernetes.io/os) because Karpenter core reads those labels to build NodePool scheduling constraints.
- Set the NodeClass hash annotation (v1alpha1.AnnotationECSNodeClassHash) on the returned NodeClaim so that IsDrifted can detect config changes.
func (*CloudProvider) Delete ¶
Delete terminates the cloud instance backing the given NodeClaim.
Karpenter contract (upstream: cloudprovider/types.go) ¶
- Return cloudprovider.NewNodeClaimNotFoundError() when the instance is already gone. Karpenter keeps retrying Delete until it receives this error, so returning nil when the instance does not exist will cause an infinite retry loop.
- Return nil only after deletion has been successfully *triggered* (not necessarily completed); Karpenter will confirm via Get/List once the node disappears.
func (*CloudProvider) GetInstanceTypes ¶
func (c *CloudProvider) GetInstanceTypes(ctx context.Context, nodePool *coreapis.NodePool) ([]*cloudprovider.InstanceType, error)
GetInstanceTypes returns instance types supported by this provider for the given NodePool.
Karpenter contract (upstream: cloudprovider/types.go, controllers/nodeclaim/disruption/drift.go) ¶
## Offerings must include temporarily-unavailable entries (Available: false)
The drift controller calls GetInstanceTypes() every 30 min (starting 1 h after node creation) and runs instanceTypeNotFound(), which checks:
- Is the node's instance type present in the returned list? (lo.Find by name)
- Does any Offering match the node's (zone, capacityType) labels? (Offerings.HasCompatible)
HasCompatible does NOT filter by Offering.Available — it only checks requirements. Therefore: if zone inventory drops to zero and we omit the offering entirely, the drift controller sees "no compatible offering" and marks the node as InstanceTypeNotFound, triggering an unnecessary replacement.
Correct behaviour: always emit an Offering for every known (zone, capacityType) combination, even when inventory is zero. Set Offering.Available = false for out-of-stock combinations.
Callers that create NEW nodes use Offerings.Available() to filter — they correctly skip Available=false offerings. Only drift detection uses HasCompatible, which must find them.
## Offering.Available controls scheduling, not drift
- Scheduling path: Offerings.Available().HasCompatible(req) — needs Available=true.
- Drift path: Offerings.HasCompatible(req) — ignores Available.
- Price sort: also filters by Available=true (cloudprovider/types.go:228).
## InstanceType.Requirements must cover all known zones and capacity types
Requirements on the InstanceType struct are used as a pre-filter by the scheduler. They should be built from ALL offerings (including Available=false ones) so that the scheduler does not silently exclude an instance type from the result set during drift checks.
func (*CloudProvider) GetSupportedNodeClasses ¶
func (c *CloudProvider) GetSupportedNodeClasses() []status.Object
GetSupportedNodeClasses returns the CloudProvider NodeClass that implements status.Object
func (*CloudProvider) IsDrifted ¶
func (c *CloudProvider) IsDrifted(ctx context.Context, nodeClaim *coreapis.NodeClaim) (cloudprovider.DriftReason, error)
IsDrifted reports whether the running node has drifted from its desired configuration.
Karpenter contract (upstream: controllers/nodeclaim/disruption/drift.go) ¶
## Call order
The drift controller calls static checks first (NodePool hash, Requirements), then calls GetInstanceTypes() + instanceTypeNotFound(), and finally calls IsDrifted() here. IsDrifted is NOT called when the instance type can no longer be found in GetInstanceTypes(), so provider-level drift checks run only for nodes whose instance type is still known.
## Return value semantics
- Return ("", nil): node is not drifted — drift condition is cleared.
- Return (reason, nil): node is drifted for reason — drift condition is set and the node becomes a disruption candidate.
- Return ("", err): transient error — drift condition unchanged, reconciler retries.
## What to check here vs. in GetInstanceTypes
IsDrifted should cover cloud-provider-specific config drift (image, vSwitch, security group, NodeClass hash). Instance-type availability drift is handled upstream via instanceTypeNotFound; do not duplicate it here. Duplicating it risks false positives when inventory is temporarily zero (the upstream check is guarded by a 1-hour delay + 30-min cache; this one is not).
## Reconcile frequency
The drift reconciler re-queues every 5 minutes, so IsDrifted is called frequently. Avoid expensive or rate-limited API calls on the hot path; cache where possible.
func (*CloudProvider) Name ¶
func (c *CloudProvider) Name() string
Name returns the CloudProvider implementation name.
func (*CloudProvider) RepairPolicies ¶
func (c *CloudProvider) RepairPolicies() []cloudprovider.RepairPolicy
RepairPolicies returns the repair policies for Alibaba Cloud nodes