Use AWS CloudFormation to create a scaling policy
The following example shows how to configure model auto scaling on an endpoint using AWS CloudFormation.
Endpoint: Type: "AWS::SageMaker::Endpoint" Properties: EndpointName:yourEndpointNameEndpointConfigName:yourEndpointConfigNameScalingTarget: Type: "AWS::ApplicationAutoScaling::ScalableTarget" Properties: MaxCapacity:10MinCapacity:2ResourceId: endpoint/my-endpoint/variant/my-variantRoleARN:arnScalableDimension: sagemaker:variant:DesiredInstanceCount ServiceNamespace: sagemaker ScalingPolicy: Type: "AWS::ApplicationAutoScaling::ScalingPolicy" Properties: PolicyName:my-scaling-policyPolicyType: TargetTrackingScaling ScalingTargetId: Ref: ScalingTarget TargetTrackingScalingPolicyConfiguration: TargetValue:70.0ScaleInCooldown:600ScaleOutCooldown:30PredefinedMetricSpecification: PredefinedMetricType: SageMakerVariantInvocationsPerInstance
For more information, see Create Application Auto Scaling resources with AWS CloudFormation in the Application Auto Scaling User Guide.