cloud-platform-skills / skills
mchittineni/cloud-platform-skills/.cursor/rules/skills/scalability-high-availability-patterns.mdc
Scalability and high availability: multi-AZ and active-active topology, global load balancing and failover, HPA and KEDA autoscaling on custom or queue-depth metrics, circuit breakers, bulkheads, retries with jitter, and rate limiting. Use when a service must survive an availability zone failure, when it collapses under traffic spikes and drags downstream services with it, or when CPU is the wrong autoscaling signal.
What's in it
- Scalability & High Availability (HA) Architecture Patterns
- When to Use This Skill
- 1. High Availability Architecture Blueprint
- 2. Horizontal Pod Autoscaler (HPA) with Custom Metrics
- 3. Best Practices & Anti-Patterns
- 4. Event-Driven Autoscaling with KEDA
---
description: "Scalability and high availability: multi-AZ and active-active topology, global load balancing and failover, HPA and KEDA autoscaling on custom or queue-depth metrics, circuit breakers, bulkheads, retries with jitter, and rate limiting. Use when a service must survive an availability zone failure, when it collapses under traffic spikes and drags downstream services with it, or when CPU is the wrong autoscaling signal."
globs:
alwaysApply: false
---
# Scalability & High Availability (HA) Architecture Patterns
## When to Use This Skill
**Triggers — load this skill when:**
- A design must survive AZ or region loss without manual intervention
- Autoscaling on custom or event-driven metrics needs configuring
- Overload protection (circuit breaker, bulkhead, rate limit, backpressure) is missing
**Route elsewhere when:**
- Recovery from catastrophic loss -> `backup-and-disaster-recovery`
- Measuring headroom empirically -> `performance-load-testing`
- Mesh-level resilience policy -> `api-gateway-service-mesh`
## 1. High Availability Architecture Blueprint
```text
[Global Anycast / Route 53]
|
+----------------+----------------+
| |
[Region 1 (Primary)] [Region 2 (Secondary)]
| |
[Application Load Balancer] [Application Load Balancer]
| |
+--------+--------+ +--------+--------+
| | | | | |
[AZ-A] [AZ-B] [AZ-C] [AZ-A] [AZ-B] [AZ-C]
| | | | | |
+--------+--------+ +--------+--------+
| |
[Aurora Multi-AZ Primary] <------ [Aurora Global Read Replica]
```
---
## 2. Horizontal Pod Autoscaler (HPA) with Custom Metrics
```yaml
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: payment-service-hpa
namespace: production
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: payment-service
minReplicas: 3
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65
- type: Pods
pods:
metric:
name: http_requests_per_second
target:
type: AverageValue
averageValue: 1k
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 15
scaleDown:
stabilizationWindowSeconds: 300 # Prevent flapping
policies:
- type: Percent
value: 10
periodSeconds: 60
```
---
## 3. Best Practices & Anti-Patterns
- **Do**: Always distribute computing across a minimum of 3 Availability Zones (AZs) per region.
- **Do**: Use exponential backoff with jitter on all downstream API retries to prevent thundering herd crashes.
- **Don't**: Never scale down immediately without a stabilization window (`stabilizationWindowSeconds: 300`).
---
## 4. Event-Driven Autoscaling with KEDA
CPU is a poor proxy for demand in queue-driven and I/O-bound systems: the work is waiting in a
broker while CPU sits at 15%. KEDA scales on the backlog itself, and to zero when there is none.
```yaml
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: order-worker
spec:
scaleTargetRef: { name: order-worker }
minReplicaCount: 0 # scale-to-zero between bursts
maxReplicaCount: 200
pollingInterval: 15
cooldownPeriod: 300
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 300 # slow down, fast up
triggers:
- type: aws-sqs-queue
metadata:
queueURL: https://sqs.eu-west-1.amazonaws.com/1234/orders
queueLength: "20" # target backlog per replica
awsRegion: eu-west-1
- type: prometheus # triggers compose; the highest demand wins
metadata:
serverAddress: http://prometheus.monitoring:9090
query: sum(rate(orders_received_total[2m]))
threshold: "50"
```
Choosing the signal: scale on **queue depth or lag** for async work (SQS, Kafka consumer lag,
RabbitMQ), on **in-flight requests or RPS** for synchronous services, and on CPU only for
genuinely CPU-bound compute. Always set `maxReplicaCount` below what the downstream datastore
can absorb — otherwise autoscaling converts a traffic spike into a database outage.
More agent context in mchittineni/cloud-platform-skills
167 other files this repository gives its agents, the first 60 shown.
AGENTS.md
CLAUDE.md
Copilot instructions
Cursor rule
- .cursor/rules/00-index.mdc
- .cursor/rules/skills/ai-agent-security-llm-threats.mdc
- .cursor/rules/skills/api-gateway-service-mesh.mdc
- .cursor/rules/skills/aws-cloud-migration-strategies.mdc
- .cursor/rules/skills/aws-eks-enterprise-patterns.mdc
- .cursor/rules/skills/aws-iam-zero-trust-policies.mdc
- .cursor/rules/skills/azure-aks-enterprise-landing-zones.mdc
- .cursor/rules/skills/azure-cloud-engineering-patterns.mdc
- .cursor/rules/skills/backup-and-disaster-recovery.mdc
- .cursor/rules/skills/chaos-engineering-resilience-testing.mdc
- .cursor/rules/skills/cicd-pipeline-design.mdc
- .cursor/rules/skills/cloud-native-microservices-patterns.mdc
- .cursor/rules/skills/cloud-security-posture-cspm-cis.mdc
- .cursor/rules/skills/configuration-management-ansible.mdc
- .cursor/rules/skills/container-runtime-security-falco.mdc
- .cursor/rules/skills/database-devops-lifecycle.mdc
- .cursor/rules/skills/detection-engineering-threat-hunting.mdc
- .cursor/rules/skills/devops-metrics-dora-kpis.mdc
- .cursor/rules/skills/docker-containerization-basics.mdc
- .cursor/rules/skills/enterprise-iac-governance-terragrunt.mdc
- .cursor/rules/skills/finops-framework-inform-optimize-operate.mdc
- .cursor/rules/skills/gcp-cloud-engineering-patterns.mdc
- .cursor/rules/skills/gcp-gke-autopilot-multi-tenant.mdc
- .cursor/rules/skills/git-branching-merge-strategies.mdc
- .cursor/rules/skills/gitops-multi-cluster-argo-flux.mdc
- .cursor/rules/skills/helm-kubernetes-deployment.mdc
- .cursor/rules/skills/incident-management-and-postmortem.mdc
- .cursor/rules/skills/infrastructure-host-monitoring.mdc
- .cursor/rules/skills/internal-developer-portal-backstage.mdc
- .cursor/rules/skills/linux-sysadmin-troubleshooting.mdc
- .cursor/rules/skills/performance-load-testing.mdc
- .cursor/rules/skills/policy-as-code-opa-kyverno.mdc
- .cursor/rules/skills/prometheus-grafana-otel-tracing.mdc
- .cursor/rules/skills/scripting-and-automation.mdc
- .cursor/rules/skills/secops-incident-triage-forensics.mdc
- .cursor/rules/skills/secrets-management-vault-kms.mdc
- .cursor/rules/skills/serverless-event-driven-architecture.mdc
- .cursor/rules/skills/shift-left-security-sast-sca.mdc
- .cursor/rules/skills/sli-slo-error-budget-design.mdc
- .cursor/rules/skills/supply-chain-security-slsa-sigstore.mdc
- .cursor/rules/skills/terraform-iac-modules.mdc
- .cursor/rules/skills/write-a-skill.mdc
- .cursor/rules/skills/zero-downtime-release-strategies.mdc
Skill
- ai-agent-security-llm-threats.agents/skills/ai-agent-security-llm-threats/SKILL.md
- api-gateway-service-mesh.agents/skills/api-gateway-service-mesh/SKILL.md
- aws-cloud-migration-strategies.agents/skills/aws-cloud-migration-strategies/SKILL.md
- aws-eks-enterprise-patterns.agents/skills/aws-eks-enterprise-patterns/SKILL.md
- aws-iam-zero-trust-policies.agents/skills/aws-iam-zero-trust-policies/SKILL.md
- azure-aks-enterprise-landing-zones.agents/skills/azure-aks-enterprise-landing-zones/SKILL.md
- azure-cloud-engineering-patterns.agents/skills/azure-cloud-engineering-patterns/SKILL.md
- backup-and-disaster-recovery.agents/skills/backup-and-disaster-recovery/SKILL.md
- chaos-engineering-resilience-testing.agents/skills/chaos-engineering-resilience-testing/SKILL.md
- cicd-pipeline-design.agents/skills/cicd-pipeline-design/SKILL.md
- cloud-native-microservices-patterns.agents/skills/cloud-native-microservices-patterns/SKILL.md
- cloud-security-posture-cspm-cis.agents/skills/cloud-security-posture-cspm-cis/SKILL.md
- configuration-management-ansible.agents/skills/configuration-management-ansible/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

