cloud-platform-skills / skills
mchittineni/cloud-platform-skills/.cursor/rules/skills/helm-kubernetes-deployment.mdc
Helm chart engineering: chart/subchart layout, values schema validation, hardened Deployment templates (probes, resources, securityContext, PDB), and release lifecycle with rollback. Use when authoring or reviewing a Helm chart, templating Kubernetes manifests, or debugging a failed or stuck Helm release.
Cursor rule1 starsChanged 45 days ago
What's in it
- Production Helm Chart Engineering & Kubernetes Packaging
- When to Use This Skill
- 1. Production Chart Architecture
- 2. Hardened Deployment Manifest Template (deployment.yaml)
- 3. Deployment Safety Rules
- 4. Probe Separation & Release Lifecycle
- Release lifecycle
---
description: "Helm chart engineering: chart/subchart layout, values schema validation, hardened Deployment templates (probes, resources, securityContext, PDB), and release lifecycle with rollback. Use when authoring or reviewing a Helm chart, templating Kubernetes manifests, or debugging a failed or stuck Helm release."
globs:
alwaysApply: false
---
# Production Helm Chart Engineering & Kubernetes Packaging
## When to Use This Skill
**Triggers — load this skill when:**
- A service needs a production chart with probes, limits, and securityContext set correctly
- Chart values need schema validation or environment layering
- A `helm upgrade` failed, hung, or must be rolled back
**Route elsewhere when:**
- Progressive canary/blue-green rollout control -> `zero-downtime-release-strategies`
- Continuous reconciliation of charts across clusters -> `gitops-multi-cluster-argo-flux`
- Managed-cluster/node-pool design -> `aws-eks-enterprise-patterns`, `azure-aks-enterprise-landing-zones`, `gcp-gke-autopilot-multi-tenant`
## 1. Production Chart Architecture
```text
chart/
├── Chart.yaml
├── values.yaml
├── values.schema.json # JSON Schema validation for inputs
└── templates/
├── _helpers.tpl # Standardized label macros
├── deployment.yaml
├── service.yaml
├── hpa.yaml
└── pdb.yaml
```
---
## 2. Hardened Deployment Manifest Template (`deployment.yaml`)
```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: {{ include "app.fullname" . }}
labels:
{{- include "app.labels" . | nindent 4 }}
spec:
replicas: {{ .Values.replicaCount }}
selector:
matchLabels:
{{- include "app.selectorLabels" . | nindent 6 }}
template:
metadata:
labels:
{{- include "app.selectorLabels" . | nindent 8 }}
spec:
securityContext:
runAsNonRoot: true
runAsUser: 10001
fsGroup: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: {{ .Chart.Name }}
image: "{{ .Values.image.repository }}:{{ .Values.image.tag | default .Chart.AppVersion }}"
imagePullPolicy: {{ .Values.image.pullPolicy }}
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
ports:
- name: http
containerPort: {{ .Values.service.port }}
protocol: TCP
livenessProbe:
httpGet:
path: /healthz
port: http
initialDelaySeconds: 15
periodSeconds: 10
readinessProbe:
httpGet:
path: /ready
port: http
initialDelaySeconds: 5
periodSeconds: 5
resources:
{{- toYaml .Values.resources | nindent 12 }}
```
---
## 3. Deployment Safety Rules
- **Pod Disruption Budgets (PDB)**: Always define a `PodDisruptionBudget` for production workloads (`minAvailable: 1` or `maxUnavailable: 25%`).
- **Resource Requests & Limits**: Never omit CPU and memory requests and limits; prevent noisy neighbor starvation.
- **Topology Spread Constraints**: Distribute pods evenly across availability zones to withstand cloud node/zone outages.
---
## 4. Probe Separation & Release Lifecycle
**Three probes, three different questions.** Reusing one endpoint for all three is the most
common cause of restart loops on slow-starting services:
```yaml
startupProbe: # "has it finished booting?" — buys slow starters time
httpGet: { path: /healthz, port: http }
failureThreshold: 30
periodSeconds: 5 # up to 150s to start; liveness stays disabled until this passes
livenessProbe: # "is it wedged and in need of a restart?" — cheap, no dependencies
httpGet: { path: /healthz, port: http }
periodSeconds: 10
failureThreshold: 3
readinessProbe: # "should it receive traffic right now?" — may check dependencies
httpGet: { path: /readyz, port: http }
periodSeconds: 5
failureThreshold: 2
```
A liveness probe that checks the database restarts every replica during a database blip. Keep
dependency checks in readiness only.
### Release lifecycle
```bash
helm upgrade --install api ./chart -f values.prod.yaml --atomic --timeout 5m --wait # --atomic auto-rolls back a failed upgrade
helm history api # revision, status, chart/app version
helm rollback api 7 --wait # deterministic return to a known-good revision
helm get values api --revision 7 # what was actually deployed then
```
Stuck in `pending-upgrade` usually means a previous run was killed mid-flight: inspect
`helm history`, then `helm rollback` to the last `deployed` revision rather than deleting the
release secret by hand.
More agent context in mchittineni/cloud-platform-skills
167 other files this repository gives its agents, the first 60 shown.
AGENTS.md
CLAUDE.md
Copilot instructions
Cursor rule
- .cursor/rules/00-index.mdc
- .cursor/rules/skills/ai-agent-security-llm-threats.mdc
- .cursor/rules/skills/api-gateway-service-mesh.mdc
- .cursor/rules/skills/aws-cloud-migration-strategies.mdc
- .cursor/rules/skills/aws-eks-enterprise-patterns.mdc
- .cursor/rules/skills/aws-iam-zero-trust-policies.mdc
- .cursor/rules/skills/azure-aks-enterprise-landing-zones.mdc
- .cursor/rules/skills/azure-cloud-engineering-patterns.mdc
- .cursor/rules/skills/backup-and-disaster-recovery.mdc
- .cursor/rules/skills/chaos-engineering-resilience-testing.mdc
- .cursor/rules/skills/cicd-pipeline-design.mdc
- .cursor/rules/skills/cloud-native-microservices-patterns.mdc
- .cursor/rules/skills/cloud-security-posture-cspm-cis.mdc
- .cursor/rules/skills/configuration-management-ansible.mdc
- .cursor/rules/skills/container-runtime-security-falco.mdc
- .cursor/rules/skills/database-devops-lifecycle.mdc
- .cursor/rules/skills/detection-engineering-threat-hunting.mdc
- .cursor/rules/skills/devops-metrics-dora-kpis.mdc
- .cursor/rules/skills/docker-containerization-basics.mdc
- .cursor/rules/skills/enterprise-iac-governance-terragrunt.mdc
- .cursor/rules/skills/finops-framework-inform-optimize-operate.mdc
- .cursor/rules/skills/gcp-cloud-engineering-patterns.mdc
- .cursor/rules/skills/gcp-gke-autopilot-multi-tenant.mdc
- .cursor/rules/skills/git-branching-merge-strategies.mdc
- .cursor/rules/skills/gitops-multi-cluster-argo-flux.mdc
- .cursor/rules/skills/incident-management-and-postmortem.mdc
- .cursor/rules/skills/infrastructure-host-monitoring.mdc
- .cursor/rules/skills/internal-developer-portal-backstage.mdc
- .cursor/rules/skills/linux-sysadmin-troubleshooting.mdc
- .cursor/rules/skills/performance-load-testing.mdc
- .cursor/rules/skills/policy-as-code-opa-kyverno.mdc
- .cursor/rules/skills/prometheus-grafana-otel-tracing.mdc
- .cursor/rules/skills/scalability-high-availability-patterns.mdc
- .cursor/rules/skills/scripting-and-automation.mdc
- .cursor/rules/skills/secops-incident-triage-forensics.mdc
- .cursor/rules/skills/secrets-management-vault-kms.mdc
- .cursor/rules/skills/serverless-event-driven-architecture.mdc
- .cursor/rules/skills/shift-left-security-sast-sca.mdc
- .cursor/rules/skills/sli-slo-error-budget-design.mdc
- .cursor/rules/skills/supply-chain-security-slsa-sigstore.mdc
- .cursor/rules/skills/terraform-iac-modules.mdc
- .cursor/rules/skills/write-a-skill.mdc
- .cursor/rules/skills/zero-downtime-release-strategies.mdc
Skill
- ai-agent-security-llm-threats.agents/skills/ai-agent-security-llm-threats/SKILL.md
- api-gateway-service-mesh.agents/skills/api-gateway-service-mesh/SKILL.md
- aws-cloud-migration-strategies.agents/skills/aws-cloud-migration-strategies/SKILL.md
- aws-eks-enterprise-patterns.agents/skills/aws-eks-enterprise-patterns/SKILL.md
- aws-iam-zero-trust-policies.agents/skills/aws-iam-zero-trust-policies/SKILL.md
- azure-aks-enterprise-landing-zones.agents/skills/azure-aks-enterprise-landing-zones/SKILL.md
- azure-cloud-engineering-patterns.agents/skills/azure-cloud-engineering-patterns/SKILL.md
- backup-and-disaster-recovery.agents/skills/backup-and-disaster-recovery/SKILL.md
- chaos-engineering-resilience-testing.agents/skills/chaos-engineering-resilience-testing/SKILL.md
- cicd-pipeline-design.agents/skills/cicd-pipeline-design/SKILL.md
- cloud-native-microservices-patterns.agents/skills/cloud-native-microservices-patterns/SKILL.md
- cloud-security-posture-cspm-cis.agents/skills/cloud-security-posture-cspm-cis/SKILL.md
- configuration-management-ansible.agents/skills/configuration-management-ansible/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

