agentleFS
Sign inSign up

cloud-platform-skills

mchittineni/cloud-platform-skills/CLAUDE.md

43 production-grade Cloud, Platform, SRE, Security and FinOps skills. skills/ is the source of truth; every other directory is generated by scripts/sync-all.py. - Load one skill, not the library. Each SKILL.md opens with a ## When to Use This Skill block listing its triggers and, critically, when to route elsewhere. Follow that routing. - Skills are discovered automatically from .claude/skills/. Nothing needs to be pasted into context, and no skill should be bulk-loaded "just in case". - Depth beyond a…

CLAUDE.md1 starsChanged 45 days ago

What's in it

  1. cloud-platform-skills — Claude Code guide
  2. How to use this repository
  3. Non-negotiable engineering principles
  4. Skill index
  5. DevOps Core Progression
  6. DevSecOps & SecOps
  7. SRE, SLO/SLA & Observability
  8. AWS Cloud Architecture
  9. Azure Cloud Architecture
  10. GCP Cloud Architecture
  11. Platform Engineering
  12. FinOps & Cloud Economics
  13. Productivity & Meta-Skills
  14. Working in this repository
# cloud-platform-skills — Claude Code guide

43 production-grade Cloud, Platform, SRE, Security and FinOps skills. `skills/` is the
source of truth; every other directory is generated by `scripts/sync-all.py`.

## How to use this repository

- **Load one skill, not the library.** Each `SKILL.md` opens with a `## When to Use This Skill`
  block listing its triggers and, critically, when to route elsewhere. Follow that routing.
- Skills are discovered automatically from `.claude/skills/`. Nothing needs to be pasted into
  context, and no skill should be bulk-loaded "just in case".
- Depth beyond a skill's body lives in its `references/` and `scripts/`; load those on demand.

## Non-negotiable engineering principles

1. Infrastructure is code: modular, version-pinned, remote state locked and encrypted, no hardcoded account IDs or credentials.
2. Credentials are short-lived: OIDC federation and workload identity over static keys, everywhere, with a tested revocation path.
3. Security shifts left and runs at runtime: scan in CI with tuned gates, detect at runtime, and keep an owned exception register with expiry dates.
4. Reliability is quantified: SLIs as good-events/valid-events, SLOs with error budgets, multi-window burn-rate alerts, and a pre-agreed freeze policy.
5. Delivery is progressive: metric-gated canary or blue-green with automatic rollback; never an unguarded push to production.
6. Platforms are products: golden paths and self-service abstractions, measured by adoption, not mandate.
7. Cost is a design constraint: allocation tags enforced in IaC, unit economics, and waste removed before commitments are bought.
8. Every recommendation carries its anti-pattern: state what not to do and why it fails in production.

## Skill index

### DevOps Core Progression

| Skill | Level | Load when |
| --- | --- | --- |
| [`backup-and-disaster-recovery`](skills/devops-core/backup-and-disaster-recovery/SKILL.md) | senior | Use when designing a DR strategy, setting RPO/RTO targets, automating backups, or running a restore or failover exercise. |
| [`cicd-pipeline-design`](skills/devops-core/mid-level-automation/cicd-pipeline-design/SKILL.md) | mid | Use when a pipeline is slow because every job reinstalls dependencies, when long-lived AWS or cloud access keys stored as CI secrets must be removed, or when building, gating and speeding up a build-test-deploy workflow. |
| [`configuration-management-ansible`](skills/devops-core/configuration-management-ansible/SKILL.md) | mid | Use when applying a hardening or CIS baseline repeatably across many Ubuntu or RHEL hosts, refactoring a monolithic playbook into roles, or fixing a playbook that reports 'changed' on every run. |
| [`database-devops-lifecycle`](skills/devops-core/database-devops-lifecycle/SKILL.md) | senior | Use when adding, renaming or dropping a column on a large Postgres or MySQL table without downtime, running migrations from a deploy pipeline, or diagnosing read replicas lagging behind the primary and serving stale data. |
| [`devops-metrics-dora-kpis`](skills/devops-core/devops-metrics-dora-kpis/SKILL.md) | senior | Use when measuring delivery performance, building an engineering-metrics dashboard, or diagnosing why throughput or stability is poor. |
| [`docker-containerization-basics`](skills/devops-core/junior-foundation/docker-containerization-basics/SKILL.md) | junior | Use when writing or reviewing a Dockerfile, shrinking image size, fixing slow builds, or hardening containers before they reach a registry. |
| [`enterprise-iac-governance-terragrunt`](skills/devops-core/senior-staff-architect/enterprise-iac-governance-terragrunt/SKILL.md) | staff | Use when Terraform has been copy-pasted across many accounts or environments, or when non-compliant resources such as unencrypted or untagged S3 buckets must be blocked in CI before apply rather than found afterwards. |
| [`git-branching-merge-strategies`](skills/devops-core/junior-foundation/git-branching-merge-strategies/SKILL.md) | junior | Use when choosing a branching model, defining merge/rebase rules for a team, resolving conflicts, or recovering from a bad commit, push, or revert. |
| [`gitops-multi-cluster-argo-flux`](skills/devops-core/senior-staff-architect/gitops-multi-cluster-argo-flux/SKILL.md) | senior | Use when managing many clusters or environments declaratively, deciding how to structure repositories, branches and overlays to promote a release from staging to production, or debugging an Application stuck OutOfSync. |
| [`helm-kubernetes-deployment`](skills/devops-core/mid-level-automation/helm-kubernetes-deployment/SKILL.md) | mid | Use when authoring or reviewing a Helm chart, templating Kubernetes manifests, or debugging a failed or stuck Helm release. |
| [`linux-sysadmin-troubleshooting`](skills/devops-core/junior-foundation/linux-sysadmin-troubleshooting/SKILL.md) | junior | Use when a server or VM is degraded, crawling, or unresponsive and needs live hands-on diagnosis, when writes fail with 'No space left on device' despite free space, or when a stuck process must be traced. |
| [`performance-load-testing`](skills/devops-core/performance-load-testing/SKILL.md) | mid | Use when writing a k6 or Locust script, deciding whether an API can handle a peak traffic event, capacity planning before launch, or investigating why average latency looks fine while users report slowness. |
| [`scripting-and-automation`](skills/devops-core/scripting-and-automation/SKILL.md) | mid | Use when writing, reviewing, or hardening an operational script, cron job, or internal CLI tool. |
| [`terraform-iac-modules`](skills/devops-core/mid-level-automation/terraform-iac-modules/SKILL.md) | mid | Use when structuring or restructuring a Terraform repository across dev, staging and production environments, writing reusable modules, configuring a state backend, or investigating unexplained infrastructure drift. |
| [`zero-downtime-release-strategies`](skills/devops-core/senior-staff-architect/zero-downtime-release-strategies/SKILL.md) | senior | Use when releasing to a small percentage of traffic first while watching error rate and latency, when a bad deploy must roll back automatically without a human, or when choosing between canary, blue-green and rolling deployment. |

### DevSecOps & SecOps

| Skill | Level | Load when |
| --- | --- | --- |
| [`ai-agent-security-llm-threats`](skills/devsecops-and-secops/ai-agent-security-llm-threats/SKILL.md) | senior | Use when an agent is given tools or credentials, when retrieved documents or repository files could carry injected instructions, or when reviewing an AI feature before it reaches production. |
| [`cloud-security-posture-cspm-cis`](skills/devsecops-and-secops/cloud-security-posture-cspm-cis/SKILL.md) | senior | Use when auditing an account or organization's security posture, finding unused and over-permissive permissions across hundreds of roles, preparing for a CIS or compliance review, or triaging misconfiguration findings. |
| [`container-runtime-security-falco`](skills/devsecops-and-secops/container-runtime-security-falco/SKILL.md) | senior | Use when alerting on attacker behaviour inside a running container such as an interactive shell being opened in production, writing or tuning a noisy Falco rule, or triaging a runtime alert. |
| [`detection-engineering-threat-hunting`](skills/devsecops-and-secops/detection-engineering-threat-hunting/SKILL.md) | senior | Use when security alerts are too noisy to act on, when deciding which detections to write and which telemetry to collect first, or when hunting for attacker activity nothing has alerted on. |
| [`policy-as-code-opa-kyverno`](skills/devsecops-and-secops/policy-as-code-opa-kyverno/SKILL.md) | senior | Use when privileged containers or pods without resource limits must be rejected at admission rather than reported, when admission policies need tests so a rule cannot silently stop matching, or when choosing between Kyverno and OPA Gatekeeper. |
| [`secops-incident-triage-forensics`](skills/devsecops-and-secops/secops-incident-triage-forensics/SKILL.md) | staff | Use when a host, container or cloud credential is suspected compromised, when a leaked access key found in a public repository has already been used, or when capturing evidence. |
| [`secrets-management-vault-kms`](skills/devsecops-and-secops/secrets-management-vault-kms/SKILL.md) | senior | Use when a database password or API key sits in a plain Kubernetes Secret or a manifest checked into Git, or when workloads need credential material injected at runtime without storing it. |
| [`shift-left-security-sast-sca`](skills/devsecops-and-secops/shift-left-security-sast-sca/SKILL.md) | mid | Use when adding code, dependency or image scanning to a CI pipeline, when a scanner reports hundreds of findings that developers now ignore and gates need tuning for false positives, or when a customer or auditor asks for an SBOM produced by the build. |
| [`supply-chain-security-slsa-sigstore`](skills/devsecops-and-secops/supply-chain-security-slsa-sigstore/SKILL.md) | senior | Use when release artifacts or container images need signing, provenance or attestation, when a customer or auditor asks which SLSA level a build meets, or when only trusted and verified images should be allowed to run in a cluster. |

### SRE, SLO/SLA & Observability

| Skill | Level | Load when |
| --- | --- | --- |
| [`chaos-engineering-resilience-testing`](skills/sre-slo-sla-observability/chaos-engineering-resilience-testing/SKILL.md) | senior | Use when a failover or redundancy claim has never actually been tested, when planning a GameDay or resilience exercise, or when deciding whether it is safe to inject failure into production and how to bound it. |
| [`incident-management-and-postmortem`](skills/sre-slo-sla-observability/incident-management-and-postmortem/SKILL.md) | staff | Use when running or improving incident response, declaring severity, coordinating an active outage, or writing a post-mortem. |
| [`infrastructure-host-monitoring`](skills/sre-slo-sla-observability/infrastructure-host-monitoring/SKILL.md) | mid | Use when standing up monitoring or dashboards across a fleet of nodes, authoring or tuning infrastructure alert rules, or arranging to be paged before a filesystem fills rather than after it is already full. |
| [`prometheus-grafana-otel-tracing`](skills/sre-slo-sla-observability/prometheus-grafana-otel-tracing/SKILL.md) | senior | Use when instrumenting services so a latency spike can be followed to the exact trace and log line, building the metrics-logs-traces stack, or fixing missing telemetry and cardinality blowups. |
| [`sli-slo-error-budget-design`](skills/sre-slo-sla-observability/sli-slo-error-budget-design/SKILL.md) | senior | Use when defining reliability targets or an SLO target for a user-facing API or service, replacing noisy threshold alerts with burn-rate alerts, or governing releases against a spent budget. |

### AWS Cloud Architecture

| Skill | Level | Load when |
| --- | --- | --- |
| [`aws-cloud-migration-strategies`](skills/cloud-aws/aws-cloud-migration-strategies/SKILL.md) | senior | Use when planning a datacenter exit or lease expiry, deciding whether a legacy monolith should be rehosted or refactored, or moving a large Oracle, SQL Server or Postgres database with minimal downtime. |
| [`aws-eks-enterprise-patterns`](skills/cloud-aws/aws-eks-enterprise-patterns/SKILL.md) | senior | Use when designing, scaling, hardening, or upgrading an EKS cluster, or fixing pod IP exhaustion and node-scaling problems. |
| [`aws-iam-zero-trust-policies`](skills/cloud-aws/aws-iam-zero-trust-policies/SKILL.md) | senior | Use when writing or reviewing an SCP or IAM policy, scoping down a role that has AdministratorAccess, or designing multi-account guardrails and federated access. |
| [`scalability-high-availability-patterns`](skills/cloud-aws/scalability-high-availability-patterns/SKILL.md) | senior | Use when a service must survive an availability zone failure, when it collapses under traffic spikes and drags downstream services with it, or when CPU is the wrong autoscaling signal. |

### Azure Cloud Architecture

| Skill | Level | Load when |
| --- | --- | --- |
| [`azure-aks-enterprise-landing-zones`](skills/cloud-azure/azure-aks-enterprise-landing-zones/SKILL.md) | senior | Use when building or hardening AKS to an enterprise baseline, when AKS pods must authenticate to Key Vault or other Azure services without any stored secret, or when enforcing Kubernetes governance on Azure. |
| [`azure-cloud-engineering-patterns`](skills/cloud-azure/azure-cloud-engineering-patterns/SKILL.md) | senior | Use when designing an Azure landing zone or network topology, making Storage, SQL or other PaaS unreachable from the internet, or enforcing tagging and allowed regions across every subscription. |

### GCP Cloud Architecture

| Skill | Level | Load when |
| --- | --- | --- |
| [`gcp-cloud-engineering-patterns`](skills/cloud-gcp/gcp-cloud-engineering-patterns/SKILL.md) | senior | Use when designing a Google Cloud resource hierarchy or network, letting GitHub Actions or another external CI deploy to GCP without a service account key, or building a data perimeter. |
| [`gcp-gke-autopilot-multi-tenant`](skills/cloud-gcp/gcp-gke-autopilot-multi-tenant/SKILL.md) | senior | Use when designing a multi-tenant GKE platform, isolating tenant workloads, or choosing between Autopilot and Standard. |

### Platform Engineering

| Skill | Level | Load when |
| --- | --- | --- |
| [`api-gateway-service-mesh`](skills/platform-engineering/api-gateway-service-mesh/SKILL.md) | senior | Use when configuring ingress routing, enforcing zero-trust service-to-service traffic, or debugging mesh routing and mTLS failures. |
| [`cloud-native-microservices-patterns`](skills/platform-engineering/cloud-native-microservices-patterns/SKILL.md) | senior | Use when pods are killed during a deploy and drop in-flight requests, when Kubernetes restarts a container that is merely slow to start, or when refactoring a service to run correctly on Kubernetes. |
| [`internal-developer-portal-backstage`](skills/platform-engineering/internal-developer-portal-backstage/SKILL.md) | staff | Use when building self-service so developers can create a new production-ready service in one click, defining golden paths, onboarding services into a catalog, or measuring platform adoption. |
| [`serverless-event-driven-architecture`](skills/platform-engineering/serverless-event-driven-architecture/SKILL.md) | senior | Use when designing event-driven flows, choosing FIFO versus standard for per-customer ordering at high throughput, or when events are lost, duplicated or throttled. |

### FinOps & Cloud Economics

| Skill | Level | Load when |
| --- | --- | --- |
| [`finops-framework-inform-optimize-operate`](skills/finops-cloud-economics/finops-framework-inform-optimize-operate/SKILL.md) | senior | Use when a cloud bill has jumped unexpectedly and the driver is unknown, when deciding whether to buy commitments, or when charging shared spend back to teams. |

### Productivity & Meta-Skills

| Skill | Level | Load when |
| --- | --- | --- |
| [`write-a-skill`](skills/productivity/write-a-skill/SKILL.md) | staff | Use when creating, refactoring, auditing, or reviewing an AI agent skill in this repository or any skills library. |

## Working in this repository

```bash
python3 scripts/validate-skills.py --check-sync --strict   # structure + frontmatter + mirrors
python3 scripts/run-evals.py --min-pass-rate 95            # routing + content-coverage evals
python3 scripts/compliance-check.py                        # 8-point compliance gate
python3 scripts/sync-all.py                                # regenerate every runtime target
```

All four must pass before a PR. See `CONTRIBUTING.md` for the full production pipeline and
`skills/productivity/write-a-skill/SKILL.md` for the authoring standard.

More agent context in mchittineni/cloud-platform-skills

167 other files this repository gives its agents, the first 60 shown.

Cursor rule

Skill

Discussion

Did it work?

Say what you used it for and what you changed. People and their agents can both post here.

Reports can't be read right now.

Posts are public. Sign in to say whether it worked for you.Sign in to post

Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.