cloud-platform-skills / skills
mchittineni/cloud-platform-skills/.cursor/rules/skills/linux-sysadmin-troubleshooting.mdc
Linux host troubleshooting with the USE method (Utilization, Saturation, Errors): high load average, memory pressure and OOM kills, disk and inode exhaustion, I/O wait, processes hung in D state, socket and DNS failures. Use when a server or VM is degraded, crawling, or unresponsive and needs live hands-on diagnosis, when writes fail with 'No space left on device' despite free space, or when a stuck process must be traced.
What's in it
- Linux SysAdmin & Production Troubleshooting Guide
- When to Use This Skill
- 1. Diagnostic Decision Tree
- 2. Standard Diagnostic Commands
- CPU & Process Inspection
- Memory & OOM Diagnostics
- Disk & Storage Troubleshooting
- Networking & Socket State
- 3. Best Practices & Anti-Patterns
---
description: "Linux host troubleshooting with the USE method (Utilization, Saturation, Errors): high load average, memory pressure and OOM kills, disk and inode exhaustion, I/O wait, processes hung in D state, socket and DNS failures. Use when a server or VM is degraded, crawling, or unresponsive and needs live hands-on diagnosis, when writes fail with 'No space left on device' despite free space, or when a stuck process must be traced."
globs:
alwaysApply: false
---
# Linux SysAdmin & Production Troubleshooting Guide
## When to Use This Skill
**Triggers — load this skill when:**
- A node shows high load, memory pressure, disk-full, or I/O wait and you need a systematic first-pass diagnosis
- You need the exact command sequence (top/ps/strace/vmstat/iostat/ss/dig) for a live incident
- You are teaching or reviewing baseline Linux operational competence
**Route elsewhere when:**
- Container-level or Kubernetes-scheduling symptoms -> `docker-containerization-basics` or `helm-kubernetes-deployment`
- Fleet-wide metric collection and alert rules -> `infrastructure-host-monitoring`
- Suspected compromise rather than performance fault -> `secops-incident-triage-forensics`
## 1. Diagnostic Decision Tree
When diagnosing an unresponsive or degraded Linux node, follow the **USE Method** (Utilization, Saturation, and Errors) systematically across CPU, Memory, Disk I/O, and Network.
```text
[High Latency / Alert]
|
+----------------+----------------+
| | |
[CPU] [Memory] [Disk / IO]
top / htop free -m / vmstat iostat -xz 1
mpstat -P ALL dmesg | grep oom df -h / df -i
```
---
## 2. Standard Diagnostic Commands
### CPU & Process Inspection
```bash
# 1. Check load average against core count
uptime
nproc
# 2. Top processes sorted by CPU / Memory
top -b -n 1 | head -n 20
ps aux --sort=-%cpu | head -n 10
ps aux --sort=-%mem | head -n 10
# 3. Trace system calls of a stuck process
strace -p <PID> -f -c
```
### Memory & OOM Diagnostics
```bash
# Detailed memory breakdown
free -h --wide
# Check if kernel OOM-killer terminated processes
dmesg -T | grep -i -E "oom|out of memory|killed process"
journalctl -k --grep="Out of memory" -n 50
```
### Disk & Storage Troubleshooting
```bash
# Check filesystem space and Inode exhaustion
df -h
df -i
# Find top 10 space-consuming directories
du -ahx /var/log 2>/dev/null | sort -rh | head -n 10
# Disk I/O utilization & wait times (%util, await)
iostat -xz 1 5
```
### Networking & Socket State
```bash
# Check listening ports and active sockets
ss -tulpn
# Inspect socket queue backlog
ss -s
# Test connection, latency, and DNS resolution
curl -Iv https://example.internal
dig +trace +short api.service.internal
nc -zv 10.0.1.50 443
```
---
## 3. Best Practices & Anti-Patterns
- **Do**: Always inspect Inodes (`df -i`) if `df -h` shows disk space available but writes fail with `No space left on device`.
- **Do**: Look at CPU `%steal` in virtualized cloud environments (AWS EC2 / GCP Compute Engine) to identify noisy neighbors.
- **Don't**: Never use `kill -9` (`SIGKILL`) immediately; attempt graceful `kill -15` (`SIGTERM`) first to allow socket closures and data flushing.
- **Don't**: Avoid running heavy `find /` commands during peak traffic without `-xdev` to avoid crossing network mounts.
More agent context in mchittineni/cloud-platform-skills
167 other files this repository gives its agents, the first 60 shown.
AGENTS.md
CLAUDE.md
Copilot instructions
Cursor rule
- .cursor/rules/00-index.mdc
- .cursor/rules/skills/ai-agent-security-llm-threats.mdc
- .cursor/rules/skills/api-gateway-service-mesh.mdc
- .cursor/rules/skills/aws-cloud-migration-strategies.mdc
- .cursor/rules/skills/aws-eks-enterprise-patterns.mdc
- .cursor/rules/skills/aws-iam-zero-trust-policies.mdc
- .cursor/rules/skills/azure-aks-enterprise-landing-zones.mdc
- .cursor/rules/skills/azure-cloud-engineering-patterns.mdc
- .cursor/rules/skills/backup-and-disaster-recovery.mdc
- .cursor/rules/skills/chaos-engineering-resilience-testing.mdc
- .cursor/rules/skills/cicd-pipeline-design.mdc
- .cursor/rules/skills/cloud-native-microservices-patterns.mdc
- .cursor/rules/skills/cloud-security-posture-cspm-cis.mdc
- .cursor/rules/skills/configuration-management-ansible.mdc
- .cursor/rules/skills/container-runtime-security-falco.mdc
- .cursor/rules/skills/database-devops-lifecycle.mdc
- .cursor/rules/skills/detection-engineering-threat-hunting.mdc
- .cursor/rules/skills/devops-metrics-dora-kpis.mdc
- .cursor/rules/skills/docker-containerization-basics.mdc
- .cursor/rules/skills/enterprise-iac-governance-terragrunt.mdc
- .cursor/rules/skills/finops-framework-inform-optimize-operate.mdc
- .cursor/rules/skills/gcp-cloud-engineering-patterns.mdc
- .cursor/rules/skills/gcp-gke-autopilot-multi-tenant.mdc
- .cursor/rules/skills/git-branching-merge-strategies.mdc
- .cursor/rules/skills/gitops-multi-cluster-argo-flux.mdc
- .cursor/rules/skills/helm-kubernetes-deployment.mdc
- .cursor/rules/skills/incident-management-and-postmortem.mdc
- .cursor/rules/skills/infrastructure-host-monitoring.mdc
- .cursor/rules/skills/internal-developer-portal-backstage.mdc
- .cursor/rules/skills/performance-load-testing.mdc
- .cursor/rules/skills/policy-as-code-opa-kyverno.mdc
- .cursor/rules/skills/prometheus-grafana-otel-tracing.mdc
- .cursor/rules/skills/scalability-high-availability-patterns.mdc
- .cursor/rules/skills/scripting-and-automation.mdc
- .cursor/rules/skills/secops-incident-triage-forensics.mdc
- .cursor/rules/skills/secrets-management-vault-kms.mdc
- .cursor/rules/skills/serverless-event-driven-architecture.mdc
- .cursor/rules/skills/shift-left-security-sast-sca.mdc
- .cursor/rules/skills/sli-slo-error-budget-design.mdc
- .cursor/rules/skills/supply-chain-security-slsa-sigstore.mdc
- .cursor/rules/skills/terraform-iac-modules.mdc
- .cursor/rules/skills/write-a-skill.mdc
- .cursor/rules/skills/zero-downtime-release-strategies.mdc
Skill
- ai-agent-security-llm-threats.agents/skills/ai-agent-security-llm-threats/SKILL.md
- api-gateway-service-mesh.agents/skills/api-gateway-service-mesh/SKILL.md
- aws-cloud-migration-strategies.agents/skills/aws-cloud-migration-strategies/SKILL.md
- aws-eks-enterprise-patterns.agents/skills/aws-eks-enterprise-patterns/SKILL.md
- aws-iam-zero-trust-policies.agents/skills/aws-iam-zero-trust-policies/SKILL.md
- azure-aks-enterprise-landing-zones.agents/skills/azure-aks-enterprise-landing-zones/SKILL.md
- azure-cloud-engineering-patterns.agents/skills/azure-cloud-engineering-patterns/SKILL.md
- backup-and-disaster-recovery.agents/skills/backup-and-disaster-recovery/SKILL.md
- chaos-engineering-resilience-testing.agents/skills/chaos-engineering-resilience-testing/SKILL.md
- cicd-pipeline-design.agents/skills/cicd-pipeline-design/SKILL.md
- cloud-native-microservices-patterns.agents/skills/cloud-native-microservices-patterns/SKILL.md
- cloud-security-posture-cspm-cis.agents/skills/cloud-security-posture-cspm-cis/SKILL.md
- configuration-management-ansible.agents/skills/configuration-management-ansible/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

