sre-agent / zava-aks-postgres
microsoft/sre-agent/labs/zava-aks-postgres/AGENTS.md
Guidance for any AI agent (GitHub Copilot CLI/IDE, Claude Code, Cursor, Codex, etc.) working in this repo. Azure SRE Agent demo — AKS + PostgreSQL e-commerce app with break/fix scenarios and autonomous incident response. When a user clones this repo, guide them through setup: 1. Check prerequisites: az, azd, pwsh, plus Owner/User Access Administrator (or equivalent role-assignment write permission) at subscription scope — install anything missing. (kubectl is not required on the operator workstation: the AKS cluster is private, operator…
# AGENTS.md — SRE Zava Demo
Guidance for any AI agent (GitHub Copilot CLI/IDE, Claude Code, Cursor, Codex, etc.) working in this repo.
Azure SRE Agent demo — AKS + PostgreSQL e-commerce app with break/fix scenarios and autonomous incident response.
## Getting Started
When a user clones this repo, guide them through setup:
1. Check prerequisites: `az`, `azd`, `pwsh`, plus Owner/User Access Administrator (or equivalent role-assignment write permission) at subscription scope — install anything missing. (`kubectl` is **not** required on the operator workstation: the AKS cluster is private, operator operations use `az aks command invoke`, and the SRE Agent uses its built-in Kubernetes tools.)
2. Run `azd up` — pick their default subscription, use `swedencentral` region
and record the responsible owner and expiry date using the subscription's
required tag keys. If deployment policy does not apply them, tag the resource
group immediately after deployment.
3. `azd up` runs `scripts/setup-sre-agent.ps1` to wait for API readiness, deploy the staged connector template, then apply skills, guarded custom agents, response plans, knowledge, global instructions, and tool settings. Run it manually only to retry or apply later configuration changes. Existing managed drift stops before configuration writes; review it before explicitly using `-UpdateExisting`. `-ResourceGroup rg-example -RenderOnly` previews the runtime configuration offline.
4. Open the storefront in browser to verify it works
5. Run a break scenario to demonstrate the SRE Agent
Use browser tools (Playwright MCP or Chrome DevTools) to visually verify the storefront after deployment.
## Non-Obvious Things (repo / IaC author concerns)
These are gotchas for someone editing this repo's IaC or Bicep — they're *not* runtime-incident knowledge. The runtime knowledge the SRE Agent itself needs (private PostgreSQL reachability, native tool operation boundaries, ECONNREFUSED vs ETIMEDOUT, NSG-vs-NetworkPolicy, App Insights filter requirements, etc.) lives in [`sre-config/knowledge-base/zava-architecture.md`](sre-config/knowledge-base/zava-architecture.md), which gets uploaded to the agent on deploy.
- **K8s manifests** use `${VAR}` placeholders — substituted by `post-provision.ps1`, not Helm/Kustomize.
- **No `azd deploy api`** — there's no `services:` block in `azure.yaml`. Container images are built by `post-provision.ps1` via `az acr build`. To iterate: `az acr build --registry $acr --image zava-api:latest ./src/api` then `az aks command invoke … kubectl rollout restart deployment …`.
- **Activity log alerts default to Sev4** because `Microsoft.Insights/activityLogAlerts` rules don't expose a severity field for the categories we use (Administrative). Response plans must include `Sev4` in `priorities[]` (this repo includes all severities) or activity-log-driven incidents won't match.
- **Response plans use three purpose-built routes and one bounded fallback.** `zava-database`, `zava-performance`, and `zava-application` run autonomously; `zava-unknown` handles other `Zava` alerts in Review mode. Keep filters mutually exclusive and explicitly exclude known routes from the fallback.
- **The demo has three enabled dispatching alerts.** `postgres-unreachable` covers database availability, `Zava-products-query-slow` covers query performance, and `Zava-http-5xx-errors` covers application failures. Supporting metrics remain available for investigation even when they do not dispatch a separate alert.
- **Dispatching scheduled-query alerts use `evaluationFrequency: PT5M`.** The deployed and validated contract for `postgres-unreachable`, `Zava-products-query-slow`, and `Zava-http-5xx-errors` is `PT5M` evaluation with a `PT5M` window. Keep new demo dispatching alerts aligned with that configuration.
- **Both database scenarios use `postgres-unreachable`.** Diagnose a stopped server versus a network block from PostgreSQL ARM state and network configuration, not error text alone. After verified recovery, close only the owned alert when the tool supports it; otherwise report blocked closure. Wait for `monitorCondition == Resolved` before the next run, even if the alert state is already Closed.
- **Updating an existing deployment leaves orphans — incremental ARM doesn't delete removed resources.** A fresh `azd up` (new RG) is clean, but applying this on top of a prior deploy keeps the old alerts/filters/skills firing. Delete the retired ones: alerts `Zava-slow-response-time`, `Zava-nsg-change`, `Zava-nsg-rule-deleted`, `postgres-server-stopped`, `postgres-server-down`, `postgres-network-blocked`, and the OLD metric `Zava-http-5xx-errors` (it's now a scheduled query of the same name); incidentFilters `zava-db-response`, `zava-app-response`; skill `db-incident-investigation`.
- **`PT3M` is an invalid `windowSize`** for Azure Monitor metric alerts — only `PT1M, PT5M, PT10M, PT15M, PT30M, PT45M, PT1H+` are accepted. Deployment fails with a misleading error.
- **Scenario 3 is tuned for the demo dataset.** Its seed size, PostgreSQL cost setting, alert threshold, and load generator work together. Review the inline comments in `seed.js`, `logger.js`, `monitoring.bicep`, and the performance runbook before changing them.
- **Scenario 3 drops both category indexes.** `fix-db-perf.ps1` recreates both indexes during cleanup.
- **Scenario 3 restarts the API deployment after dropping the indexes.** This refreshes PostgreSQL query plans and the telemetry exporter before the load run.
- **Scenario 3 includes a telemetry precheck.** `break-db-perf.ps1` verifies recent `zava-api` `AppRequests` data before injecting the fault. Pass `-SkipTelemetryCheck` only when the workspace is intentionally empty.
- **Scenario 3 load runs as a Kubernetes Job.** `break-db-perf.ps1` renders `k8s/jobs/load-categories.yaml`; `fix-db-perf.ps1` removes the Job after restoring the indexes.
- **Scenario 5 (`break-compound.ps1`) is a bounded correlation proof of concept.** It overlaps Scenario 3 and Scenario 4 by 90 seconds so the sample can compare nearby alerts against dependency, deployment, and database telemetry. The causes are intentionally independent. Run the database-performance fault first because it restarts the API deployment; the bad-deploy revision must remain the latest rollout for the application investigation. Do not present this scenario as a comprehensive correlation benchmark or a guaranteed model outcome.
- **Skills are split by incident domain.** Keep database, performance, application, general triage, proactive health, and correlation procedures separate. Domain skills may use the read-only correlation skill when needed.
- **Alert `description` strings in `monitoring.bicep` are agent-readable payload, not cosmetic Bicep strings.** Azure Monitor includes the alert description in the incident context the SRE Agent reads. They MUST stay symptom-only — never re-add "Likely cause: …", "Remediation: …", "(Scenario N)", or any specific resource name (table, index, NetworkPolicy) the agent could pattern-match instead of diagnosing. The descriptions describe what was observed; the runbook + KB explain how to investigate.
- **Correlation guidance is split between global instructions and an on-demand skill.** `sre-config/custom-instructions.md` identifies when a wider review may be useful; `sre-config/skills/incident-correlation.md` contains the read-only procedure. Keep alert descriptions symptom-focused and avoid duplicating correlation instructions across response plans.
- **Correlation requires mechanism evidence.** Alert timestamps identify candidate overlap but not causal order. Establish onset from raw telemetry and compare dependency targets, result codes, deployment history, and PostgreSQL metrics before assigning a shared cause.
- **`Zava-db-cpu-saturation` is disabled by default (`enableDbCpuSaturationAlert=false`).** It supports the Scenario 5 demonstration of checking relevant disabled rules and querying their underlying metrics. Set the parameter to `true` to include the database CPU alert. Metric and scheduled-query rule inventories use different APIs, so the correlation skill checks both.
- **Use Azure Service Health as a correlation source.** Query subscription events with `Microsoft.ResourceHealth/events` and per-resource state with `availabilityStatuses`. The correlation skill uses the list APIs supported by this sample.
- **Keep the knowledge base focused on environment-specific facts.** Include resource names, access paths, and configuration details the investigation cannot discover reliably. Put reusable procedures in skills and link to public Azure documentation instead of duplicating general guidance.
- **The deployed demo intentionally does not connect this source repository.** The break scripts and application source contain the injected faults, so repository indexing would expose the answer key. Keep the README explicit that this is a lab constraint, not production guidance. Production agents should correlate high-quality telemetry and alerts with connected source code, incident access, curated response plans, targeted skills, maintained runbooks, and concise durable custom instructions.
- **The Log Analytics workspace has no `dailyQuotaGb` cap.** Demo alerts must always be able to fire, so workspace ingestion is uncapped. Volume is bounded at the SDK layer instead: `logger.js` only ships warn/error to OTel (HTTP auto-instrumentation already records access logs as `AppRequests`), and high-volume App Insights tables are pinned to the 4-day per-table minimum.
- **The api workload identity needs `Monitoring Metrics Publisher` on the AI resource.** The `@azure/monitor-opentelemetry` SDK uses AAD auth (via the workload-identity-injected federated token) for the breeze ingestion endpoint when `AZURE_CLIENT_ID` is set in the pod env, even if a connection string is also provided. Without the role, exports are silently rejected and no telemetry flows. The role is granted in `identity.bicep`'s `appMetricsPublisher` resource (scoped to the AI component, not the RG) — don't drop it.
- **Scenario 2 uses the NetworkPolicy name `database-tier-isolation`.** Investigations should inspect its egress rules and destination ranges rather than infer behavior from the resource name.
- **Scenario 2 includes both NSG and NetworkPolicy changes.** PostgreSQL Flexible Server private access uses a delegated subnet, so the investigation must determine which control is enforcing the block. Keep the knowledge-base links to the Azure delegation and PostgreSQL networking documentation.
- **No `SendOutlookEmail` tool wired into the skills.** Email-out requires an OAuth consent flow that has to be completed by an interactive user in the agent's portal — there is no Bicep/ARM verb to provision it on the agent's behalf. Adding it to `skill.tools[]` without that consent makes the skill *fail to load*. If you want post-remediation email for your own deployment, follow the public docs ([Microsoft 365 connector for SRE Agent](https://learn.microsoft.com/azure/sre-agent/)) to grant consent in the portal, then re-add `'SendOutlookEmail'` to the relevant skill's `tools` array and a "send a summary email" line to the runbook.
- **The network is hub-and-spoke (three VNets), not one flat VNet.** `vnet.bicep` deploys a **hub** (`vnet-Zava-hub-*`, 10.10.0.0/22 — `AzureFirewallSubnet` + the Azure Firewall, a reserved `GatewaySubnet` for a future ExpressRoute/VPN gateway, and `pe-subnet` for the AMPLS private endpoint), a **platform spoke** (`vnet-Zava-platform-*`, 10.20.0.0/16 — `aks-subnet` + delegated `db-subnet`), and an **agent spoke** (`vnet-Zava-agent-*`, 10.30.0.0/24 — the delegated `agent-subnet` at 10.30.0.0/27). `/27` is the minimum accepted size: 32 total addresses minus Azure's 5 reserved addresses leaves the required 27 usable addresses. The firewall lives in the **hub**; the agent subnet force-tunnels to it via a UDR (`0.0.0.0/0` → firewall private IP `10.10.0.4`) over VNet peering, and every firewall rule takes its source from the single `agentSubnetPrefix` variable — keep that variable at `/27` unless the product requirement changes. Consequence for scripts: anything that picks "the VNet" must select the one containing its target subnet — `break-network.ps1` now queries `[?subnets[?name=='aks-subnet']]`, not `[0]`.
- **Agent VNet injection is REGIONAL; cross-region reach is via PEERING.** The `agent-subnet` (delegated to `Microsoft.App/environments`) **must be in the same region as the `Microsoft.App/agents` resource** — VNet injection is regional, not a tuning knob ([SRE Agent subnet requirements](https://learn.microsoft.com/azure/sre-agent/network-integration#configure-azure-vnet-mode): *"The subnet must be in the same region as your SRE Agent resource"*). A single-region `azd up` satisfies this automatically — `vnet.bicep` deploys all three VNets and `sre-agent.bicep` deploys the agent with the same `location: location`; don't move the agent subnet to another region expecting injection to work. The agent's **reach is NOT regional**, though: peered to the hub, it routes to anything the hub peers to — **other Azure regions over global VNet peering**, **on-prem over the reserved `GatewaySubnet` gateway** (*"as long as your network routes and rules allow it"*). The on-prem example is just one instance of this. `vnet.bicep` ships a **commented `remote-region` global-peering example** (after the local peerings) and the README's *"Reaching other regions and on-premises"* section is the narrative. Because the agent is force-tunneled to the hub firewall, reaching a new peered range also needs a firewall network rule (`agent-subnet → that range`), not just the peering.
- **Agent skills use the built-in Kubernetes system tools.** Use `RunKubectlReadCommand` and `RunKubectlWriteCommand` directly; do not make incident runbooks depend on terminal-native kubectl.
- **Agent PostgreSQL access is native and bounded.** Use `QueryZavaPostgres` for fixed read-only diagnostics and `RepairZavaPostgresIndexes` for its fixed maintenance operations. Keep Kubernetes and `az aks command invoke` for workload operations and operator fault injection; do not route agent database evidence through pod `exec`. This self-contained lab makes the SRE Agent UMI the PostgreSQL Entra administrator, with the fixed-operation allow-list as its trust boundary. Production should use a dedicated least-privilege database principal instead of copying this administrator grant.
- **Sandbox package bootstrap uses the SRE Agent infrastructure network.** Keep `pypi` in `sandboxConfiguration.egress.allowedRegistries` for the pinned `pg8000` and `azure-identity` packages. Do not add PyPI to the customer hub firewall. Keep package-manager access enabled because the platform can rebuild the sandbox base image after workspace or package changes.
- **`pg_stat_statements` has two provisioning gates.** Keep it in the PostgreSQL `azure.extensions` allowlist, then let `post-provision.ps1` create and verify it in `zava_store` after API rollout and before SRE Agent configuration. A failed or ambiguous readback must stop setup.
- **The Microsoft Learn MCP connector uses the hub firewall path.** Keep `allowHttpMcpServerNetworkAccess: false`, allow-list `learn.microsoft.com` and `raw.githubusercontent.com`, and use the `learn-docs` connector name and selected tool IDs declared in Bicep.
- **Agent self-management is explicitly controlled.** `firewall-agent-dataplane.bicep` allows the agent data-plane FQDN only when `allowAgentSelfManagement=true`; otherwise it deploys an empty rule collection to revoke that path.
- **Stage connectors after API readiness.** Keep connector resources out of `sre-agent.bicep` and the core `main.bicep` graph. `setup-sre-agent.ps1` applies `sre-agent-connectors.bicep` from manifest definitions only after an authenticated configuration GET succeeds. Do not replace this gate with fixed sleeps or make the firewall module depend on connector completion. ARM redacts some connector settings; verify exposed fields without claiming complete readback.
- **Keep custom instructions concise and broadly applicable.** `sre-config/custom-instructions.md` identifies when to use the correlation skill and when independent evidence paths justify parallel subagents. Describe the behavior rather than a tool name: use an explicit count and scope, run independent tracks concurrently, wait for all results, verify claims, and synthesize before acting. Keep detailed procedures in skills and keep writes/remediation out of parallel fan-out.
- **The hub Azure Firewall is the demo's "network device".** `firewall-diagnostics.bicep` ships its logs to Log Analytics as resource-specific `AZFW*` tables (`logAnalyticsDestinationType: 'Dedicated'`) so the agent can interrogate it *indirectly* (KQL on `AZFWNetworkRule` / `AZFWApplicationRule` / …) as well as *directly* (ARM reads of its policy/rules). Don't drop the diagnostic setting or the `Dedicated` flag — the KB points the agent at those tables, which only exist in Dedicated mode. The agent already holds Reader/Monitoring Reader on the RG, so no new role is needed for the direct path.
- **Agent AMPLS lockdown is ON by default (`lockAgentToPrivateMonitor = true`).** `monitor-private-link.bicep` always creates the Azure Monitor Private Link Scope, scoped resources (LA + App Insights), the private endpoint, and the five `privatelink.*` DNS zones (linked to the hub). By default it ALSO links those zones to the **agent** spoke, and `vnet.bicep` drops the public `AzureMonitor` service tag from the firewall L4 rule, so the agent reaches Log Analytics / App Insights only over the AMPLS private endpoint (maximum restraint). The agent remains fully functional under it: it queries Log Analytics / App Insights and remediates incidents end-to-end through Monitor and the built-in Kubernetes tools. The Monitor query connector is platform-brokered, so dropping the public `AzureMonitor` tag from the agent-VNet firewall doesn't gate it. Set `lockAgentToPrivateMonitor = false` to keep the public Monitor path. The **platform/workload** spoke is a separate concern: `linkWorkloadVnetsToPrivateMonitor` stays **false** by default because linking it forces the app's App Insights traffic onto the private endpoint — and the regional ingestion host (`<region>-N.in.applicationinsights.azure.com`, from the component's connection string) can resolve into the private zone without a matching record → NXDOMAIN → the app silently stops shipping telemetry (a documented private-link DNS pitfall; this lab doesn't validate the workload's private path). The agent's lockdown is independent (it only queries Monitor, over its own spoke). Don't switch the AMPLS access mode to `PrivateOnly` (resource-level) without testing — that can block operator public queries region-wide.
- **Use a distinct resource group for each demo environment.** The default `rg-$AZURE_ENV_NAME` and resource-name suffix isolate deployments. Create a new environment name after teardown rather than reusing a deleted Log Analytics workspace identity.
- **Teardown includes resources outside the resource group.** Run `azd down --force --purge` from the same environment so the `predown` hook removes the runtime identity's subscription Reader assignment and AMPLS links. Verify the resource group and assignment are gone before considering teardown complete. Keep ownership and expiry tags current until then.
- **Post-provision must fail on required application steps.** Wait for concrete Kubernetes permissions, not just `kubectl version`, before applying manifests. Use the known operator object ID and principal type for role assignment to avoid a redundant directory lookup. Do not ignore failed ingress, manifest, rollout, or endpoint commands.
## Project-local skills (Copilot CLI)
For agents that support Copilot CLI's project-local skills under `.github/skills/`:
- `deploying-demo` — Full deployment workflow (prerequisites, azd up, SRE Agent config, verification)
- `running-demo` — Break/fix scenarios with browser verification (Scenarios 1-4 single-fault, Scenario 5 compound)
- `managing-sre-agent` — Create/manage SRE Agent skills, response plans, knowledge files
## Agent configuration
`scripts/setup-sre-agent.ps1` applies the staged connector Bicep template after
API readiness, then synchronizes skills, named evidence agents with child-specific hooks, response plans, knowledge files,
Microsoft Learn MCP tool enablement, and agent-global custom instructions.
`sre-config/agent-config.json` is authoritative for skill descriptions and
`properties.tools` dependencies, and for connector definitions consumed by the
staged Bicep template. Agent references load `sre-config/agents/` and
embed the one `sre-config/hooks/readonly-evidence.py` source; never make this a
global hook. Keep the agents' useful explicit `ReadFile` base and nonempty skill
selection. Empty tool lists can restore workspace defaults, so successful
Monitor use alone is not proof of skill-owned loading. Keep the existing response
plans and authorized parent remediation intact.
Use `python -B -m unittest discover -s .\tests -p 'test_*.py' -v` from the lab
directory for offline configuration/hook contracts. The
[`configuration verification guide`](docs/skills-agents-validation.md)
requires separate authorization for live changes and is not replaced by local
tests. Private snapshots, traces, and deployment identifiers must not be committed.
The source of truth for custom instructions is
[`sre-config/custom-instructions.md`](sre-config/custom-instructions.md). Keep
that file concise because its contents are applied globally. Detailed incident
procedures belong in the domain or correlation skills.
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
No one has posted yet. Be the first.

