gpu-live-metric-validation
DataDog/datadog-agent/.agents/skills/gpu-live-metric-validation/SKILL.md
Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures. Use when asked to validate GPU metrics on live clusters, check a GPU Agent release, or run dda inv gpu.validate-metrics.
Skill3.8k starsChanged 2 days ago
- Reads credentials
What's in it
- GPU Live Metric Validation
- Purpose
- Required input
- Validation workflow
- Investigating findings
- Candidate investigations by failure type
- Constraints
---
name: gpu-live-metric-validation
description: Validate live GPU metrics on clusters running the Agent version under test and investigate missing metrics or tag failures. Use when asked to validate GPU metrics on live clusters, check a GPU Agent release, or run dda inv gpu.validate-metrics.
---
<!-- @format -->
# GPU Live Metric Validation
## Purpose
Run GPU metric validation against Kubernetes clusters that are exclusively
running the Agent version under test during the selected validation window,
then investigate any findings.
## Required input
Ask the user which Datadog orgs to validate before running any queries. The
supported values match `tasks/gpu.py`:
- `prod` (`app.datadoghq.com`)
- `staging` (`ddstaging.datadoghq.com`)
## Validation workflow
1. Ask the user which supported orgs to validate.
2. Run the validator for each selected org:
```bash
dda inv gpu.validate-metrics \
--org <prod-or-staging> \
--lookback-seconds <window-seconds>
```
By default, the task derives the Agent image-tag wildcard from the current
release branch and its latest release-candidate tag. It passes that wildcard
to the validator, which queries `datadog.agent.running` grouped by
`kube_cluster_name,image_tag`, selects clusters whose nonempty image tags
all match the wildcard, and ANDs that cluster selection with the GPU
configuration filters.
Use `--agent-version <wildcard>` to override the derived version; it is
required outside release branches (`N.N.x`), where derivation fails. Use
`--metric-filter <filter>` only for an additional scope; it is ANDed with
the version-derived cluster filter.
3. Record the selected org, Agent version wildcard, lookback window, and
validation output before interpreting any findings.
## Investigating findings
For follow-up Datadog queries, use `dd-auth` to select the target. `pup`
consumes the injected `DD_API_KEY`, `DD_APP_KEY`, and `DD_SITE`; do not
combine this workflow with `pup --org`.
1. Start from a failing GPU metric and group it by the smallest useful set of
dimensions, normally `gpu_uuid`, `host`, and `kube_cluster_name`. Add the
failed tag as a group-by dimension when investigating a tag failure.
```bash
dd-auth --domain <target-domain> -- \
pup --no-agent metrics query \
--query='count:gpu.<metric>{<validation-filter>} by {gpu_uuid,host,kube_cluster_name,<failed-tag>}' \
--from=<validation-window> --to=now
```
2. Identify patterns before drawing conclusions: whether failures are limited
to a GPU architecture, device mode, GPU model, host, cluster, workload, or
Agent image tag. Confirm the affected cluster's Agent image tag with:
```bash
dd-auth --domain <target-domain> -- \
pup --no-agent metrics query \
--query='sum:datadog.agent.running{<cluster-filter>} by {kube_cluster_name,image_tag}' \
--from=<validation-window> --to=now
```
3. Preserve the exact queries and affected GPU/host/cluster identifiers in
the summary.
## Candidate investigations by failure type
- Missing metric: compare the affected GPU configuration with the metric's
expected support and query `gpu.device.total` for the same scope to confirm
that devices were present.
- Missing required tag: group the failing metric by the required tag and
affected GPU/host/cluster. Compare a related GPU metric to determine whether
the absence is metric-specific.
- Invalid tag value: group by the invalid tag and affected GPU/host/cluster.
Check whether a series has multiple values for the same tag key before
treating a comma-separated value as a single emitted tag value.
- Unknown or extra tag: query a workload/container metric such as
`container.cpu.usage` for the same pod and container. Compare its tags with
the GPU metric to determine whether the tag originates from the workload.
- If a workload tag points to a Kubernetes resource, inspect the corresponding
Kubernetes object and its labels and annotations before assigning a source.
## Constraints
- Use an explicit, bounded validation window.
- Use `dd-auth` to select the target for `pup` queries.
- Do not validate clusters that reported a different nonempty Agent image tag
in the validation window.
- Do not use `pup --org` with `dd-auth`.
More agent context in DataDog/datadog-agent
42 other files this repository gives its agents.
CLAUDE.md
Cursor rule
Skill
- agent-supply-chain-newsletter.agents/skills/agent-supply-chain-newsletter/SKILL.md
- allium.agents/skills/allium/SKILL.md
- auto-jira.agents/skills/auto-jira/SKILL.md
- create-component.agents/skills/create-component/SKILL.md
- create-config-field.agents/skills/create-config-field/SKILL.md
- create-core-check.agents/skills/create-core-check/SKILL.md
- create-epic-recap.agents/skills/create-epic-recap/SKILL.md
- create-go-module.agents/skills/create-go-module/SKILL.md
- create-invoke-task.agents/skills/create-invoke-task/SKILL.md
- create-pr.agents/skills/create-pr/SKILL.md
- create-release-note.agents/skills/create-release-note/SKILL.md
- create-runtime-setting.agents/skills/create-runtime-setting/SKILL.md
- create-status-provider.agents/skills/create-status-provider/SKILL.md
- create-subcommand.agents/skills/create-subcommand/SKILL.md
- cws-btfhub-sync.agents/skills/cws-btfhub-sync/SKILL.md
- cws-iouring-coverage.agents/skills/cws-iouring-coverage/SKILL.md
- e2e-audit.agents/skills/e2e-audit/SKILL.md
- explain-lading-config.agents/skills/explain-lading-config/SKILL.md
- follow-pr.agents/skills/follow-pr/SKILL.md
- handle-pr-ci-failure.agents/skills/handle-pr-ci-failure/SKILL.md
- injector-dev.agents/skills/injector-dev/SKILL.md
- locate-config-setting.agents/skills/locate-config-setting/SKILL.md
- quality-gate-size-analysis.agents/skills/quality-gate-size-analysis/SKILL.md
- review-pr-comments.agents/skills/review-pr-comments/SKILL.md
- run-e2e.agents/skills/run-e2e/SKILL.md
- run-jira.agents/skills/run-jira/SKILL.md
- run-windows-e2e.agents/skills/run-windows-e2e/SKILL.md
- triage-ci-failure.agents/skills/triage-ci-failure/SKILL.md
- update-3rd-party-libs.agents/skills/update-3rd-party-libs/SKILL.md
- update-otel-deps.agents/skills/update-otel-deps/SKILL.md
- write-e2e.agents/skills/write-e2e/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.

